Intelligent learning partner toy interaction method based on large model and intelligent toy
By building a knowledge network system and large language model, capturing user data in real time and analyzing it, the problems of unstable quality of intelligent learning companion toy content generation and limited learning ability are solved, and high-quality, personalized learning experience and strong interactive user interaction are achieved.
Patent Information
- Application Number
- CN202510257254.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing intelligent learning companion toys have unstable content generation quality, limited learning ability, single interactive experience and poor user experience.
By building a knowledge network system in the target field, building a target large language model based on this system, capturing user multi-source data in real time, analyzing and generating answer text, and performing hallucination detection through preset detection models, generating interactive strategies for multimodal feedback.
It improves the quality and relevance of content generation, enhances the personalization and interactivity of learning experience, and enhances user trust and satisfaction.
Smart Images

Figure CN120087466A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of intelligent interaction technologies, and in particular, to an intelligent learning companion toy interaction method and an intelligent toy based on a large model. Background Art
[0002] With the rapid development of artificial intelligence technology, the intelligent toy market is experiencing booming growth. However, there are still many technical defects in current intelligent learning companion toys.
[0003] First of all, the unstable quality of content generation is a major problem. Intelligent toys based on large language models often have hallucination phenomena when generating educational content, that is, the generated content is inconsistent with facts or logically rigorous, which seriously affects the learning effect and user experience.
[0004] Secondly, the learning ability of intelligent toys also appears limited. Most toys rely on fixed knowledge bases and models and lack the ability to dynamically adjust content according to children's learning progress and interests, which limits the effect of personalized learning. At the same time, the existing toy interaction experiences are mostly simple question-and-answer or fixed-mode games, lacking in-depth interaction and personalized guidance, and unable to provide a richer learning experience. Summary of the Invention
[0005] In view of this, embodiments of the present disclosure provide an intelligent learning companion toy interaction method and an intelligent toy based on a large model, which can solve the problems in the prior art such as the inability to ensure the quality of output content, single interaction, and poor user experience.
[0006] In a first aspect, embodiments of the present disclosure provide an intelligent learning companion toy interaction method based on a large model, which specifically includes: constructing a knowledge network system for the target field; the knowledge network system includes various theoretical knowledge, knowledge entities, association relationships between different knowledge entities, and credibility weights corresponding to the association relationships in the target field; constructing a target large language model based on the knowledge network system; real-time capturing multi-source data of the user; the multi-source data includes one or more of user voice commands and user interaction actions; analyzing the multi-source data based on the target large language model to generate a response text; performing hallucination detection on the response text based on a preset detection model, and generating an interaction strategy based on the response text that meets the conditions; performing multi-modal feedback of the intelligent learning companion toy based on the interaction strategy.
[0007] Optionally, analyzing the multi-source data based on the target large language model to generate an answer text includes: preprocessing the multi-source data, and converting the processed data into first text information based on a deep learning model; preprocessing the first text information by using natural language processing techniques to obtain second text information; extracting key features from the second text information; selecting a target large language model from a target database based on the key features; and analyzing the key features based on the target large language model to generate an answer text.
[0008] Optionally, extracting key features from the second text information includes: performing word segmentation on the second text information to obtain a number of sub-words; assigning word labels to each of the sub-words to obtain a number of sub-labels; identifying entity information based on the number of sub-labels, where the entity information includes one or more of a person name, a place name, an organization, and a technical term; obtaining a user intention based on the number of sub-words and the number of sub-labels; and the user intention and the entity information constitute the key features.
[0009] Optionally, analyzing the key features based on the target large language model to generate an answer text includes: creating a user profile based on the key features; determining whether the key features belong to learning-related features. If so, based on the historical interaction records, determining whether the corresponding learning is the first time. If so, analyzing the user profile and the key features based on the target large language model to generate an answer text; if the corresponding learning is not the first time, obtaining the associated learning information in the historical interaction records; generating a personalized learning strategy based on the associated learning information and the user profile, and generating an answer text for the current progress based on the personalized learning strategy; if the key features do not belong to learning-related features, analyzing the user profile and the key features based on the target large language model to generate an answer text.
[0010] Optionally, performing hallucination detection on the answer text based on a preset detection model and generating an interaction strategy based on the answer text that meets the conditions includes: obtaining a number of query word vectors based on the multi-source data; obtaining a number of answer word vectors based on the answer text; analyzing the number of query word vectors and the number of answer word vectors based on a preset strategy to obtain correlation information; analyzing the correlation information based on a preset analysis strategy to obtain a final score; if the final score is not less than a preset score threshold, the answer text does not have hallucinations, and generating an interaction strategy based on the answer text; if the final score is less than the preset score threshold, triggering the execution of a correction strategy, regenerating the answer text based on the correction strategy, performing hallucination detection on the regenerated answer text, and generating an interaction strategy based on the answer text that meets the conditions.
[0011] Optionally, analyzing the relevance information based on a preset analysis strategy to obtain a final score includes: obtaining the initial relevance between each of the query word vectors and the associated answer word vectors; constructing a relevance matrix based on a plurality of the initial relevances; updating the relevance matrix layer by layer based on a preset relevance upward propagation rule; when a preset convergence condition is satisfied, stopping the relevance propagation process and outputting the final relevance matrix to obtain the final score.
[0012] Optionally, constructing a relevance matrix based on a plurality of the initial relevances includes: determining target keywords based on the relevance matrix and the answer word vectors; constructing a target graph with each of the target keywords as a node; establishing connection relationships between different nodes according to semantic similarity and positional relationships, where the connection relationships include connection edges between different nodes and corresponding weights; assigning word embedding vectors to each node and assigning features to each connection edge.
[0013] Optionally, if the final score is less than a preset score threshold, triggering the execution of a correction strategy includes: obtaining the final nodes based on the final relevance matrix; obtaining the connection paths in the relevance matrix; determining the propagation path from the corresponding query node to the answer node based on the final nodes and the connection paths; obtaining node connectivity and relevance scores based on the propagation path, identifying answer word vectors with low relevance or unexplainable ones, and recording the corresponding answer word vectors as risk answer word vectors; obtaining the context information of the risk answer word vectors based on the answer text; obtaining the types of the risk answer word vectors based on the context information, where the types include one or more of false information and irrelevant content; obtaining the corresponding training database and usage database in the target large language model based on the risk answer word vectors, and correcting the training database and the usage database based on the types; retraining the target large language model based on the corrected training database, regenerating the answer text based on the trained target large language model; performing hallucination detection on the regenerated answer text, and generating an interaction strategy based on the answer text that meets the conditions.
[0014] Optionally, analyzing the relevance information based on a preset analysis strategy to obtain a final score includes: transmitting the relevance information layer by layer from the bottom layer to the top layer based on the hierarchical structure of a preset neural network to obtain the relevance score of each layer, and controlling the transmitted content based on a preset gating mechanism; obtaining the final score based on the relevance score of each obtained layer; if the final score is less than a preset score threshold, triggering the execution of a correction strategy, including: if the final score is less than the preset score threshold, obtaining the response word vector corresponding to the minimum relevance score and denoting it as the risk response word vector; obtaining the context information of the risk response word vector based on the response text; obtaining the type of the risk response word vector based on the context information, where the type includes one or more of false information and irrelevant content; obtaining the corresponding training database and usage database in the target large language model based on the risk response word vector, and correcting the training database and the usage database based on the type; retraining the target large language model based on the corrected training database, regenerating the response text based on the trained target large language model; performing hallucination detection on the regenerated response text, and generating an interaction strategy based on the response text that meets the conditions.
[0015] In a second aspect, an embodiment of the present disclosure further provides an intelligent toy, including a total control center and a toy body with a voice interaction function; the total control center integrates the intelligent learning companion toy interaction method based on a large model for controlling the toy body to perform intelligent interaction with a user.
[0016] In a third aspect, an embodiment of the present disclosure further provides an intelligent learning companion toy interaction system based on a large model, including: a knowledge network system construction module for constructing a knowledge network system of a target domain; the knowledge network system includes various theoretical knowledge, knowledge entities, association relationships between different knowledge entities, and credibility weights corresponding to the association relationships in the target domain; a target large language model construction module for constructing a target large language model based on the knowledge network system; a capture module for capturing multi-source data of a user in real time; the multi-source data includes one or more of user voice commands and user interaction actions; a response text generation module for analyzing the multi-source data based on the target large language model to generate a response text; an interaction strategy generation module for performing hallucination detection on the response text based on a preset detection model and generating an interaction strategy based on the response text that meets the conditions; a feedback module for performing multi-modal feedback of the intelligent learning companion toy based on the interaction strategy.
[0017] In a fourth aspect, an embodiment of the present disclosure further provides a computer device, adopting the following technical solution:
[0018] The computer device includes:
[0019] At least one processor; and,
[0020] A memory communicatively connected to the at least one processor; wherein,
[0021] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute any one of the above-described intelligent learning companion toy interaction methods based on a large model.
[0022] In a fifth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium storing computer instructions for causing a computer to execute any one of the above-described intelligent learning companion toy interaction methods based on a large model.
[0023] In a sixth aspect, an embodiment of the present disclosure further provides a computer program product including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of any one of the above-described methods are implemented.
[0024] The intelligent learning companion toy interaction method based on a large model disclosed in the present application, by constructing a detailed target domain knowledge network system, can deeply understand the theoretical knowledge in a specific domain and the relationships between knowledge entities, support accurate understanding of user needs and personalized responses, thereby providing a tailored learning experience, that is, the toy can provide intelligent learning support; the real-time capture of multi-source data enables the toy to timely capture the user's voice instructions and interaction actions, and this instant feedback ability can improve the naturalness and fluency of the interaction, enabling the user to obtain a quick and relevant response, making the interaction more natural and fluent; through the analysis of multi-source data by the target large language model and combined with rich domain knowledge, more accurate and relevant answers can be generated, effectively reducing the problems of vague or irrelevant answers commonly found in traditional interaction systems; after generating the answer text, using a preset detection model for hallucination detection can identify and eliminate inaccurate or false information, and this detection mechanism can effectively ensure that the provided feedback is more credible, thereby improving the user's trust and satisfaction; the multi-modal feedback generated based on the interaction strategy enables the intelligent learning companion to interact in multiple ways such as voice, vision, and touch, and this rich feedback form can enhance the user's immersion and participation, improve the overall learning experience, increase the fun and participation of learning / activities, and provide a customized learning experience; with the continuous interaction of the user and the accumulation of data, the intelligent learning companion can continuously optimize and adjust its knowledge network and interaction strategy, and this adaptive ability enables the system to continuously improve its performance and user experience over time.
[0025] The above description is only an overview of the technical solution of the present disclosure. In order to understand the technical means of the present disclosure more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present disclosure more obvious and understandable, the following preferred embodiments are specifically given and described in detail in conjunction with the accompanying drawings as follows. Description of the Drawings
[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings required to be used in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0027] Figure 1 It is a schematic flow chart of an intelligent learning companion toy interaction method based on a large model provided by an embodiment of the present disclosure.
[0028] Figure 2 It is a schematic flow chart of a method for generating an answer text provided by an embodiment of the present disclosure.
[0029] Figure 3 It is a schematic flow chart of a method for extracting key features from second text information provided by an embodiment of the present disclosure.
[0030] Figure 4 It is a schematic flow chart of a method for generating an answer text provided by an embodiment of the present disclosure.
[0031] Figure 5 It is a schematic flow chart of a method for generating an interaction strategy provided by an embodiment of the present disclosure.
[0032] Figure 6 It is a schematic flow chart of a method for obtaining the final score of the first embodiment provided by an embodiment of the present disclosure.
[0033] Figure 7 It is a schematic flow chart of a method for constructing a correlation matrix provided by an embodiment of the present disclosure.
[0034] Figure 8 It is a schematic flow chart of a method for generating an interaction strategy when the final score is less than a preset score threshold in the first embodiment provided by an embodiment of the present disclosure.
[0035] Figure 9 It is a schematic flow chart of a method for obtaining the final score of the second embodiment provided by an embodiment of the present disclosure.
[0036] Figure 10 It is a schematic flow chart of a method for generating an interaction strategy when the final score is less than a preset score threshold in the second embodiment provided by an embodiment of the present disclosure.
[0037] Figure 11 The structural schematic diagram of a computer device provided by an embodiment of the present disclosure. Detailed implementation manners
[0038] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0039] It should be clear that the following uses specific specific examples to illustrate the implementation manners of the present disclosure, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific implementation manners, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without making creative efforts belong to the scope of protection of the present disclosure.
[0040] It should be noted that the following describes various aspects of embodiments within the scope of the appended claims. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. In addition, this device and / or practice this method can be implemented using other structures and / or functions in addition to one or more of the aspects described herein.
[0041] It should also be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present disclosure in a schematic manner, and only the components related to the present disclosure are shown in the drawings, rather than being drawn according to the number, shape and size of the components in actual implementation. The type, quantity and proportion of each component in its actual implementation can be an arbitrary change, and the component layout type may also be more complex.
[0042] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.
[0043] Referring to Figure 1 , the first aspect of the present application discloses an intelligent learning companion toy interaction method based on a large model, which specifically includes the following steps:
[0044] S100, construct a knowledge network system for the target domain; the knowledge network system includes various theoretical knowledge, knowledge entities, the association relationships between different knowledge entities, and the credibility weights corresponding to the association relationships in the target domain.
[0045] Among them, the knowledge entity is the entity corresponding to the knowledge point.
[0046] For example, when the toy user is a child, the constructed knowledge network system is a knowledge graph covering all aspects of children's education.
[0047] Specifically, the theoretical knowledge of the target domain, such as mathematics, science, or language learning, can be collected through expert systems, academic resources, and educational materials; knowledge graph tools (such as Neo4j) can be used to organize the collected knowledge into a graph, define knowledge entities (such as concepts, terms), the relationships between entities (such as causal relationships, generic relationships), and assign credibility weights to each relationship; the knowledge graph is updated regularly from new materials to ensure its accuracy and comprehensiveness.
[0048] S200, construct a target large language model based on the knowledge network system.
[0049] Specifically, the construction method of the target large language model specifically includes the following steps:
[0050] S210, obtain a target historical database based on the knowledge network system; specifically referring to historical data related to the target domain, including literature, textbooks, case studies, user interaction records, etc. Through the knowledge network system, it is possible to determine which historical data is useful for constructing the model.
[0051] For example, if the user's target domain is "medicine", a database containing medical textbooks, research papers, clinical case records, etc. can be obtained.
[0052] S220, determine an initial model based on the knowledge network system; specifically, before constructing a language model, it is first necessary to select a suitable pre-trained model as the initial model. For example, an open-source pre-trained language model (such as GPT, BERT, etc.) can be selected. These models have been widely trained and have basic language processing capabilities. Then, according to the knowledge structure and requirements in the knowledge network, the initial model is configured so that it can better adapt to the knowledge of the target domain, which can include setting model parameters, selecting appropriate training objectives, etc.
[0053] For example, if the user's target domain is "history", a pre-trained language model can be selected and adjusted according to the characteristics of the history domain (such as eras, events, figures, etc.).
[0054] S230. Train the initial model based on the target historical database, and use the initial model that meets the training requirements as the target large language model.
[0055] By training the model on data in a specific domain, the model will learn the specific knowledge and language patterns of that domain. During the training process, some criteria (such as accuracy, loss function, etc.) can be set to ensure that the training results meet the requirements. When the model performs well on these criteria, it is considered a suitable target large language model.
[0056] Continuing with the "medical" domain as an example, the model can be fine-tuned using a medical-related database (such as medical journal articles). After training, the model can generate accurate medical-related answers, and such a model can be used as the target large language model.
[0057] Systematize the knowledge in the target domain so that the large language model can understand and generate relevant content more accurately. The entities and relationships in the knowledge network provide rich background information, which helps the model generate more in-depth and targeted answers. Through verification and update, the knowledge network can continuously evolve and improve its accuracy.
[0058] In this step, through the knowledge network system, determine and obtain the historical database related to the target domain to ensure that the model can access domain-specific information; select a suitable initial pre-trained model and configure it according to the needs of the target domain to adapt to the domain characteristics; use the historical database to fine-tune the initial model so that it can accurately process and generate content in the target domain, thus becoming the target large language model.
[0059] The large language model can generate more accurate answers for a specific domain, enhancing the user experience, effectively utilizing the rich information in the knowledge network, and improving the knowledge coverage and answer quality of the model.
[0060] S300. Real-time capture the multi-source data of the user; the multi-source data includes one or more of the user's voice commands and user interaction actions.
[0061] Specifically, by integrating a speech recognition system and an action capture system, the user's voice commands and interaction actions (such as gestures, facial expressions, etc.) can be captured in real time. The captured data is pre-processed, such as speech-to-text conversion and action recognition, to ensure the accuracy and integrity of the data. Then, the user's historical data and real-time behavior can be combined to understand the user's context needs.
[0062] By capturing data in real time, a more natural and smooth interaction experience can be provided, accurately understanding the user's needs and context, and thus generating more relevant responses.
[0063] S400 analyzes multi-source data based on a target large language model to generate response text.
[0064] Specifically, the user's multi-source data (voice text, action data) is input into the large language model. The model analyzes the user input based on the knowledge network, understands the user's intention, can generate a text response according to the analysis result, and adjusts it considering the context information.
[0065] In this step, the model can generate intelligent responses that meet the user's needs, can dynamically adjust the response according to the user's actual interaction, and improve the relevance and accuracy of the interaction.
[0066] S500 performs hallucination detection on the response text based on a preset detection model and generates an interaction strategy based on the response text that meets the conditions.
[0067] By performing hallucination detection on the response text, it can be detected whether there is false or inaccurate information in the response text. According to the detection result, the response text that meets the conditions is obtained, and then an interaction strategy is generated based on the response text that meets the conditions to ensure the accuracy and reliability of the response text, reduce misleading information, and adjust the interaction strategy according to the detection result to improve the user experience and satisfaction.
[0068] S600 performs multi-modal feedback on the intelligent learning companion toy based on the interaction strategy.
[0069] Among them, the multi-modal feedback can include one or more of voice output and emotional expression display.
[0070] Specifically, according to the interaction strategy, various modalities such as sound, image, and touch can be used for feedback. For example, the user can be responded to by voice broadcast, screen display, vibration, etc. The feedback method is adjusted according to the user's reaction and interaction situation to ensure the effectiveness and adaptability of the feedback.
[0071] In this step, the multi-modal feedback enhances the user's interaction experience, makes the toy more attractive and engaging, can provide personalized feedback according to the interaction strategy, and enhances the learning effect and user satisfaction.
[0072] If the expression type indicated by the interaction strategy is a smile, control the emotional expression display device set on the intelligent learning companion toy to display a smile expression; if the expression type indicated by the interaction strategy is a wink, control the emotional expression display device set on the intelligent learning companion toy to perform a winking operation; if the expression type indicated by the interaction strategy is a frown, control the emotional expression display device set on the intelligent learning companion toy to display a frown expression; if the expression type indicated by the interaction strategy is a tongue sticking out, control the emotional expression display device set on the intelligent learning companion toy to perform a tongue sticking out operation.
[0073] The intelligent learning companion toy interaction method disclosed in this application can deeply understand the relationship between theoretical knowledge and knowledge entities in a specific domain by constructing a detailed target domain knowledge network system, support accurate understanding of user needs and personalized responses, so as to provide a tailored learning experience, that is, the toy can provide intelligent learning support; the real-time capture of multi-source data enables the toy to timely capture the user's voice commands and interaction actions, and this instant feedback ability can improve the naturalness and fluency of the interaction, enabling the user to obtain quick and relevant responses, making the interaction more natural and fluent; through the analysis of multi-source data by the target large language model and combined with rich domain knowledge, more accurate and relevant answers can be generated, which can effectively reduce the problems of vague or irrelevant answers commonly found in traditional interactive systems; after generating the answer text, using a preset detection model for hallucination detection can identify and eliminate inaccurate or false information, and this detection mechanism can effectively ensure that the provided feedback is more credible, thereby improving the user's trust and satisfaction; the multi-modal feedback generated based on the interaction strategy enables the intelligent learning companion to interact in multiple ways such as voice, vision, and touch. This rich feedback form can enhance the user's immersion and participation, improve the overall learning experience, increase the fun and participation of learning / activities, and provide a customized learning experience; with the continuous interaction of users and the accumulation of data, the intelligent learning companion can continuously optimize and adjust its knowledge network and interaction strategy. This adaptive ability enables the system to continuously improve its performance and user experience over time.
[0074] In summary, this method provides a more intelligent, personalized and reliable learning companion solution through in-depth knowledge network construction, efficient real-time data processing, accurate answer generation, rigorous hallucination detection and rich multi-modal feedback.
[0075] Refer to Figure 2 , the method for generating the answer text specifically includes the following steps:
[0076] S410, preprocess the multi-source data, and convert the processed data into the first text information based on a deep learning model.
[0077] In this step, the preprocessing preferably includes one or more of noise data removal, unified format, and missing value processing.
[0078] Specifically, irrelevant characters and error data can be detected and removed through rules and algorithms (such as regular expressions); for example, delete the advertisement links or expired information in the comment data.
[0079] For unified format, data from different sources can be standardized, such as unifying the date format (YYYY-MM-DD), formatting the phone number, etc.
[0080] For missing value handling, the missing data can be imputed or deleted, and methods such as mean, median, or mode can be used to fill in the missing numerical values.
[0081] Among them, the first text information is the user instruction information, such as a problem or requirement description.
[0082] S420, preprocess the first text information using natural language processing technology to obtain the second text information.
[0083] The preprocessing in this step includes one or more of text cleaning, standardization, and denoising.
[0084] Among them, text cleaning includes removing extra spaces, punctuation marks, and special characters to ensure the cleanliness of the input text.
[0085] For standardization, the text can be converted to lowercase letters to ensure it is not affected by case, for example, unifying "Data" to "data".
[0086] For denoising, some common irrelevant words (such as "um", "ah", excessive conjunctions, etc.) can be identified and removed to reduce unnecessary noise in the text.
[0087] S430, extract key features from the second text information.
[0088] Specifically, methods such as TF-IDF and LDA topic modeling can be used to identify the most important words and phrases from the text; NLP libraries (such as spaCy, NLTK) can be used to identify names, places, and professional terms in the text; the emotional tendency contained in the text can be analyzed to provide context information for subsequent answer generation.
[0089] In this step, by extracting key features, subsequent analysis can focus on the most important information. Key features can provide context to help the model understand the user's true intention, and the extraction of important features can effectively simplify the data structure of subsequent processing.
[0090] S440, select a target large language model from the target database based on the key features.
[0091] Specifically, different models stored in the database are evaluated, and the model with the best performance is selected according to indicators such as accuracy and recall rate of the model, and the model that best matches the target features is selected, such as a dedicated model for specific industry applications (such as medical, financial).
[0092] In this step, accurately selecting the model can accelerate the processing process, ensuring a quick response to user needs; selecting the most suitable model based on key features provides more accurate answers, and choosing a dedicated model can effectively improve the processing ability for specific domain questions.
[0093] S450. Analyze the key features based on the target large language model to generate the response text.
[0094] Specifically, the extracted key features are passed as input to the selected large language model, and the model generates the corresponding response text according to the input features, providing a structured answer or further explanation.
[0095] Furthermore, the generated response can be formatted and verified to ensure that the text is fluent and meets user needs.
[0096] In this embodiment, each step has clear goals and tasks, standardizing the whole process and enhancing repeatability; through preprocessing and feature extraction, the efficiency and accuracy of subsequent analysis are significantly improved; based on features for model selection, computing resources can be optimized and the response speed of the model can be increased; the generated high-quality response text improves user satisfaction and optimizes the interaction process. This solution, through scientific and reasonable step design, from data preprocessing to response generation, ensures the efficiency, accuracy and user-friendliness of the process, and can greatly meet user needs and improve the processing effect.
[0097] Refer to Figure 3 , the method for extracting key features from the second text information includes the following steps:
[0098] S431. Perform word segmentation on the second text information to obtain a number of sub-words.
[0099] Specifically, based on word segmentation tools (such as word_tokenize of NLTK, spacy.load('en_core_web_sm').tokenizer of SpaCy), perform word segmentation on the second text information, that is, split the second text information into individual words or phrases to obtain a number of sub-words. For example, when the second text information is "I want to learn the basic knowledge of Python language", after word segmentation, the obtained number of sub-words is: ['I', 'want', 'to learn', 'Python language', 'of', 'basic knowledge'].
[0100] Splitting the long text into sub-words makes subsequent processing more precise, and the sub-words after word segmentation are convenient for further analysis and processing.
[0101] S432. Assign a lexical label to each sub-word to obtain a number of sub-labels.
[0102] Specifically, a part-of-speech tagging tool (such as NLTK's pos_tag, SpaCy's spacy.load('en_core_web_sm').tag_) is used to assign part-of-speech tags to each word. For example: the part-of-speech tagged text after word segmentation is: ['I' (pronoun), 'want' (verb), 'learn' (verb), 'Python language' (noun phrase), 'of' (auxiliary word), 'basic knowledge' (noun)]. Word tagging can clarify the role and function of each word in a sentence, and through word tags, the semantic structure and content of the text can be better understood.
[0103] S433. Based on a number of sub-tags, entity information is identified.
[0104] Among them, the entity information includes one or more of personal names, place names, organizations, and technical terms.
[0105] NER tools (such as SpaCy's spacy.load('en_core_web_sm').ents, NLTK's nltk.chunk.ne_chunk) can be used to identify entities (such as personal names, place names, organizations, technical terms, etc.) in the text.
[0106] For example: identifying "Python language" as an entity from the text can be marked as "course" or "technology".
[0107] S434. Based on a number of sub-words and a number of sub-tags, the user's intention is obtained; the user's intention and entity information constitute key features.
[0108] In a preferred embodiment, a dependency parsing tool (such as SpaCy's spacy.load('en_core_web_sm').dep_, the dependency parser of Stanford NLP) can be used to analyze the syntactic relationships between words.
[0109] For example: analyzing the syntactic structure of "I want to learn the basic knowledge of Python language", determining that "want" is the predicate verb of the subject-predicate structure, "learn" is the object of the verb, and "the basic knowledge of Python language" is the object complement of the verb, that is, the identified user intention is "learn", and the extracted entity information "Python language" is the course objective.
[0110] By identifying the user's intention, the user's needs and expectations can be accurately grasped. A clear user intention helps generate a response text that better meets the needs. Intention recognition can help the system dynamically adjust the interaction method and better respond to the user's needs.
[0111] In this embodiment, each step defines the tasks to ensure the integrity and accuracy of the feature extraction process; through word segmentation, tagging, entity recognition, and intent analysis, the system can deeply understand the text content and improve its ability to handle complex queries; clear user intents and structured entity information help generate more accurate and relevant response texts; clear feature extraction and intent analysis enhance the intelligence of the system in conversations, making the user interaction experience more natural and smooth; by accurately extracting key features, computing resources can be utilized more efficiently, improving the system response speed and processing ability.
[0112] By implementing these steps, the overall solution can effectively extract valuable features from the text information, providing a reliable basis for subsequent response generation and enhancing the user experience and system performance.
[0113] Refer to Figure 4 , the method for generating response text specifically includes the following steps:
[0114] Create a user profile based on the key features;
[0115] Determine whether the key features belong to learning-related features. If so, obtain the historical interaction records;
[0116] Based on the historical interaction records, determine whether the corresponding learning is the first time. If so, analyze the user profile and key features based on the target large language model to generate the response text;
[0117] If the corresponding learning is not the first time, obtain the associated learning information in the historical interaction records;
[0118] Generate a personalized learning strategy based on the associated learning information and user profile, and generate the response text for the current progress based on the personalized learning strategy;
[0119] If the key features do not belong to learning-related features, analyze the user profile and key features based on the target large language model to generate the response text.
[0120] Among them, the user profile includes user age, interests, learning style, learning emotional state, historical behaviors, past learning records, degree of attention to a certain field, etc.; the user profile can provide an in-depth understanding of user needs and interests, helping to provide services that better meet user expectations. For example, the key feature is "Python programming"; the user profile can include the user's basic information (age, occupation), learning interests (programming, data analysis), and learning methods (video tutorials, practical projects).
[0121] When the key feature belongs to the learning type of feature, for example, the feature "Python programming" obviously belongs to the learning type of feature. Obtaining the historical interaction record includes retrieving the user's historical learning record from the database to understand the user's previous learning activities; then judging whether the corresponding learning is the first time based on the historical interaction record, that is, checking whether the user has learned similar content; for example, the user's historical record shows that the user has learned "Python Basics" before, but has not learned "Python programming". Since the user has not learned "Python programming" before, it is considered as the first learning.
[0122] By judging whether it belongs to the learning type of feature, it is possible to distinguish the user's learning needs from other types of needs, so as to provide targeted services, which helps to avoid duplicate learning content, save the user's time, and improve learning efficiency.
[0123] If it is the first time, use the target large language model to analyze the user profile and the key feature. Based on the analysis results, generate learning suggestions or response texts suitable for the user. For example, the target large language model analyzes the user profile and the feature, and concludes that the user is interested in Python programming, and generates a response text: "Hello! To help you quickly get started with Python programming, we suggest that you start learning from basic syntax, data types, and control structures. Here are some high-quality courses and resources recommended for you."
[0124] The response text for the first learning can provide personalized learning materials or suggestions according to the user profile; customizing content according to the user's interests and needs helps to improve the learning effect and user satisfaction.
[0125] If it is not the first time, obtain the relevant learning information in the user's historical record, such as the courses studied before, the projects completed, etc.; design a personalized learning strategy based on the historical record and the user profile. For example: The user has learned "Python Basics" and completed some small projects, and it is recommended that the user continue to study more advanced topics, such as "Python Data Analysis" or "Machine Learning Basics".
[0126] If the corresponding learning is not the first time, obtaining the associated learning information in the historical interaction record can track the user's learning progress and historical performance, and provide continuous support; using historical data to generate more targeted feedback and suggestions to help the user further learn based on the existing knowledge.
[0127] Combining the historical learning information and the user profile, use the target large language model to generate personalized response texts, such as learning suggestions or learning guidance.
[0128] Providing customized strategies based on the user's past learning helps improve learning efficiency and effectiveness, can adjust the learning plan according to the user's progress and needs, and enhances the pertinence and practicality of learning.
[0129] If the key features do not belong to learning features (such as querying the weather, getting news, etc.), directly generate the response text based on the user profile and features. For example, when the user inputs "I want to know today's weather", use the target large language model to generate the response: "Hello! Today's weather forecast shows that the temperature will be between 20°C and 25°C, mainly sunny, suitable for outdoor activities."
[0130] In this embodiment, creating a user profile based on the key features enables the response to reflect the user's personalized needs and interests, thereby improving the relevance and satisfaction of the response; understanding the user's learning progress and interest changes through historical interaction records makes the suggestions more in line with the user's actual needs.
[0131] For non-learning features (such as consultation questions), directly use the target large language model to analyze the user profile and key features to generate relevant response text. For example, "Based on your interests and hobbies, the following are some activity recommendations that you may be interested in." The response generated by the model can include suggestions, information query results, or other relevant content, ensuring that no matter what type of user needs are, effective responses can be provided, while providing diverse responses to meet the different needs of users and enhancing the overall user experience.
[0132] For the first-time learning situation, by analyzing the user profile and key features through the large language model, targeted initial learning suggestions can be generated to help users quickly get started. For non-first-time learning processing, based on the user's existing learning foundation, advanced suggestions are provided based on historical records and associated learning information to avoid repetitive content and improve learning efficiency. The generated personalized learning strategies make users feel the attention and understanding of the system, thereby increasing user participation and learning motivation; adjusting the suggestions according to the user's learning progress and historical records maintains the user's learning interest and motivation.
[0133] Through detailed analysis of the user profile and historical records, accurate responses can be provided, avoiding irrelevant or general suggestions; generating responses based on the user's actual needs and background enhances the user's satisfaction and trust in the system.
[0134] By generating responses automatically through systematic steps and the large language model, it saves the time and effort of manual processing, improves the response speed, automatically generates personalized learning strategies, helps users efficiently find the most suitable learning resources and paths, and avoids the cumbersome process of manual screening.
[0135] By analyzing the user's historical records and feedback, continuously optimize the user profile and recommendation strategy, enabling the system to continuously improve its service quality over time, and be able to evaluate and adjust the generated suggestions according to the user's learning effect, further enhancing the accuracy and practicality of the answers.
[0136] The advantage of this answer text generation method is that it can provide personalized, accurate, and effective answers, enhancing the user experience, improving learning efficiency, and saving time and energy. This method ensures the relevance and pertinence of the answers through systematic steps and data-driven analysis, thus meeting the specific needs of users.
[0137] Specifically, the personalized learning strategy includes a learning plan that matches the user; for example, when the user learns the basic knowledge of Python through the platform, the personalized learning strategy includes a plan of corresponding difficulty set according to the user's answering accuracy rate and learning speed. If the user performs well on a certain knowledge point, the personalized learning strategy will include more advanced content to accelerate the learning progress; if encountering difficulties, the personalized learning strategy will reduce the difficulty to ensure that the user can understand and master.
[0138] Refer to Figure 5 , the generation method of the interaction strategy specifically includes the following steps:
[0139] S510, Based on multi-source data, obtain a number of query word vectors.
[0140] Specifically, based on the second text information, extract a number of query keywords, and process the number of query keywords based on a pre-trained word embedding model to obtain a number of query word vectors.
[0141] For example, if the query is "How to use graph neural networks?", the query keywords include: "How", "use", "graph neural networks".
[0142] The pre-trained word embedding model includes one or more of Word2Vec, BERT. Through the pre-trained word embedding model, a fixed vector can be generated for each word, enabling the model to use these vectors for various downstream tasks, such as text classification, entity recognition, etc.
[0143] S520, Based on the answer text, obtain a number of answer word vectors.
[0144] Specifically, based on the answer text, extract a number of answer keywords, and process the number of answer keywords based on a pre-trained word embedding model to obtain a number of answer word vectors.
[0145] Convert both the query and the answer into word vectors, which is convenient for comparison and analysis in a unified vector space. The word vectors can retain the semantic information of the text, helping to accurately evaluate the relevance between the query and the answer.
[0146] S530 analyzes a number of query word vectors and a number of answer word vectors based on a preset strategy to obtain correlation information.
[0147] The correlation information includes the relative importance or degree of correlation of each part.
[0148] Specifically, the preset strategy includes: obtaining correlation information based on the attention mechanism.
[0149] Furthermore, the cosine similarity method or the dot product method can be used to calculate the correlation information.
[0150] For example, when processing tasks such as a question-and-answer system, the attention mechanism can help the model calculate the correlation between each part of the "query" (i.e., the user's question) and the "answer" (i.e., the answer generated by the model). For example, if the query is "How do I install Python?", then the attention mechanism will help the model identify which parts of the answer (such as "install" or "Python") are most relevant to the query.
[0151] Specifically, the cosine similarity is a metric for measuring the similarity between two vectors. It calculates the cosine value of the angle between two vectors to quantify their degree of similarity. The cosine similarity value ranges from -1 to 1, where 1 indicates complete similarity, -1 indicates complete opposition, and 0 indicates no similarity.
[0152] The dot product is another method for calculating correlation. It calculates the inner product of two vectors. The larger the dot product value, the stronger the correlation between the two vectors. In some models, the dot product is used to measure the matching degree between two vectors.
[0153] S540 analyzes the correlation information based on a preset analysis strategy to obtain a final score.
[0154] S550, if the final score is not less than the preset score threshold, the answer text has no hallucination, and an interaction strategy is generated based on the answer text.
[0155] If the final score is not lower than the threshold, it is considered that the answer text is highly relevant to the query and has no hallucination. An appropriate interaction strategy is generated based on this answer text, such as asking further questions, providing relevant information, etc.; generating an interaction strategy based on the highly relevant answer text can improve user satisfaction and engagement.
[0156] S560, if the final score is less than the preset score threshold, trigger the execution of a correction strategy, regenerate the answer text based on the correction strategy, perform hallucination detection on the regenerated answer text, and generate an interaction strategy based on the answer text that meets the conditions.
[0157] Through the correction strategy and re-detection mechanism, it is possible to self-correct low-quality response texts; the response texts after correction and re-detection are more accurate, which helps to generate more effective interaction strategies.
[0158] By setting a preset score threshold, identify low-correlation content (i.e., possible hallucinations), and implement a mechanism for correcting or regenerating content.
[0159] In this embodiment, through hallucination detection and correction strategies, it is possible to screen out high-quality response texts and avoid misleading users; generating interaction strategies based on highly relevant response texts can improve user satisfaction and engagement; the system has the ability of self-correction and re-detection, which can continuously improve the accuracy and relevance of response texts; parameters such as the similarity calculation method and score threshold can be adjusted according to different application scenarios and requirements, making the system more flexible and adaptable.
[0160] Refer to Figure 6 , in the first embodiment, analyze the relevant information based on a preset analysis strategy to obtain a final score. That is, the method for obtaining the final score specifically includes:
[0161] S541, obtain the initial correlation between each query word vector and its associated response word vector.
[0162] Specifically, the initial correlation between each query word vector and its associated response word vector can be calculated using cosine similarity, Euclidean distance, or other correlation measurement methods. For example, for each pair of query word vectors and response word vectors, calculate their cosine similarity and record these initial correlation values.
[0163] The initial correlation calculation provides a preliminary matching degree between the query and the response, which helps to understand their direct relationship and lays a data foundation for subsequent construction of the correlation matrix and application of the propagation rules.
[0164] S542, construct a correlation matrix based on a number of initial correlations.
[0165] Among them, the correlation matrix can be a graph structure.
[0166] Specifically, organize all the calculated initial correlation values into a correlation matrix. The rows and columns of the matrix represent the query word vectors and response word vectors respectively, and the values in the matrix represent the initial correlations between these word vectors. The correlation matrix organizes the initial correlations in a structured way, facilitating subsequent analysis and processing. The matrix-form data is convenient for further mathematical operations and correlation propagation calculations.
[0167] S543, update the correlation matrix layer by layer based on the preset upward correlation propagation rules.
[0168] Specifically, applying the relevance upward propagation rules (e.g., propagation algorithms based on graph theory such as PageRank, propagation algorithms, etc.), the relevance values are updated layer by layer in the relevance matrix. These rules utilize the existing relevance information in the matrix to gradually propagate and update the values in the matrix, reflecting the indirect relationships between word vectors.
[0169] Through the propagation rules, the indirect relevance between the query and the answer can be captured, thereby obtaining more comprehensive relevance information. Layer-by-layer update can eliminate the noise and incompleteness in the initial calculation and improve the accuracy of the final relevance evaluation.
[0170] S544. When the preset convergence condition is met, stop the relevance propagation process and output the final relevance matrix to obtain the final score.
[0171] By stopping the propagation by meeting the convergence condition, it can be ensured that the final matrix and score reflect stable relevance information, avoiding the errors caused by overpropagation; the setting of the convergence condition can effectively control the computational complexity, reduce unnecessary calculations, and optimize the overall processing efficiency.
[0172] Specifically, the preset convergence conditions include: setting a threshold, when the change of each element in the relevance matrix is less than this threshold, stop the propagation; or, setting the maximum number of iterations, when this number is reached, even if the convergence condition is not met, stop the propagation.
[0173] Refer to Figure 7 , the construction method of the relevance matrix includes:
[0174] S5421. Based on the relevance matrix and the answer word vector, determine the target keywords.
[0175] Specifically, screen out the keywords with high relevance from the relevance matrix. These keywords are usually the most relevant words to the query term. Combining the characteristics of the answer word vector, determine the most important target keywords. These keywords play a key role in the answer.
[0176] By determining the target keywords, it is possible to focus on the words that have a key impact on the answer quality, thereby simplifying the subsequent processing; selecting the most relevant keywords as the target helps to improve the accuracy and effectiveness of the relevance analysis.
[0177] S5422. Construct a target graph with each target keyword as a node.
[0178] Specifically, regard each target keyword as a node in the graph and create a target graph, where each node represents a target keyword.
[0179] Representing target keywords through a graph structure can visually display the relationships between keywords, facilitating subsequent analysis. The graph structure is conducive to applying various algorithms in graph theory, such as the shortest path and centrality analysis, to deeply understand the relationships between keywords.
[0180] S5423. Establish connection relationships between different nodes according to semantic similarity and positional relationship. The connection relationship includes the connection edges between different nodes and the corresponding weights.
[0181] Specifically, calculate the semantic similarity (such as calculating the cosine similarity using word vectors) and positional relationship (such as the relative position in the answer text) between target keywords. Based on the calculation results, establish connection edges in the target graph and assign weights to these edges to represent the correlation strength between nodes.
[0182] By combining semantic similarity and positional relationship, the relationships between target keywords can be comprehensively captured, reflecting their semantic and contextual relevance. The weights of the edges can quantify the connection strength, providing more abundant information, which is helpful for subsequent correlation analysis and propagation.
[0183] S5424. Assign word embedding vectors to each node and assign features to each connection edge.
[0184] Specifically, assign word embedding vectors to each target keyword node. These vectors represent the semantic features of the nodes; assign features (such as edge weights, similarities, etc.) to each connection edge to represent the strength and nature of the connection.
[0185] The word embedding vectors provide detailed semantic representations for the nodes, helping to more accurately understand the meanings of the nodes; the edge features provide more detailed information about the connections, making the graph analysis more comprehensive and able to better reflect the relationships between nodes.
[0186] In this embodiment, through the construction of the target graph and the assignment of features of nodes and edges, the relationships of target keywords can be represented and analyzed in a structured manner; the word embedding vectors and edge features provide rich semantic and relationship information, which is helpful for more accurately evaluating the correlation; the target graph can visually display the relationships between keywords, making the analysis and understanding of the interactions of target keywords clearer; by combining semantic similarity and positional relationship to establish the weights of connection edges, the actual relationships between target keywords can be more comprehensively captured, enhancing the effect of the final correlation analysis. This graph-based representation method provides an efficient and comprehensive way to process and analyze the correlation between keywords, laying a solid foundation for subsequent correlation propagation and final score calculation.
[0187] Refer to Figure 8, in the first embodiment, if the final score is less than the preset score threshold, the correction strategy is triggered to be executed. Specifically, the method for generating the interaction strategy includes:
[0188] A110, obtaining the final nodes based on the final correlation matrix.
[0189] Specifically, according to the final correlation matrix, the key nodes in the graph are determined, and these nodes play a core role in node connection and propagation paths.
[0190] Identifying the final key nodes can focus on dealing with those parts that play a decisive role in the answer quality and improve the processing efficiency.
[0191] A120, obtaining the connection paths in the correlation matrix.
[0192] Specifically, all valid connection paths are extracted from the correlation matrix, and these paths show the relationships between nodes and information flow.
[0193] Obtaining the connection paths can help to comprehensively understand the relationships between nodes and the routes of information propagation, and contribute to analyzing potential problems.
[0194] A130, determining the propagation path from the corresponding query node to the answer node based on the final nodes and connection paths.
[0195] Specifically, based on the final nodes, the connection paths in the correlation matrix are obtained, and the propagation paths of information from the query node to the answer node are found, as well as the contributions of these paths to the final correlation score, that is, calculating the correlation between the query and the answer based on the representation of the final nodes.
[0196] By analyzing the final nodes and connection paths, determining the propagation path from the query node to the answer node, and identifying the specific route of the information flow, it is possible to accurately trace the propagation path of information from the query to the answer, which helps to locate the position and cause of the problem.
[0197] A140, obtaining the node connectivity and correlation scores based on the propagation path, identifying the answer word vectors with low correlation or unexplainable ones, and recording the corresponding answer word vectors as risk answer word vectors.
[0198] Evaluating the node connectivity and correlation scores on the propagation path, identifying the nodes with low correlation or not meeting expectations, and marking the answer word vectors in these nodes as risk answer word vectors.
[0199] Through the analysis of connectivity and correlation scores, it is possible to identify the low-quality or unexpected answer word vectors, which helps to accurately locate the problem.
[0200] A150, obtaining the context information of the risk answer word vectors based on the answer text.
[0201] Extract the context information of the risk response word vector from the response text so as to further analyze its questions. Understanding the context information of the risk response word vector can help to more accurately locate the problems and provide a basis for subsequent corrections.
[0202] A160. Obtain the type of the risk response word vector based on the context information. The types include one or more of false information and irrelevant content.
[0203] Identifying the type of the risk word vector can perform targeted correction processing, improving the efficiency and effectiveness of the correction strategy.
[0204] A170. Obtain the corresponding training database and usage database in the target large language model based on the risk response word vector, and correct the training database and usage database based on the type.
[0205] According to the identified type, adjust the training database and usage database of the target large language model, and correct the false information or irrelevant content therein. By updating and correcting the training database, improve the training quality of the large language model and reduce the occurrence of similar problems.
[0206] A180. Retrain the target large language model based on the corrected training database, and regenerate the response text based on the trained target large language model.
[0207] The retrained model can better identify and generate high-quality answers, reduce the generation of hallucinations and incorrect information, and the generated response text will be more accurate and relevant, improving the overall response quality.
[0208] A190. Perform hallucination detection on the regenerated response text, and generate an interaction strategy based on the response text that meets the conditions.
[0209] Perform hallucination detection on the newly generated response text to ensure that it does not contain false or inaccurate information; generate corresponding interaction strategies based on the response text that meets the conditions, such as asking further questions or providing supplementary information.
[0210] Through hallucination detection, ensure that the finally generated response text meets the high-quality standard; generate an interaction strategy based on the high-quality response text to provide accurate and useful responses and improve the user experience.
[0211] In this embodiment, by identifying and correcting risk answer word vectors, the accuracy and relevance of the answer text can be comprehensively improved, false information or irrelevant content in the answer can be classified and corrected, ensuring that the generated answer meets the requirements of authenticity and relevance, updating the training database and retraining the model, improving the overall performance of the large language model, reducing the occurrence of similar problems, and improving user experience and satisfaction by generating high-quality answer text and interactive strategies.
[0212] Reference Figure 9 In the second embodiment, the correlation information is analyzed based on a preset analysis strategy to obtain a final score, that is, the method for obtaining the final score specifically includes:
[0213] A210, based on the hierarchical structure of a preset neural network, transmits relevance information layer by layer from the bottom layer to the top layer, obtains the relevance score of each layer, and controls the transmitted content based on the preset gating mechanism.
[0214] The layer-by-layer transmission includes: processing at each layer and transmitting the processed information to the upper layer; the processing includes one or more of feature transformation (linear transformation, nonlinear transformation), feature extraction, feature fusion, and feature recombination.
[0215] Specifically, the relevance information is transmitted layer by layer from the bottom layer to the top layer, including: the bottom-level neural network is responsible for the preliminary processing of the input data, extracting basic features and initial relevance scores; the middle-level neural network further integrates the bottom-level information, refines and enhances the relevance score, and combines more contextual information; the top-level neural network performs high-level abstraction and synthesis, integrates information from all levels, and obtains the final relevance score.
[0216] Controlling the transmission content based on the preset gating mechanism includes: applying the gating mechanism in each layer to control the transmission of information to ensure that important information is retained and irrelevant or noise information is suppressed. This can be achieved through the gating units in the neural network (such as the gating structure of LSTM); the gating mechanism can dynamically adjust the weight of information transmission and optimize the transmission process based on the current input and context information.
[0217] A220, based on the relevance scores obtained for each layer, a final score is obtained.
[0218] The relevance scores of each layer are aggregated, and the information of each layer is combined to calculate the final relevance score, which reflects the final evaluation of the input data by the entire network.
[0219] Furthermore, the obtained relevance scores of each layer can be analyzed, that is, hallucination detection and correction can be performed based on a layer-by-layer propagation mechanism to ensure the accuracy of the generated content.
[0220] Specifically, by analyzing the relevance scores of each layer layer by layer, if the score is less than the preset score value, external knowledge can be introduced to detect and correct the information hallucinations in the generated content, thereby improving the quality and reliability of the generated content.
[0221] In this embodiment, the layer-by-layer processing from the bottom layer to the top layer allows the model to gradually refine and enhance the relevant information in each layer, thereby capturing more complex and high-level relationships; neural networks at different levels can extract features at different levels, making the final score more comprehensive and accurate.
[0222] By transmitting and integrating information layer by layer, the features of each layer can be more effectively synthesized, enhancing the model's ability to analyze relevance; the hierarchical structure and gating mechanism enable the model to capture complex, non-linear relevance relationships, improving the depth and breadth of analysis.
[0223] The gating mechanism helps to suppress irrelevant information and noise, ensuring that important relevant information is transmitted and retained; dynamically adjusting the weight of information transmission according to different inputs and contexts improves the accuracy and flexibility of analysis.
[0224] Based on the hierarchical scoring method, multi-level information can be combined to generate a more accurate final relevance score, reflecting the true relevance of the data; the information transmitted and integrated layer by layer helps to improve the consistency and reliability of the final score, reducing the bias that may be brought by a single layer.
[0225] The hierarchical structure and gating mechanism endow the model with good flexibility, enabling it to be adjusted and optimized according to different application scenarios and requirements, capable of processing complex data structures and relationships, adapting to more diverse application scenarios, and providing efficient relevance analysis and scoring.
[0226] Transmitting the relevance information layer by layer based on the hierarchical structure of the preset neural network and controlling the transmitted content through the gating mechanism can significantly improve the accuracy, comprehensiveness, and efficiency of relevance analysis. The hierarchical processing and dynamic adjustment mechanism enable the model to better capture complex relevance relationships, providing a more reliable and consistent final score, thereby supporting more accurate decision-making and analysis.
[0227] Refer to Figure 10 , in the second embodiment, if the final score is less than the preset score threshold, a correction strategy is triggered, that is, the generation method of the interaction strategy includes:
[0228] B110, if the final score is less than the preset score threshold, obtain the answer word vector corresponding to the minimum relevance score and denote it as the risk answer word vector.
[0229] From the final analysis result, identify the answer word vector with the lowest relevance score and mark it as the risk answer word vector.
[0230] Directly processing the word vectors with the lowest scores can quickly focus on the parts most likely to have problems, improve processing efficiency, and the identified risky response word vectors can directly guide subsequent correction and optimization strategies, making the improvement measures more targeted.
[0231] B120, obtaining the context information of the risky response word vectors based on the response text.
[0232] Extract the context information containing the risky response word vectors from the original response text to understand their roles and positions in the overall response.
[0233] The context information helps to deeply understand the actual usage scenarios of the risky response word vectors, identify their specific problem sources; provides detailed context information, which helps to formulate improvement strategies more accurately.
[0234] B130, obtaining the types of the risky response word vectors based on the context information, and the types include one or more of false information and irrelevant content.
[0235] Clarifying the types of the risk word vectors can solve specific problems targeted, improve the accuracy of correction. Different types of problems require different processing strategies, and classification helps to formulate specific improvement measures.
[0236] B140, obtaining the corresponding training database and usage database in the target large language model based on the risky response word vectors, and correcting the training database and usage database based on the types.
[0237] Correcting the training database and usage database can improve the quality of the data, reduce the occurrence of similar problems. Data correction directly affects the training effect of the model, and helps to improve the accuracy and relevance of the answers generated by the model.
[0238] The retrained model can better process high-quality data, improve the accuracy and reliability of the generated answers. The updated training data reduces false information and irrelevant content, and improves the overall performance of the model.
[0239] B150, retraining the target large language model based on the corrected training database, and regenerating the response text based on the trained target large language model.
[0240] The generated response text will be more accurate and relevant, reducing the problems existing in the previous model. High-quality answers improve user satisfaction and trust.
[0241] B160, performing hallucination detection on the regenerated response text, and generating an interaction strategy based on the response text that meets the conditions.
[0242] Hallucination detection ensures that the final answer text does not contain false or inaccurate information, improving the credibility of the answer; formulating an interaction strategy based on the detection results can further improve the user experience by providing targeted supplementary information or clarification.
[0243] In this embodiment, by gradually identifying and correcting the risk answer word vectors, the performance of the model can be systematically improved, and the answer quality can be enhanced; processing the answer word vectors with the minimum correlation score ensures the efficient positioning and solution of problems; classifying and correcting the types of risk word vectors makes the improvement measures more accurate and can effectively enhance the overall performance of the model; updating and correcting the training database reduces false information and irrelevant content, improving the quality of data and the accuracy of the model; by generating high-quality answer text and formulating effective interaction strategies, the user experience and satisfaction are improved. This detailed correction strategy ensures a significant improvement in the quality of the answer text generated by the model through a systematic method, thereby enhancing the overall application effect and user experience.
[0244] For example, when the toy is explaining the concept of "photosynthesis", if the generated content falsely claims that "plants perform photosynthesis at night", the layer-by-layer relevance propagation detection system will quickly identify this hallucination. The system will trace the relationships of the concepts "photosynthesis", "plants", and "time" in the knowledge graph and find that there is a low correlation between "night" and the occurrence time of "photosynthesis", thereby marking and correcting this error to ensure the transmission of accurate scientific knowledge to children.
[0245] The following describes how to use the BERT model to implement a system for layer-by-layer relevance propagation and detecting hallucination content in combination with examples. The following is a detailed explanation of each step and its purpose:
[0246] Step A1: Prepare data;
[0247] Step A1.1: Load the pre-trained BERT model; Purpose: Load the pre-trained BERT model, which has powerful language understanding capabilities and can be used to process the input queries and generated answers. Implementation: Usually use a library (such as the Transformers library) to load the BERT model and its corresponding weights.
[0248] Step A1.2: Tokenize the input queries and generated answers; Purpose: Convert natural language text (queries and answers) into a format that the model can process. The BERT model uses the WordPiece tokenizer to split the text into sub-word units. Implementation: Use the BERT tokenizer to tokenize the input text.
[0249] Step A1.3: Convert the word segmentation results into the input format of BERT (including token_ids, segment_ids, and attention_mask); Purpose: Convert the word segmentation results into the input format required by the BERT model, which includes: token_ids, which is the ID representation of the segmented text, segment_ids, which is used to identify the segmentation of sentences (for handling question-answer pairs), and attention_mask, which is used to identify which tokens are actual inputs and which are padding parts. Implementation: Use the tools of BERT to convert the word segmentation results into the corresponding input format.
[0250] Step A2: Construct a relevance propagation layer;
[0251] Step A2.1: Use the output of each layer of BERT as the basis for relevance propagation; Purpose: Utilize the outputs of each layer of the BERT model as the basic data for relevance propagation. Each layer of BERT extracts different levels of text features; Implementation: Obtain the output results of each layer of the BERT model.
[0252] Step A2.2: Implement a custom attention mechanism to calculate the relevance between the query and the answer; Purpose: Calculate the relevance between the query and the answer through a custom attention mechanism to track the information flow during the layer-by-layer propagation process; Implementation: Design a mechanism to calculate the relevance score between the query and the answer using the output of each layer.
[0253] Step A2.3: Design a gated update unit to control the propagation of relevance information; Purpose: Use a gating mechanism to control the propagation of relevance information and determine which information needs to be retained and which needs to be suppressed; Implementation: Design and implement a gating mechanism that can dynamically adjust the information transmission according to the context and relevance scores.
[0254] Step A3: Implement a layer-by-layer propagation algorithm;
[0255] Step A3.1: Start from the bottom layer of BERT and calculate and update the relevance layer by layer; Purpose: Start from the bottom layer of the BERT model, process the relevance information layer by layer, and update the relevance scores of each layer; Implementation: Process each layer of the BERT model in sequence and calculate the relevance layer by layer.
[0256] Step A3.2: Use the attention mechanism and the gated update unit at each layer; Purpose: Apply the attention mechanism and the gated update unit at each layer to ensure the effective propagation of information; Implementation: Apply the attention mechanism to calculate the relevance at each layer and use the gated update unit to adjust the information transmission.
[0257] Step A3.3: Aggregate the relevance results of each layer to obtain the final relevance matrix; Purpose: To synthesize the relevance results of each layer to obtain an overall relevance matrix that reflects the relevance of the entire model to the input; Implementation: Aggregate the relevance scores of each layer to generate the final relevance matrix.
[0258] Step A4: Detect hallucinations;
[0259] Step A4.1: Analyze the final relevance matrix to identify low-relevance regions; Purpose: To analyze the low-relevance regions in the relevance matrix, which usually correspond to a relatively low relevance between the answers generated by the model and the queries. Low-relevance regions may indicate a weak connection between the answers generated by the model and the input queries, which may be a manifestation of hallucinations; Implementation: Through threshold analysis of the relevance matrix, find the regions where the relevance scores are lower than the preset threshold, and mark these regions as potential hallucination content.
[0260] Step A4.2: Use the threshold method to mark potential hallucination content; Purpose: To mark the content that may have hallucinations by setting a threshold, which is determined based on previous analysis and experiments and is used to distinguish normal low-relevance regions from possible hallucination regions; Implementation: According to the threshold setting in the relevance matrix, mark the segments with relevance scores lower than the threshold as potential hallucination content.
[0261] Step A4.3: Calculate the hallucination probability of each marked segment; Purpose: To calculate the hallucination probability for the content marked as potential hallucinations to quantify the severity of the authenticity problem. This step usually involves deeper analysis to determine the likelihood of hallucinations; Implementation: Calculate the hallucination probability of the marked segments, which may be based on factors such as relevance scores, context analysis, and the previous performance of the model. Additional models or statistical methods can be used to further analyze these segments.
[0262] Use the BERT model for layer-by-layer relevance propagation and hallucination detection. Use BERT for text tokenization and convert it into the input format of the model; Based on the outputs of each layer of BERT, calculate and propagate relevance information through a custom attention mechanism and gated update unit; Process the relevance information layer by layer starting from the bottom layer of the model, and synthesize the results of each layer to generate the final relevance matrix; Analyze the relevance matrix, mark the low-relevance regions as potential hallucination content, and calculate the hallucination probability of this content to detect and evaluate the authenticity and accuracy of the answers generated by the model.
[0263] By analyzing and updating relevance layer by layer, it is possible to deeply understand the authenticity of the content generated by the model, effectively detect and mark potential hallucination content, which helps to improve the reliability of the generation model and the user experience. By identifying hallucination content, the model can be further optimized to enhance the quality and relevance of the generated content.
[0264] Furthermore, the intelligent learning companion toy interaction method based on a large model disclosed in this application further includes: identifying structural damage of the toy based on a preset identification strategy, specifically including the following steps:
[0265] Collect data of the toy body based on a multi-modal sensor network; for example, vibration sensors, temperature sensors, strain gauges, etc. can be used to provide various data types related to the structural health status.
[0266] Develop or deploy a data acquisition system to transmit sensor data to a storage and processing platform through a network, which specifically involves the configuration of hardware and software to ensure that the data can be transmitted and stored in real time and stably.
[0267] Clean and standardize the collected data to remove noise, handle missing values, unify the data format, etc., and prepare high-quality data for model training. Specifically, data cleaning techniques such as denoising, imputing missing values, and data normalization can be used to ensure the accuracy and consistency of the data. Standardization processing can include converting the data into a unified scale and format for subsequent analysis.
[0268] Construct a deep learning model architecture; for example, design a deep learning model architecture for structural damage identification. Select a suitable network structure, such as a convolutional neural network (CNN), a recurrent neural network (RNN), or other models suitable for processing sensor data; define the hierarchical structure of the model, including the input layer, hidden layer, and output layer, and select activation functions, loss functions, and optimization algorithms, etc.
[0269] Use the preprocessed data for model training; specifically, train the deep learning model with the preprocessed data so that the model can identify structural damage; divide the data into a training set and a validation set, use the training set to train the model, and evaluate the performance of the model through the validation set. Adjust hyperparameters to optimize the performance of the model.
[0270] Model performance evaluation and optimization; specifically include evaluating the performance of the model in the damage identification task to ensure its accuracy and robustness, and performing optimization. Use evaluation metrics (such as accuracy, recall rate, F1 score, etc.) to evaluate the performance of the model. Adjust model parameters, structure, or training strategies according to the evaluation results to improve performance.
[0271] Implementation of a continuous learning mechanism; specifically, it includes designing an incremental learning algorithm that enables the model to be updated when new data arrives without having to retrain the entire model. This helps to gradually improve the model during long-term monitoring; defining the strategy for incremental learning, including how to integrate new data into the existing model, how to update model parameters, and how to handle the problem of forgetting.
[0272] Implement an online learning update mechanism; specifically, it includes establishing an online learning mechanism that enables the model to be dynamically updated when real-time data arrives, thereby continuously improving its recognition ability. Develop an online learning system that can automatically acquire new data, update the model, and perform real-time learning in the data stream. This may involve online training algorithms and system integration, etc.
[0273] Build a knowledge distillation system to retain historical knowledge; specifically, through knowledge distillation technology, transfer historical knowledge from the old model to the new model, retain the previous learning results, and at the same time incorporate new knowledge. Implement the knowledge distillation method by training the new model to inherit the knowledge of the old model, ensuring that the new model not only learns new data but also retains and utilizes the previous knowledge. This helps to prevent the model from forgetting the knowledge it has learned.
[0274] This continuous learning-based structural damage identification framework aims to establish a damage identification system that can monitor in real time and continuously improve, collect and process data from multi-modal sensors to ensure data quality; build and train a deep learning model for structural damage identification; design and implement an incremental learning and online update mechanism so that the model can maintain and enhance its recognition ability while continuously acquiring new data. By combining real-time data and continuous learning, this framework can improve the accuracy and adaptability of structural damage identification and is applicable to long-term monitoring and dynamic structural health management.
[0275] Example: Application of the continuous learning-based structural damage identification framework in an intelligent learning companion toy; in this example, we will elaborate on how to implement the continuous learning-based structural damage identification framework in an intelligent learning companion toy called "SmartBuddy".
[0276] Step A1: Data collection and preprocessing. Step A1.1: Design a multi-modal sensor network, specifically including: Step A1.1.1: Install micro acceleration sensors at key structural parts of SmartBuddy to detect the motion state and impacts of the toy; Step A1.1.2: Install strain sensors at parts prone to compression or deformation (such as joints) to monitor the deformation of the toy structure; Step A1.1.3: Install temperature sensors in areas where electronic components are concentrated to detect abnormal temperature changes; Step A1.1.4: Install humidity sensors inside the shell to monitor the impact of environmental humidity on the toy materials.
[0277] Step A1.2: Implement a real-time data acquisition system. Specifically, it includes: Step A1.2.1: Develop a low-power microcontroller program to regularly collect data from various sensors; Step A1.2.2: Implement a data caching mechanism to prevent data loss; Step A1.2.3: Design a trigger-based data acquisition mode to increase the sampling frequency when an anomaly is detected.
[0278] Step A1.3: Data cleaning and standardization. Specifically, it includes: Step A1.3.1: Implement a noise filtering algorithm to remove high-frequency noise from sensor data; Step A1.3.2: Perform data standardization to unify data from different sensors to the same scale; Step A1.3.3: Implement an outlier detection and handling mechanism to eliminate or correct significantly abnormal data points.
[0279] Step A2: Initial model training.
[0280] Step A2.1: Build a deep learning model architecture; Step A2.1.1: Design a multi-layer perceptron (MLP) network as the basic model. Specifically, it includes Step A2.1.2: Add a long short-term memory (LSTM) layer on top of the MLP to capture temporal features; Step A2.1.3: Design a multi-modal fusion layer to integrate data from different types of sensors.
[0281] Step A2.2: Use the preprocessed data for model training. Specifically, it includes: Step A2.2.1: Collect laboratory simulation data, including normal usage and various damage conditions; Step A2.2.2: Use the cross-validation method to divide the training set and the validation set; Step A2.2.3: Implement the batch gradient descent algorithm to optimize the model parameters.
[0282] Step A2.3: Model performance evaluation and optimization. Specifically, it includes: Step A2.3.1: Use a confusion matrix to evaluate the recognition accuracy of the model for different damage types; Step A2.3.2: Implement the early stopping method to prevent overfitting; Step A2.3.3: Use a learning rate decay strategy to improve the convergence stability of the model.
[0283] Step A3: Implement a continuous learning mechanism. Step A3.1: Design an incremental learning algorithm. Specifically, it includes: Step A3.1.1: Implement a sliding window mechanism to retain the usage data for the most recent N days; Step A3.1.2: Design a sample importance weighting strategy to emphasize newly emerging damage patterns; Step A3.1.3: Implement a dynamic model structure expansion mechanism to adapt to new damage types.
[0284] Step A3.2: Implement an online learning and updating mechanism, specifically including: Step A3.2.1: Design a trigger-based update strategy to initiate an update when enough new samples have been accumulated; Step A3.2.2: Implement a gradient accumulation method to update the model under limited computing resources; Step A3.2.3: Develop a model rollback mechanism to restore the previous version when the update leads to a performance decline.
[0285] Step A3.3: Build a knowledge distillation system to retain historical knowledge, specifically including: Step A3.3.1: Implement a teacher-student network architecture to guide the new model's learning with the old model; Step A3.3.2: Design a soft label mechanism to retain the probability output information of the old model; Step A3.3.3: Implement a joint training strategy to balance new data learning and historical knowledge retention.
[0286] Application example: Suppose that after Smart Buddy has been used for some time, the user reports that the right arm joint of the toy has become loose. Our continuous learning framework will operate as follows: 1) Data collection: The acceleration sensor detects an abnormal vibration pattern in the right arm, and the strain sensor records an abnormal stress distribution at the joint; 2) Preprocessing: The system filters out noise and normalizes the collected data, identifying this as a new potential damage pattern; 3) Model update trigger: The accumulated amount of new data reaches the preset threshold, triggering the model update process; 4) Incremental learning: The system uses a sliding window mechanism to update the model by combining the usage data of the last 30 days and the newly collected abnormal data. A higher sample importance weight is assigned to the newly emerged right arm joint loosening pattern; 5) Online learning: Using edge computing devices, the model parameters are updated during the periods when the toy is not in use at night. The gradient accumulation method is used to process the data in batches to adapt to limited computing resources. 6) Knowledge distillation: During the update process, the original model is used as the teacher network to guide the new model's learning, ensuring that the recognition ability of other damage patterns learned previously (such as neck torsion, leg bending, etc.) is not forgotten; 7) Performance verification: The updated model is tested on the validation set to confirm that it has a high recognition accuracy for the newly emerged right arm joint loosening pattern, while maintaining the recognition ability for other damage types; 8) Deployment and monitoring: The new model is successfully deployed to Smart Buddy. The system continuously monitors the model performance, collects user feedback, and prepares for the next round of updates.
[0287] Through this continuous learning process, Smart Buddy can adapt to newly emerged damage patterns and early warn of possible safety hazards. For example, when early signs of right arm joint loosening are detected, the system will prompt parents to check the toy's condition through voice, or automatically stop the movement function of the right arm in severe cases to prevent further damage and ensure children's safety.
[0288] This continuous learning-based structural damage identification framework enables Smart Buddy to continuously evolve, adapt to various usage environments and user behaviors, greatly enhancing the safety, durability, and user experience of the toy.
[0289] Example 1: Smart Math Tutoring: Suppose a 10-year-old child is using this learning companion toy to learn math.
[0290] 1) Speech Recognition and Natural Language Understanding: The child says, "I don't quite understand how to solve fraction addition problems." The system recognizes the keyword "fraction addition" and understands that the user's intention is to seek help. 2) Personalized Learning Path Generation: The system queries the user model and finds that the child has a good foundation in fraction concepts but is weak in operations. The dynamic curriculum planning algorithm decides to start with simple addition of like fractions and gradually transition to addition of unlike fractions. 3) Multimodal Feedback: The speech synthesis module generates encouraging speech: "It's okay, let's learn fraction addition together! We'll start with the easy ones." The LCD screen displays a friendly smiling face emoji, and at the same time, the vibration motor vibrates slightly to convey encouraging tactile feedback. 4) Interactive Learning Process: The system presents a simple addition problem of like fractions: "1 / 4 + 2 / 4 =?" The child answers, "3 / 4"; the system recognizes the correct answer and gives positive feedback: "Great! You answered correctly." At the same time, a happy expression is shown. The system gradually increases the difficulty, introduces addition of unlike fractions, and uses visualization methods (such as showing fraction graphics on the screen) to assist understanding.
[0291] Example 2: Adaptive Language Learning: Suppose a 12-year-old child is using this learning companion toy to learn English.
[0292] 1) Speech recognition and natural language understanding: The child says, "Let's practice English conversation." The system recognizes that the user wants to practice English conversation. 2) Personalized learning path generation: The system queries the user model and finds that the child has good listening skills but needs to improve oral expression. The dynamic curriculum planning algorithm decides to start from daily conversation scenarios and focus on training oral expression. 3) Multimodal feedback: The speech synthesis module responds with natural English speech: "Great!Let's start with a conversation at a restaurant. I'll be the waiter, and you can be the customer." The LCD screen displays a simple animation of a restaurant scene to enhance the sense of context. 4) Interactive learning process: The system acts as the waiter: "Good evening!Welcome to our restaurant. What would you like to order?" The child answers, "I want...uh...hamburger." The system recognizes a grammar error but does not directly point it out. Instead, it guides by repeating the correct sentence pattern: "Certainly!You would like a hamburger. Would you like any drinks with that?" The system continuously adjusts the conversation difficulty, introduces new vocabulary and more complex sentence patterns, and gives encouragement through facial expressions and tactile feedback.
[0293] These examples demonstrate how an interactive learning system utilizes natural language understanding, personalized learning path generation, and multimodal feedback to create an engaging and efficient learning experience. The system can adapt to the user's needs and performance in real time, providing targeted guidance and encouragement, thus maximizing the learning effect.
[0294] The intelligent learning companion toy interaction method based on a large model disclosed in this application further includes: a security and privacy protection mechanism.
[0295] Specifically, it includes: Step 5.1: Localization processing; Step 5.1.1: Implement local storage of sensitive data. Purpose: Ensure that sensitive data (such as personal identity information, health data, etc.) is stored locally instead of being transmitted to a remote server. This can reduce the risk of data leakage during transmission and enhance data control and protection. Implementation: Store the data on the user device or a local server instead of sending it over the network. A secure storage mechanism needs to be configured to ensure the protection of data during storage.
[0296] Step 5.1.2: Develop an edge computing processing pipeline. Purpose: To process data at the source of data generation (i.e., edge devices), reducing dependence on remote servers. This not only improves processing speed but also enhances data privacy protection as data processing occurs locally. Implementation: Establish an edge computing architecture, deploy computing and data processing tasks to edge devices (such as sensors, smart terminals, etc.), ensure that edge devices can process data efficiently and securely, and reduce the amount of data transmission.
[0297] Step 5.1.3: Design a data minimization transmission strategy. Purpose: To reduce the amount of data transmission to lower the risk of data leakage. Only necessary data will be transmitted to remote servers or other systems. Implementation: Formulate a data transmission strategy, including the types, frequencies, and amounts of data transmitted. Methods such as data aggregation, compression, or screening can be used to transmit only information necessary for analysis or processing.
[0298] Step 5.2: Data encryption system; Step 5.2.1: Implement an end-to-end encryption mechanism. Purpose: To ensure that data remains encrypted throughout the transmission process from the source to the destination, preventing data from being stolen or tampered with during transmission. Implementation: Use encryption algorithms (such as AES, RSA, etc.) to encrypt data, ensure that data is encrypted during both transmission and storage, and implement the processes of encryption and decryption, and ensure the secure management of keys.
[0299] Step 5.2.2: Develop a secure key management system. Purpose: To manage the keys required for encryption, ensure that the processes of key generation, storage, distribution, and update are secure, preventing key leakage or abuse; Implementation: Establish a key management system responsible for key generation, storage, and distribution. The system should have key lifecycle management functions to ensure the security and effectiveness of keys.
[0300] Step 5.2.3: Implement data anonymization processing. Purpose: To anonymize data to prevent personal identities from being identified through data, which helps protect user privacy, especially during data analysis and sharing. Implementation: Apply data anonymization techniques (such as de-identification, data masking, etc.) to ensure that data cannot be directly associated with personal identities during processing and analysis, and consider how to maintain the usefulness of data while protecting privacy.
[0301] Step 5.3: Access Control and Auditing; Step 5.3.1: Design a multi-factor authentication system. Purpose: To enhance the security of system access by verifying the user's identity through multiple authentication factors and prevent unauthorized access. Implementation: Implement multi-factor authentication (such as password + SMS verification code, fingerprint recognition, etc.) to enhance the security of user authentication and ensure that the authentication process is both convenient and secure. Step 5.3.2: Implement fine-grained permission control. Purpose: To define and manage user permissions in the system, ensure that users can only access the data and functions within their authorized scope, and prevent unauthorized operations. Implementation: Design a permission management system to configure access permissions based on user roles and requirements, and models such as Role-Based Access Control (RBAC) or Attribute-Based Access Control (ABAC) can be used.
[0302] Step 5.3.3: Develop a comprehensive log auditing system. Purpose: To record the operations and events of the system, provide a basis for audit tracking and problem troubleshooting. Audit logs help to detect and analyze security incidents and ensure the compliance and security of the system. Implementation: Establish a log recording and management system to record information such as user operations, system events, and data access. It is necessary to ensure the integrity and security of the logs and provide audit analysis tools for event detection and investigation.
[0303] These steps describe how to implement a comprehensive security and privacy protection mechanism, including: Localization processing: Reduce data transmission and enhance the security of local data storage and processing; Data encryption system: Ensure the confidentiality of data during storage and transmission; Access control and auditing: Improve the security and compliance of the system through multi-factor authentication, permission control, and log auditing. These measures work together to protect users' privacy and data security, prevent data leakage and unauthorized access, and ensure the security and reliability of the system.
[0304] The second aspect of this application discloses an intelligent toy, including a central control center and a toy body with voice interaction function. The central control center integrates the intelligent learning companion toy interaction method based on the large model, which is used to control the toy body to perform intelligent interaction with users.
[0305] Specifically, the central control center, as the core control unit of the entire intelligent toy, is responsible for coordinating the operations of the toy body and the interaction with users; the central control center integrates the intelligent learning companion toy interaction method based on the large model, and this large model can be a language model similar to GPT, which can understand and generate natural language to provide an intelligent interaction experience. The central control center is usually equipped with powerful computing and storage capabilities to handle complex computing tasks and store interaction data. The central control center communicates with the toy body through wireless or wired networks to ensure the real-time transmission of data and instructions, and can receive updates through the network to continuously improve and optimize the functions and interaction strategies of the intelligent learning companion.
[0306] The toy body is integrated with voice recognition and voice synthesis functions, enabling it to communicate with users via voice, and execute specific actions or reactions according to the instructions of the central control center, such as making sounds, moving, flashing lights, etc. Specifically, the toy body can be built-in with sensors (such as microphones, cameras, accelerometers, etc.) and actuators (such as speakers, motors, LED lights, etc.), enabling the toy to perceive the environment and respond.
[0307] By integrating the central control center and the toy body with voice interaction functions, this intelligent toy can provide a rich and intelligent user interaction experience. The central control center realizes intelligent dialogue management and learning through an advanced large model, while the toy body interacts with users through voice interaction and various sensors. The design of the entire system aims to provide a highly interactive and intelligent toy experience that can continuously learn and adapt to the needs of users.
[0308] In a third aspect, the present application discloses an intelligent learning companion toy interaction system based on a large model, including: a knowledge network system construction module for constructing a knowledge network system of the target domain; the knowledge network system includes various theoretical knowledge, knowledge entities, association relationships between different knowledge entities, and credibility weights corresponding to the association relationships in the target domain; a target large language model construction module for constructing a target large language model based on the knowledge network system; a capture module for capturing multi-source data of the user in real time; the multi-source data includes one or more of user voice instructions and user interaction actions; an answer text generation module for analyzing the multi-source data based on the target large language model and generating an answer text; an interaction strategy generation module for performing hallucination detection on the answer text based on a preset detection model and generating an interaction strategy based on the answer text that meets the conditions; a feedback module for performing multi-modal feedback on the intelligent learning companion toy based on the interaction strategy.
[0309] It should be noted that the solutions in the intelligent learning companion toy interaction method based on a large model disclosed in the first aspect of the present application are all applicable to the intelligent learning companion toy interaction system based on a large model disclosed in the third aspect of the present application.
[0310] The computer device according to an embodiment of the present disclosure includes a memory and a processor. The memory is used to store non-temporary computer-readable instructions. Specifically, the memory may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0311] The processor can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the computer device to perform desired functions. In one embodiment of the present disclosure, the processor is used to run the computer-readable instructions stored in the memory, so that the computer device executes all or part of the steps of the large model-based intelligent learning companion toy interaction method according to the embodiments of the present disclosure described above.
[0312] Those skilled in the art should understand that, in order to solve the technical problem of how to obtain good user experience effects, well-known structures such as communication buses and interfaces may also be included in this embodiment, and these well-known structures should also be included in the protection scope of the present disclosure.
[0313] As Figure 11 FIG. [ID] is a schematic structural diagram of a computer device provided by an embodiment of the present disclosure. It shows a schematic structural diagram of a computer device suitable for implementing the computer device in the embodiments of the present disclosure. Figure 11 The shown computer device is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0314] As Figure 11 As shown, the computer device may include a processor (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) or the program loaded from the storage device into the random access memory (RAM). In the RAM, various programs and data required for the operation of the computer device are also stored. The processor, ROM, and RAM are connected to each other through a bus. The input / output (I / O) interface is also connected to the bus.
[0315] Generally, the following devices may be connected to the I / O interface: an input device including, for example, a sensor or a visual information acquisition device; an output device including, for example, a display screen; a storage device including, for example, a magnetic tape, a hard disk, etc.; and a communication device. The communication device can allow the computer device to communicate wirelessly or wiredly with other devices (such as edge computing devices) to exchange data. Although Figure 11 FIG. [ID] shows a computer device having various devices, but it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0316] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processor, all or part of the steps of the intelligent learning companion toy interaction method based on a large model according to the embodiments of the present disclosure are performed.
[0317] For a detailed description of this embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, and details are not repeated herein.
[0318] A computer-readable storage medium according to an embodiment of the present disclosure stores non-temporary computer-readable instructions. When the non-temporary computer-readable instructions are run by a processor, all or part of the steps of the intelligent learning companion toy interaction method based on a large model according to the foregoing embodiments of the present disclosure are performed.
[0319] The above computer-readable storage media include, but are not limited to: optical storage media (e.g., CD-ROMs and DVDs), magneto-optical storage media (e.g., MOs), magnetic storage media (e.g., magnetic tapes or external hard drives), media with built-in rewritable non-volatile memories (e.g., memory cards), and media with built-in ROMs (e.g., ROM cartridges).
[0320] For a detailed description of this embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, and details are not repeated herein.
[0321] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above specific details are only for illustrative and facilitating understanding purposes, rather than limitations, and the above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.
[0322] In this disclosure, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. The block diagrams of the devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used herein refer to the phrase "and / or", and can be used interchangeably with it, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with it.
[0323] In addition, as used herein, the "or" used in the listing of items starting with "at least one" indicates a disjunctive listing, so that for example, the listing of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the term "exemplary" does not mean that the described examples are preferred or better than other examples.
[0324] It should also be noted that in the systems and methods of this disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this disclosure.
[0325] Various changes, substitutions, and alterations to the technologies described herein can be made without departing from the teachings defined by the appended claims. In addition, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and acts described above. Current or later-developed processes, machines, manufactures, compositions of events, means, methods, or acts that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Thus, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or acts within their scope.
[0326] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
[0327] The foregoing description has been presented for purposes of illustration and description. In addition, this description is not intended to limit embodiments of the present disclosure to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.
Claims
1. A method for interacting with a smart learning companion toy based on a large model, characterized in that: include: Build a knowledge network system in the target field; The knowledge network system includes various theoretical knowledge in the target field, knowledge entities, associations between different knowledge entities, and credibility weights corresponding to the associations; Building a target large language model based on the knowledge network system; Capturing multi-source data of the user in real time; the multi-source data includes one or more of user voice commands and user interactive actions; Analyze the multi-source data based on the target large language model to generate a response text; Performing hallucination detection on the answer text based on a preset detection model, and generating an interactive strategy based on the answer text that meets the conditions; Multimodal feedback of the intelligent learning companion toy is performed based on the interactive strategy.
2. The method for interacting with a large model-based intelligent learning companion toy according to claim 1, characterized in that: The step of analyzing the multi-source data based on the target large language model to generate a response text includes: Preprocessing the multi-source data, and converting the processed data into first text information based on a deep learning model; Preprocessing the first text information using natural language processing technology to obtain second text information; Extracting key features from the second text information; Selecting a target large language model from a target database based on the key features; The key features are analyzed based on the target large language model to generate a response text.
3. The method for interacting with a large model-based intelligent learning companion toy according to claim 2, characterized in that: The extracting key features from the second text information includes: Performing word segmentation processing on the second text information to obtain a plurality of sub-words; Assigning a vocabulary label to each of the sub-vocabulary to obtain a plurality of sub-labels; Based on the plurality of sub-tags, identifying entity information; the entity information includes one or more of a person's name, a place name, an organization, and a professional term; Based on the plurality of sub-vocabularies and the plurality of sub-tags, the user intention is obtained; the user intention and the entity information constitute the key feature.
4. The method for interacting with a large model-based intelligent learning companion toy according to claim 3, characterized in that: The step of analyzing the key features based on the target large language model to generate a response text includes: Based on the key features, create a user profile; Determine whether the key feature belongs to a learning feature. If so, determine whether the corresponding learning is the first time based on historical interaction records. If so, analyze the user portrait and the key feature based on the target large language model to generate a response text. If the corresponding learning is not the first time, obtaining the associated learning information in the historical interaction record; Generate a personalized learning strategy based on the associated learning information and the user portrait, and generate an answer text of the current progress based on the personalized learning strategy; If the key feature does not belong to a learning feature, the user portrait and the key feature are analyzed based on the target large language model to generate a response text.
5. The method for interacting with a large model-based intelligent learning companion toy according to claim 2, characterized in that: The method of performing hallucination detection on the answer text based on a preset detection model and generating an interactive strategy based on the answer text that meets the conditions includes: Based on the multi-source data, obtain several query word vectors; Based on the answer text, obtain a number of answer word vectors; Analyze the query word vectors and the answer word vectors based on a preset strategy to obtain correlation information; Analyze the correlation information based on a preset analysis strategy to obtain a final score; If the final score is not less than a preset score threshold, then there is no hallucination in the answer text, and an interaction strategy is generated based on the answer text; If the final score is less than a preset score threshold, a correction strategy is triggered to be executed, and an answer text is regenerated based on the correction strategy. Hallucination detection is performed on the regenerated answer text, and an interactive strategy is generated based on the answer text that meets the conditions.
6. The method for interacting with a large model-based intelligent learning companion toy according to claim 5, characterized in that: Analyzing the correlation information based on a preset analysis strategy to obtain a final score includes: Obtaining an initial correlation between each of the query word vectors and the answer word vectors associated therewith; constructing a correlation matrix based on the initial correlations; Based on a preset correlation upward propagation rule, updating the correlation matrix layer by layer; When the preset convergence condition is met, the correlation propagation process is stopped, and the final correlation matrix is output to obtain the final score.
7. The method for interacting with a large model-based intelligent learning companion toy according to claim 6, characterized in that: The constructing a correlation matrix based on the initial correlations comprises: Determine target keywords based on the correlation matrix and the answer word vector; Constructing a target graph with each of the target keywords as a node; Establishing a connection relationship between different nodes according to the word meaning similarity and the position relationship, wherein the connection relationship includes connection edges between different nodes and corresponding weights; Assign a word embedding vector to each node and a feature to each connecting edge.
8. The method for interacting with a large model-based intelligent learning companion toy according to claim 6, characterized in that: If the final score is less than the preset score threshold, triggering execution of a correction strategy includes: Based on the final correlation matrix, the final nodes are obtained; Obtaining a connection path in the correlation matrix; Determine a propagation path from the corresponding query node to the answer node based on the final node and the connection path; Obtaining node connectivity and relevance scores based on the propagation path, identifying answer word vectors with low relevance or unexplainable, and recording the corresponding answer word vectors as risk answer word vectors; Acquire context information of the risk answer word vector based on the answer text; Acquire the type of the risk answer word vector based on the context information, where the type includes one or more of false information and irrelevant content; Acquire a corresponding training database and a usage database in the target large language model based on the risk answer word vector, and correct the training database and the usage database based on the type; Retrain the target large language model based on the corrected training database, and regenerate the answer text based on the trained target large language model; Perform hallucination detection on the regenerated answer text, and generate an interactive strategy based on the answer text that meets the conditions.
9. The method for interacting with a large model-based intelligent learning companion toy according to claim 5, characterized in that: The analyzing the correlation information based on the preset analysis strategy to obtain the final score includes: based on the hierarchical structure of the preset neural network, transferring the correlation information layer by layer from the bottom layer to the top layer, obtaining the correlation score of each layer, and controlling the transferred content based on the preset gating mechanism; obtaining the final score based on the obtained correlation score of each layer; If the final score is less than the preset score threshold, triggering execution of a correction strategy includes: If the final score is less than the preset score threshold, the answer word vector corresponding to the minimum relevance score is obtained and recorded as the risk answer word vector; Acquire context information of the risk answer word vector based on the answer text; Acquire the type of the risk answer word vector based on the context information, where the type includes one or more of false information and irrelevant content; Acquire a corresponding training database and a usage database in the target large language model based on the risk answer word vector, and correct the training database and the usage database based on the type; Retrain the target large language model based on the corrected training database, and regenerate the answer text based on the trained target large language model; Perform hallucination detection on the regenerated answer text, and generate an interactive strategy based on the answer text that meets the conditions.
10. A smart toy, characterized in that: Including a master control center and a toy body with voice interaction function; The master control center is integrated with the large-model-based intelligent learning companion toy interaction method according to any one of claims 1 to 9, and is used to control the toy body to perform intelligent interaction with the user.
Citation Information
Patent Citations
Method and system for solving illusion problem of large legal language model
CN117744802A
Large model illusion relieving method and device, equipment and storage medium
CN118964583A
Child education deep interaction method and device based on multi-modal large model
CN119179386A
Answer generation method and device for intelligent customer service question-answering system and storage medium
CN119248918A
Multi-source hybrid question answering method and system thereof
KR101662450B1
Cited By
Open source large model-based poem beginner multi-task cooperative training method
CN121301534A
A poem beginner multi-task cooperative training method based on an open source large model
CN121301534B