Robot interaction method and interaction system, robot and computer program product

By extracting long-term memory and combining short-term memory, the robot can remember historical interaction information and user preferences in multiple rounds of conversations, solving the problem of repeatedly elaborating background information in traditional systems and improving communication efficiency and coherence.

CN120045751APending Publication Date: 2025-05-27UBTECH ROBOTICS CORP LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411994402.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Traditional robot chat systems need to repeatedly elaborate on the same background information in multiple rounds of conversations, resulting in the impact of communication coherence and response efficiency.

Method used

Long-term memory is extracted through the agent, remember historical interaction information, user preferences and detailed content in specific scenarios, and extract dialogue information in historical dialogues in combination with short-term memory to generate personalized and accurate responses.

Benefits of technology

It improves the communication efficiency and coherence of the robot interaction system, allowing the robot to quickly recall interactive topics and related points in multiple rounds of dialogue, and provides more personalized and accurate services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045751A_ABST
    Figure CN120045751A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of robots, and provides a robot interaction method and system, a robot and a computer program product, and the method comprises the steps: obtaining interaction content inputted by a user; preprocessing the interaction content, and outputting long-term memory related information corresponding to the interaction content; under the condition that the interaction type is determined to be a session type, extracting related information of a preset session round number according to the interaction content; and generating prompt information according to the related information of the long-term memory and the related information of the preset dialogue round number, and inputting the prompt information into the large language model to obtain reply content corresponding to the prompt information. According to the method, the robot can effectively remember historical interaction information, user preferences and detail contents in a specific scene, dialogue information in historical dialogues is extracted in combination with short-term memory, the robot can quickly recall interaction topics and key points in multiple rounds of dialogues, and the communication efficiency and continuity are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of robots, and particularly relates to an interaction method and system for robots, a robot, and a computer program product. Background Art

[0002] With the continuous development of robot technology, intelligent robots are widely used in various fields, such as hotel services, mall guidance, hospital guidance, and so on. In the actual application process, robots can interact with users through an interaction system. Traditional chat systems often only reply based on the current input. When users have multi-round conversations, they need to repeatedly elaborate on the same background information, which will undoubtedly affect the coherence and response efficiency of the communication. Therefore, how to improve the response ability of the robot interaction system is an urgent problem to be solved. Summary of the Invention

[0003] The embodiments of this application provide an interaction method and system for robots, a robot, and a computer program product, which can effectively improve the communication efficiency and coherence.

[0004] In a first aspect, the embodiments of this application provide an interaction method for a robot, including:

[0005] Obtain the interaction content input by the user;

[0006] Preprocess the interaction content and output the relevant information of the long-term memory corresponding to the interaction content;

[0007] When it is determined that the interaction type is a conversation type, extract the relevant information of a preset number of dialogue rounds according to the interaction content;

[0008] Generate a prompt message according to the relevant information of the long-term memory and the relevant information of the preset number of dialogue rounds, and input the prompt message into a large language model to obtain a reply content corresponding to the prompt message;

[0009] Feed back the reply content to the user.

[0010] In this application, by extracting the long-term memory through an agent, the robot can effectively remember historical interaction information, user preferences, and details in specific scenarios. When interacting with users, it can provide more personalized and accurate services and responses, and combines short-term memory to extract dialogue information from historical conversations, enabling the robot to quickly recall the interaction topics and relevant key points in multi-round conversations, thereby effectively improving the communication efficiency and coherence.

[0011] In a possible implementation manner of the first aspect, the relevant information of the long-term memory is extracted by a long-term memory retrieval agent.

[0012] In a possible implementation manner of the first aspect, after obtaining the interaction content input by the user, the method further includes:

[0013] Determine the preprocessing type according to the configuration information.

[0014] In the embodiments of the present application, the content of the preprocessing that needs to be performed can be determined through the configuration information, and the preprocessing type can be better selected according to the actual application situation of the robot, so that the robot has higher adaptability.

[0015] In a possible implementation manner of the first aspect, after obtaining the interaction content input by the user, the method further includes:

[0016] Determine the interaction type according to the interaction content.

[0017] In a possible implementation manner of the first aspect, after determining the interaction type according to the interaction content, the method further includes:

[0018] In the case where the determined interaction type is a task type, perform task decomposition and execute the decomposed subtasks.

[0019] The interaction method provided by the embodiments of the present application can also perform task decomposition and execution based on a large model when responding to tasks, which can effectively improve the response efficiency and accuracy.

[0020] In a possible implementation manner of the first aspect, the generating the prompt information according to the relevant information of the long-term memory and the relevant information of the preset number of dialogue turns includes:

[0021] Determine the input template according to the configuration information;

[0022] According to the format of the input template, splice the relevant information of the long-term memory and the relevant information of the preset number of dialogue turns into the input template to obtain the prompt information.

[0023] In a possible implementation manner of the first aspect, the long-term memory retrieval agent includes multiple simple agents.

[0024] In the present application, long-term memory retrieval is set throughout the communication process to assist in the generation of conversations. By setting rules for simple agents and replacing the task decomposition part with the rules in the profile, the invocation of the large language model is reduced, and the time consumption and resources are reduced.

[0025] In a second aspect, an interaction system for a robot provided by the embodiments of the present application includes:

[0026] A configuration module, configured to configure the robot interaction system according to configuration information;

[0027] A preprocessing module, configured to preprocess the interaction content input by the user according to the preprocessing type determined by the configuration information, so as to obtain respective preprocessing results;

[0028] A session processing module, configured to extract relevant information of a preset number of dialogue turns according to the interaction content when it is determined that the interaction type is a session type;

[0029] A template splicing module, configured to generate a prompt message according to the relevant information in long-term memory and the relevant information of the preset number of dialogue turns, and input the prompt message into a large language model to obtain a reply content corresponding to the prompt message;

[0030] A reply module, configured to feedback the reply content to the user.

[0031] Thirdly, an embodiment of the present application provides a robot, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, where when the processor executes the computer program, the method described in the first aspect is implemented.

[0032] Fourthly, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect is implemented.

[0033] Fifthly, an embodiment of the present application provides a computer program product, when the computer program product runs on a robot, the robot is enabled to execute the method described in the first aspect above.

[0034] It can be understood that the beneficial effects of the above second aspect to fifth aspect can refer to the relevant descriptions in the first aspect above, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0036] Figure 1 It is a schematic flowchart of the implementation of an interaction method of a robot provided by an embodiment of the present application;

[0037] Figure 2 It is a schematic flowchart of the implementation of another interaction method of a robot provided by an embodiment of the present application;

[0038] Figure 3It is a schematic diagram of the application of the robot interaction method provided by the embodiments of the present application;

[0039] Figure 4 It is a schematic diagram of the implementation process of another robot interaction method provided by the embodiments of the present application;

[0040] Figure 5 It is a schematic diagram of the structure of a robot interaction system provided by the embodiments of the present application;

[0041] Figure 6 It is a schematic diagram of the structure of another robot interaction system provided by the embodiments of the present application

[0042] Figure 7 It is a schematic diagram of the structure of a robot provided by the embodiments of the present application;

[0043] Figure 8 It is a schematic diagram of the structure of another robot provided by the embodiments of the present application. Detailed implementation manners

[0044] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0045] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0046] It should also be understood that the term "and / or" as used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0047] As used in the specification of the present application and the appended claims, the term "if" can be interpreted as "when...", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined", "in response to determining", "once detecting [the described condition or event]", or "in response to detecting [the described condition or event]" according to the context.

[0048] In addition, in the description of the specification and the appended claims of this application, the terms "first", "second", "third", etc. are only used for differential description and should not be construed as indicating or implying relative importance.

[0049] The reference to "one embodiment" or "some embodiments" etc. described in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprise", "include", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0050] First, the relevant concepts involved in the embodiments of this application are described as follows:

[0051] 1. Agent:

[0052] It refers to an entity that can perceive the environment, make decisions, and take actions. It can be a software program, a robot, or other autonomous computing units. The behavior of an Agent can be based on simple rules or have a complex internal structure. For example, a simple automatic temperature regulation Agent that senses the ambient temperature through a temperature sensor and turns on the refrigeration equipment when the temperature is higher than the set value is an Agent based on simple rules. Its core lies in autonomy and the ability to interact with the environment.

[0053] 2. AI Agent:

[0054] It is a subset of Agent that emphasizes the use of artificial intelligence technologies to achieve perception, decision-making, and action. AI Agents usually use artificial intelligence methods such as machine learning, deep learning, and natural language processing to handle complex tasks and environments. For example, an intelligent customer service AI Agent that uses natural language processing technology to understand customers' questions, classifies the questions through a machine learning model, generates answers based on the learned knowledge, and then takes actions (such as sending the answer to the customer, disassembling the task and completing it).

[0055] AI Agents need to use complex artificial intelligence technologies. In the perception stage, convolutional neural networks (CNNs) in deep learning may be used for image perception, or language models based on the Transformer architecture may be used for natural language perception. In the decision-making stage, reinforcement learning algorithms or classification and generation models based on deep learning are used to make decisions.

[0056] 3. The Large Language Model (LLM) is an artificial intelligence technology based on deep learning.

[0057] The large language model uses a large-scale text dataset to train the model. Through a stacked neural network structure, it learns and simulates the complex rules of human language, enabling it to generate natural language text or understand the meaning of language text. The Transformer architecture is an important foundation for large language models. It adopts a self-attention mechanism, which can better capture long-range dependencies in language, solves the limitations of recurrent neural networks in parallel processing, and significantly improves the model's ability to process large-scale datasets, greatly enhancing the performance of large language models.

[0058] With the continuous development of robot technology, intelligent robots are widely used in various fields, such as hotel services, mall guidance, hospital guidance, etc. In the actual application process, robots can interact with users through an interaction system. Traditional chat systems often only reply based on the current input. When users have multiple rounds of conversations, they need to repeatedly elaborate on the same background information, which will undoubtedly affect the coherence and response efficiency of the communication. Therefore, how to improve the response ability of the robot interaction system is an urgent problem to be solved.

[0059] To improve the response ability of a robot interaction system using a large model, the embodiments of the present application provide an interaction method and interaction system for a robot, a robot, and a computer program product. It can extract long-term memory through an agent, enabling the robot to effectively remember historical interaction information, user preferences, and details in specific scenarios. When interacting with users, it can provide more personalized and accurate services and responses, and combines short-term memory to extract conversation information from historical conversations, enabling the robot to quickly recall interaction topics and relevant key points in multiple rounds of conversations, thereby effectively improving communication efficiency and coherence.

[0060] Figure 1 The figure shows a schematic implementation flowchart of an interaction method for a robot provided by the embodiments of the present application. Among them, the execution entity is a robot, such as Figure 1 As shown, the above-mentioned interaction method for a robot may include S101 to S105, which are described in detail as follows:

[0061] S101: Obtain the interaction content input by the user.

[0062] In specific applications, users can interact with the robot through the interaction entry provided by the robot. The interaction content input by the user may include, but is not limited to: pictures, texts, audios, videos, and other contents.

[0063] S102: Preprocess the interaction content and output information related to the long-term memory corresponding to the interaction content.

[0064] In the embodiments of the present application, the long-term memory can be user portraits, key object information, image information, etc. extracted from a large amount of historical data accumulated by the interaction system in the past.

[0065] In specific applications, an intelligent agent for long-term memory retrieval can be used to retrieve information related to the long-term memory corresponding to the interaction content.

[0066] In some embodiments, the above intelligent agent for long-term memory retrieval may include multiple simple intelligent agents. Among them, a simple intelligent agent refers to an intelligent agent that only includes a role definition (profile) and an action.

[0067] Generally, an intelligent agent can include four parts: role definition (profile), memory extraction and storage (memory), task decomposition (planning), and action. In the embodiments of the present application, the long-term memory retrieval is set throughout the communication process to assist in the generation of conversations. By setting rules for simple intelligent agents and replacing the task decomposition part with the rules in the profile, the invocation of large language models is reduced, and the time consumption and resources are reduced.

[0068] In specific applications, the above multiple simple intelligent agents may specifically include: portrait intelligent agent, key object intelligent agent, image intelligent agent, evaluation intelligent agent, and forgetting summary intelligent agent.

[0069] Among them, the portrait intelligent agent can use a large language model to extract portrait information and user tags of the user and user-related people from the user's multi-round conversations. User tags include but are not limited to name, gender, age, birthday, location, occupation, personality, interpersonal relationships, skills, pets, hobbies, behaviors, plans, needs, etc.

[0070] The key object intelligent agent is used to extract information about non-human related objects, such as information about animals, plants, furniture, exhibits, etc. from the user's multi-round conversations, and generate corresponding object tags. Object tags include but are not limited to weight, age, location, activities, important events, health status, hobbies, habits, etc.

[0071] The image intelligent agent can use a vision-language model to extract visual information of people and objects that appear in the images captured by the robot, such as information about clothing, relative positions, activities, states, etc.

[0072] The evaluation agent has additional content related to rule processing compared to other agents. For example, it can include removing the generated "none" information in statements, etc. It can cooperate with the large language model to align the user statements and the objects in the picture, and evaluate the information extracted by the above portrait agent, key object agent, and image agent to confirm whether the information is credible, whether there are contradictions between the information, or whether there is unreasonable information. Summarize according to the evaluation. At the same time, it is necessary to determine the specific storage location and corresponding execution actions of the information extracted by the above multiple agents. For example, determine which table in the database the information needs to be stored in and execute the action.

[0073] The forgetting summary agent can be started regularly. It judges whether the relevant information in the database needs to be forgotten through the large language model. For example, it can forget the plan a month ago and the relevant information of the plan, and summarize some information, such as summarizing the user's personality, and then update it to the database.

[0074] It should be noted that the above database can be the database in the large language model used to store various types of information, such as the database for storing the above long-term memory information.

[0075] Among these agents, the profile stores the input prompts (prompts) of the agent and the attributes and limitations of the agent. These attributes define how the above large language model understands and responds to tasks.

[0076] Exemplarily, the prompt of the above agent can be described as: You are an extremely accurate portrait information extraction expert, and your task is to record the portrait information of the user and the user's relatives and friends from the historical conversation records between the user and the robot.

[0077] Format requirements: List the clear information in the format of "[serial number].[The user and the relevant humans of the user, not pets. If it is the user himself, fill in 'user', if it is a user relative, fill in the relevant address such as father, mother, and if it is a friend, fill in the friend's name]-[tag]-[specific information]", with a maximum of 2 "-" signs. If there is no clear information, write "Clear information:\nnone". The tag range in the clear information: 'name', 'nickname','sex', 'age', 'birthday', 'geographical location', 'occupation', 'educational background', 'personality characteristics', 'emotional status', 'family members', 'interpersonal relationships', 'work and rest patterns', 'dietary preferences', 'pet situation', 'travel mode', 'hobbies', 'needs', 'plans', 'behavior events','mood', 'emergency events', 'tasks for the robot', etc.

[0078] Rule description:

[0079] 1. The conversation history is a chat session between the user and the robot. The "you" in the user's words refers to the robot.

[0080] 2. Only extract the portrait information of people related to the user (the user, the user's relatives, friends) from the user's conversation records. Do not extract the portrait information of non - humans (such as animals, plants) and robots.

[0081] 3. Do not extract information about people who are not relatives or friends (such as historical figures, political figures, etc.). If the relevant person is not specified (such as the user mentions him, but does not say who he is), then there is no need to extract.

[0082] Exemplarily, the attributes and limitations stored in the profile for the agent can include:

[0083] The user portrait agent is only used when there is a result in face recognition.

[0084] The image summary can be initiated not only when the user is having a conversation, but also actively when there is no conversation, to extract information about changes in the surrounding environment.

[0085] The execution module mainly operates on the database, storing or extracting.

[0086] In specific applications, the above - mentioned portrait agent and key object agent can adopt text large models, and the image agent and evaluation agent can use multi - modal models. The number of parameters used can be within 32b. Of course, each of the above agents can also use multi - modal models. Of course, other agents such as video extraction agents and audio extraction agents can also be added to achieve information extraction, and the functions of multiple agents can also be integrated into one agent for processing. This application does not make a unique limitation on this.

[0087] In specific applications, after obtaining the interaction content, it is possible to determine whether to obtain relevant information from the long - term memory in the database through the input prompt information or interaction similarity of the large - language model, and specifically extract which tags of which objects as the long - term memory corresponding to the interaction content. For example, if the user asks who likes to eat fish, the interaction system can extract from the database the similar history that the user or pet's dietary preference is to like eating fish. These historical information of the user or pet can be output as the relevant content of the long - term memory related to the interaction content.

[0088] It can be understood that the extraction and storage of the above - mentioned long - term memory can be performed regularly in the background. Specifically, it can be to convert the relevant information extracted from the short - term history into long - term memory and store it in the database.

[0089] The forgetting of memory information can specifically be the summary and induction of key information and the forgetting of unimportant information.

[0090] S103. When it is determined that the interaction type is the conversation type, relevant information for a preset number of dialogue turns is extracted according to the interaction content.

[0091] In specific applications, the preset number of dialogue turns can be set according to the actual application situation. For example, it can be 20 dialogue turns, that is, the information corresponding to the interaction content in 20 dialogue turns can be extracted to achieve the extraction of short-term memory. The extracted information is the key details in the 20-round interaction process, such as the needs, preferences mentioned by the user in the historical conversation, or the intermediate results of a certain task just completed, etc. For example, in a shopping mall robot providing the function of locating stores in the shopping mall, the short-term memory can be the store names involved in the user's current consultation question, etc.

[0092] S104. Generate a prompt message based on the relevant information of the long-term memory and the relevant information of the preset number of dialogue turns, and input the prompt message into the large language model to obtain the reply content corresponding to the prompt message.

[0093] In specific applications, after extracting the short-term memory, that is, the relevant information of the above-mentioned preset number of dialogue turns, data preprocessing and integration can also be performed on the relevant information of the preset number of dialogue turns. For example, incomplete or incorrect data formats are cleaned to enable effective fusion with other data. At the same time, it is merged with the relevant information of the long-term memory, and input templates are spliced to obtain a prompt message.

[0094] In specific applications, the input template of the large language model can be pre-set with different template structures according to the task and scenario. Based on the template structure, the integrated data is arranged and combined according to a specific format and logic to generate a prompt message prompt that meets the input requirements of the large language model.

[0095] Exemplarily, in a text generation task, the input template can stipulate that the beginning part is the background introduction of the task. The relevant content of the background introduction can be filled with the relevant information retrieved from the long-term memory and the relevant information extracted from the short-term memory. The middle part is the specific requirements and constraints, which can be determined according to knowledge retrieval and current needs. The ending part is some prompts for formats or styles, such as language style, word count limit, etc. In this way, the chaotic data is transformed into a large language model input with clear structure and explicit semantics, enabling the large language model to better understand the background and objectives of the task, and thus generating more accurate and user-expected reply content.

[0096] S105. Feed back the reply content to the user.

[0097] In specific applications, the robot can feed back the reply content to the user through interaction devices such as a screen and a speaker.

[0098] As can be seen from the above, the robot interaction method provided by the embodiments of the present application can extract long-term memory through an agent, enabling the robot to effectively remember historical interaction information, user preferences, and details in specific scenarios. When interacting with users, it can provide more personalized and accurate services and responses. Moreover, by combining short-term memory to extract dialogue information from historical conversations, the robot can quickly recall the topics and relevant key points of the interaction in multi-round conversations, thereby effectively improving the communication efficiency and coherence.

[0099] Please refer to Figure 2 , in another embodiment of the present application, different from the previous embodiment, as Figure 2 shown, the interaction method of the robot provided by the embodiments of the present application further includes S201, which is described in detail as follows:

[0100] S201: Determine the preprocessing type according to the configuration information.

[0101] In the embodiments of the present application, after obtaining the interaction content input by the user, the configuration information of the robot can be determined first. Here, the configuration information refers to the pre-configured information related to the interaction process of the robot, and the configuration information of the robot can be set based on the configuration xml file. The configuration information can specifically include information such as which content preprocessing the robot is configured to perform and which input template of the large language model to select.

[0102] In specific applications, the preprocessing of the above interaction content may also include other contents. For example, the interaction type can also be determined based on the interaction content, the robot's emotion can be recognized based on the interaction content, the robot's persona can be recognized based on the interaction content, the robot's expression can be determined based on the interaction content, the robot's action can be determined based on the interaction content, and the knowledge base can be retrieved based on the interaction content, etc.

[0103] In specific applications, the above determination of the interaction type based on the interaction content can use a model or rules to determine whether the interaction content input by the user belongs to the task type or the conversation type.

[0104] In specific applications, the above recognition of the robot's emotion based on the interaction content can specifically be to judge and generate the robot's emotion through a large model, that is, what kind of emotion the robot needs to use to reply to the user. This emotion recognition can assist in expression output, audio generation, and text generation, etc.

[0105] In specific applications, the above recognition of the robot's persona based on the interaction content can specifically be to determine the current persona of the robot based on the user's settings.

[0106] In specific applications, for a robot that can reply with expressions, such as a robot that can display a smiling face, a crying face, a shy expression, etc. on the display screen, the expression that the robot needs to reply can be determined based on the interaction content.

[0107] In a specific application, for a robot that can perform action responses, such as a robot that can perform operations like hugging and clapping, the actions that the robot needs to adopt can be determined based on the interaction content.

[0108] In a specific application, for content that requires professional knowledge for a response, the knowledge base can be called to perform a knowledge base retrieval to find relevant information in the knowledge base.

[0109] Through the configuration information, it can be determined what preprocessing the user can perform, so as to call the relevant modules to implement the preprocessing function.

[0110] Figure 3 The application schematic diagram of the interaction method of the robot provided by the embodiment of the present application is shown, as Figure 3 shown, the preprocessing of the interaction content may include: determining the interaction type based on the interaction content, identifying the robot's emotion based on the interaction content, identifying the robot's persona based on the interaction content, determining the robot's expression based on the interaction content, determining the robot's actions based on the interaction content, performing a knowledge base retrieval based on the interaction content, performing a long-term memory retrieval based on the interaction content, and so on.

[0111] Correspondingly, the above S104 may include: generating a prompt message according to the relevant information of the long-term memory, the relevant information of the knowledge base, and the relevant information of the preset number of dialogue turns, and inputting the prompt message into the large language model to obtain a response content corresponding to the prompt message.

[0112] Merge the relevant information of the short-term memory with the relevant information of the long-term memory and the relevant information retrieved from the knowledge base, and perform input template splicing to obtain the prompt message.

[0113] In a specific application, the input template of the large language model can be set with different template structures in advance according to the task and scenario. Based on the template structure, the integrated data is arranged and combined according to a specific format and logic to generate a prompt message prompt that meets the input requirements of the large language model. Which input template to select to generate the prompt message can be determined according to the above configuration information.

[0114] Based on this, the above S104 may specifically include:

[0115] Determine the input template according to the configuration information;

[0116] According to the format of the input template, splice the relevant information of the above long-term memory, the relevant information of the knowledge base, and the relevant information of the preset number of dialogue turns into the input template to obtain the prompt message.

[0117] Exemplarily, in a text generation task, the input template can specify that the beginning part is the background introduction of the task. The relevant content of the background introduction can be filled with the relevant information retrieved from long-term memory and the relevant information extracted from short-term memory. The middle part is the specific requirements and constraints, which can be determined according to the information retrieved from knowledge (i.e., the relevant information of the above knowledge base) and the current requirements. The ending part is some hints on format or style, such as language style, word count limit, etc. In this way, the chaotic data is transformed into the input of the large language model with clear structure and definite semantics, enabling the large language model to better understand the background and objectives of the task, and thus generating more accurate and more in line with the user's expectations of the response content.

[0118] As can be seen from the above, the interaction method of the robot provided by the embodiments of the present application can determine the content of the preprocessing to be performed through configuration information, and can better select the preprocessing type according to the actual application situation of the robot, making the robot have higher adaptability.

[0119] Please refer to Figure 4 , in another embodiment of the present application, different from the previous embodiment, as Figure 4 shown, the interaction method of the robot provided by the embodiments of the present application further includes S401 to S402, which are described in detail as follows:

[0120] In S401, determine the interaction type according to the interaction content.

[0121] For the relevant content of determining the interaction type according to the interaction content, reference can be made to the relevant description of S201, which will not be elaborated here.

[0122] In S402, when it is determined that the interaction type is a task type, perform task decomposition and execute the decomposed subtasks.

[0123] In specific applications, the above S402 can be executed by a task agent. The task agent can use a large model to split the user's requirements into multiple subtasks and execute them in sequence, and give feedback according to the results of each execution and perform further task decomposition until the task is completed. Exemplarily, the task can be a simple task, such as navigation, charging, or a complex task, such as executing a meal ordering task. Then, the meal ordering task can be split into multiple subtasks such as opening the meal ordering page, automatically / manually selecting the meal, confirming the meal, and placing an order, and the above multiple subtasks are executed in sequence until the meal ordering task is completed.

[0124] As can be seen from the above, the interaction method provided by the embodiments of the present application can also implement task decomposition and execution based on a large model when responding to tasks, which can effectively improve the response efficiency and accuracy.

[0125] The above describes the interaction method of the robot provided by the embodiments of the present application. Next, the execution subject applicable to the interaction method of the robot provided by the embodiments of the present application, that is, the interaction system of the robot, will be described:

[0126] As Figure 5 shown, the interaction system 50 provided by the embodiments of the present application may include a configuration module 51, a preprocessing module 52, a session processing module 53, a template splicing module 54, a reply module 55, a post-processing module 56, and a database 57.

[0127] In the embodiments of the present application, the above configuration module 51 may configure the robot interaction system according to the configuration information.

[0128] For the content of configuring the robot interaction system according to the configuration information, reference may be made to the relevant description of S201, which will not be repeated here.

[0129] The above preprocessing module 52 may preprocess the interaction content input by the user according to the preprocessing type determined by the configuration information to obtain various preprocessing results.

[0130] In specific applications, the preprocessing module 52 may include, but is not limited to: determining the interaction type based on the interaction content, recognizing the robot's emotion based on the interaction content, recognizing the robot's persona based on the interaction content, determining the robot's expression based on the interaction content, determining the robot's action based on the interaction content, performing a knowledge base retrieval based on the interaction content, etc. After determining the required preprocessing type according to the configuration information, the corresponding preprocessing can be performed to extract the corresponding information, such as information related to long-term memory, information related to the knowledge base, the interaction type, etc.

[0131] The above session processing module 53 is used to extract the relevant information of the preset number of dialogue turns according to the interaction content when the interaction type is determined to be the session type.

[0132] For the relevant content of the session processing module 53, reference may be specifically made to the content of S103, which will not be repeated here.

[0133] The above template splicing module 54 is used to generate a prompt message according to the relevant information of long-term memory and the relevant information of the preset number of dialogue turns, and input the prompt message into a large language model to obtain a reply content corresponding to the prompt message.

[0134] Of course, the above template splicing module 54 may also be used to generate a prompt message according to the relevant information of long-term memory, the relevant information of the knowledge base, and the relevant information of the preset number of dialogue turns, and input the prompt message into a large language model to obtain a reply content corresponding to the prompt message.

[0135] For the relevant content of the template splicing module 54, please refer to the content of S104 for details, which will not be repeated here.

[0136] The above reply module 55 can be used to feedback the reply content to the user.

[0137] For the relevant content of the reply module 55, please refer to the content of S105 for details, which will not be repeated here.

[0138] The above post-processing module 56 is used to extract relevant information of the long-term memory through the long-term memory retrieval agent.

[0139] For the relevant content of the post-processing module 56, please refer to the relevant description of the long-term memory retrieval agent, which will not be repeated here. It can be understood that the above long-term memory retrieval agent can be configured in the post-processing module 56.

[0140] The above database 57 is used to store relevant information of the above long-term memory.

[0141] As can be seen above, the interaction system provided by the embodiments of the present application can also extract long-term memory through the agent, enabling the robot to effectively remember historical interaction information, user preferences, and details in specific scenarios. When interacting with the user, it can provide more personalized and accurate services and responses, and combines short-term memory to extract dialogue information in historical conversations, enabling the robot to quickly recall interaction topics and related key points in multi-round conversations, thereby effectively improving communication efficiency and coherence.

[0142] Please refer to Figure 6 , Figure 6 which shows a schematic structural diagram of another interaction system provided by the embodiments of the present application. As shown in Figure 6 the interaction system 50 provided by the embodiments of the present application may further include a task processing module 58.

[0143] Among them, the task processing module 58 can be used to disassemble tasks and execute the disassembled subtasks when it is determined that the interaction type is a task type.

[0144] As can be seen above, the interaction system provided by the embodiments of the present application can also implement task disassembly and execution based on the large model when responding to tasks, which can effectively improve the response efficiency and accuracy.

[0145] Corresponding to the interaction method of the robot described in the above embodiments, Figure 7 which shows a block diagram of the structure of the robot provided by the embodiments of the present application. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown.

[0146] Refer to Figure 7, the robot 70 includes: an acquisition module 701, a long-term memory extraction module 702, a short-term memory extraction module 703, a generation module 704, and a feedback module 705.

[0147] The acquisition module 701 is used to acquire the interaction content input by the user;

[0148] The long-term memory extraction module 702 is used to preprocess the interaction content and output the relevant information of the long-term memory corresponding to the interaction content;

[0149] The short-term memory extraction module 703 is used to extract the relevant information of a preset number of dialogue turns according to the interaction content when it is determined that the interaction type is a conversation type;

[0150] The generation module 704 is used to generate a prompt message according to the relevant information of the long-term memory and the relevant information of the preset number of dialogue turns, and input the prompt message into a large language model to obtain a reply content corresponding to the prompt message;

[0151] The feedback module 705 is used to feedback the reply content to the user.

[0152] In a possible implementation manner, the relevant information of the long-term memory is extracted by a long-term memory retrieval agent.

[0153] In a possible implementation manner, the above-mentioned generation module 704 may include a determination unit and a splicing unit.

[0154] The determination unit is used to determine an input template according to configuration information;

[0155] The splicing unit is used to splice the relevant information of the long-term memory and the relevant information of the preset number of dialogue turns into the input template according to the format of the input template to obtain the prompt message.

[0156] In a possible implementation manner, the robot further includes a determination module, and the above-mentioned determination module is used to determine a preprocessing type according to configuration information.

[0157] In a possible implementation manner, the robot further includes an interaction type determination module, and the interaction type determination module is used to determine an interaction type according to the interaction content.

[0158] In a possible implementation manner, the robot further includes a task execution module, and the above-mentioned task execution module is used to disassemble a task and execute the disassembled subtasks when it is determined that the interaction type is a task type.

[0159] It should be noted that for the content such as information interaction and execution process between the above-mentioned devices / units, since it is based on the same concept as the method embodiment of the present application, for its specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details will not be elaborated here.

[0160] Based on this, the robot provided by the embodiment of the present application can also extract long-term memory through an agent, enabling the robot to effectively remember historical interaction information, user preferences, and details in specific scenarios. When interacting with users, it can provide more personalized and accurate services and responses. Moreover, by combining short-term memory to extract dialogue information in historical conversations, the robot can quickly recall interaction topics and relevant key points in multi-round conversations, thereby effectively improving communication efficiency and coherence.

[0161] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments, and details will not be elaborated here.

[0162] Figure 8 It is a schematic structural diagram of a robot provided by another embodiment of the present application. As Figure 8 shown, the robot 8 of this embodiment includes: at least one processor 80 ( Figure 8 only one is shown in the figure), a processor, a memory 81, and a computer program 82 stored in the memory 81 and executable on the at least one processor 80. When the processor 80 executes the computer program 82, it implements the steps in any of the above-mentioned interaction method embodiments of the robot.

[0163] The robot may include, but is not limited to, a processor 80 and a memory 81. Those skilled in the art can understand that Figure 8 this is only an example of the robot 8 and does not constitute a limitation on the robot 8. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0164] The so-called processor 80 may be a Central Processing Unit (CPU), and the processor 80 may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0165] In some embodiments, the memory 81 may be an internal storage unit of the robot 8, such as the hard disk or memory of the robot 8. In other embodiments, the memory 81 may also be an external storage device of the robot 8, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the robot 8. Further, the memory 81 may also include both the internal storage unit of the robot 8 and the external storage device. The memory 81 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of the computer program. The memory 81 may also be used to temporarily store data that has been output or will be output.

[0166] The embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.

[0167] The embodiment of the present application provides a computer program product, and when the computer program product runs on a robot, the robot can execute to implement the steps in the above-mentioned various method embodiments.

[0168] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of this application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the robot, a recording medium, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electrical carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0169] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0170] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this document can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0171] In the embodiments provided in this application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.

[0172] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0173] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included within the protection scope of the present application.

Claims

1. A robot interaction method, characterized in that: include: Get the interactive content entered by the user; Preprocessing the interactive content, and outputting relevant information of the long-term memory corresponding to the interactive content; When it is determined that the interaction type is a conversation type, extracting relevant information of a preset number of dialogue rounds according to the interaction content; Generate prompt information according to the relevant information of the long-term memory and the relevant information of the preset number of dialogue turns, and input the prompt information into the large language model to obtain reply content corresponding to the prompt information; The reply content is fed back to the user.

2. The robot interaction method according to claim 1, characterized in that: The relevant information of the long-term memory is extracted by a long-term memory retrieval agent.

3. The robot interaction method according to claim 1, characterized in that: After obtaining the interactive content input by the user, the method further includes: Determine the preprocessing type based on the configuration information.

4. The robot interaction method according to any one of claims 1 to 3, characterized in that: After obtaining the interactive content input by the user, it also includes: An interaction type is determined according to the interaction content.

5. The robot interaction method according to claim 4, characterized in that: After determining the interaction type according to the interaction content, the method further includes: When the interaction type is determined to be a task type, the task is decomposed and the decomposed subtasks are executed.

6. The robot interaction method according to any one of claims 1 to 5, characterized in that: The generating of prompt information according to the relevant information of the long-term memory and the relevant information of the preset number of dialogue rounds includes: Determine the input template according to the configuration information; According to the format of the input template, the relevant information of the long-term memory and the relevant information of the preset number of dialogue rounds are spliced ​​into the input template to obtain the prompt information.

7. The robot interaction method according to claim 2, characterized in that: The long-term memory retrieval agent includes a plurality of simple agents.

8. A robot interaction system, characterized in that: include: A configuration module, used to configure the robot interaction system according to the configuration information; A preprocessing module, used to preprocess the interactive content input by the user according to the preprocessing type determined by the configuration information to obtain various preprocessing results; A conversation processing module, configured to extract relevant information of a preset number of conversation rounds according to the interaction content when the interaction type is determined to be a conversation type; A template splicing module is used to generate prompt information based on the relevant information of the long-term memory and the relevant information of the preset number of dialogue turns, and input the prompt information into the large language model to obtain the reply content corresponding to the prompt information; The reply module is used to feed back the reply content to the user.

9. A robot comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer program product, characterized in that When the computer program product runs on a robot, the robot is caused to perform the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Memory-based man-machine interaction method, electronic equipment, medium and program product

    CN121560166A