Human-machine interaction method, apparatus, electronic device and storage medium

By integrating a memory bank system with long-term and short-term components, large-scale language models can enhance interaction efficiency and personalization, addressing memory limitations and dialogue inconsistencies.

JP7799929B2Active Publication Date: 2026-01-16BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024146577
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-08-28
Publication Date
2026-01-16
Estimated Expiration
2044-08-28

AI Technical Summary

Technical Problem

Current large-scale language models lack sufficient memory capacity, limiting their ability to maintain long-term interactions and provide personalized experiences, leading to inconsistencies in dialogue and reduced user engagement.

Method used

Integrate a memory bank system comprising long-term and short-term memory banks to store user-specific information, allowing large-scale language models to retrieve and combine historical data for more personalized and consistent responses.

Benefits of technology

Enhances interaction efficiency and personalization by providing large-scale language models with memory capabilities, improving dialogue consistency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007799929000001
    Figure 0007799929000001
  • Figure 0007799929000002
    Figure 0007799929000002
  • Figure 0007799929000003
    Figure 0007799929000003
Patent Text Reader

Abstract

To provide a human-machine interaction method capable of grasping a complex language structure using a large-scale language model and generating smooth and consistent text, and indicating outstanding ability in text processing and generation, an apparatus, an electronic device, and a storage medium.SOLUTION: A human-machine interaction method includes: step 101 of acquiring a question input by a user during conversation with a large-scale language model; step 102 of searching for memory information in a memory bank, the memory information being historical memory information about the user; and step 103 of generating answer information corresponding to the question using the large-scale language model in conjunction with the matched memory information in response to retrieved memory information required to generate the answer information corresponding to the question, taking the searched memory information as matched memory information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to the field of artificial intelligence, and in particular to human-machine interaction methods, apparatuses, electronic devices, and storage media in fields such as natural language processing, large-scale language models, and deep learning. [Background technology]

[0002] Currently, large language models (LLMs) are able to understand complex linguistic structures and generate smooth, coherent text, demonstrating exceptional capabilities in text processing and generation. Summary of the Invention [Problem to be solved by the invention]

[0003] The present invention provides a human-machine interaction method, apparatus, electronic device and storage medium. [Means for solving the problem]

[0004] There is provided a human-machine interaction method including: acquiring a question input by a user in an interaction with a large-scale language model; searching for stored information in a memory bank, the stored information being historical stored information related to the user; and, in response to the search for stored information necessary for generating response information corresponding to the question, using the searched stored information as matching stored information and combining the matching stored information with the large-scale language model to generate the response information.

[0005] a question acquisition module that acquires a question input by a user in interaction with a large-scale language model; an information processing module for searching stored information in a memory bank, said stored information being historical stored information relating to said user; A human-machine interaction device is provided, which includes: a response generation module that, in response to retrieval of memory information necessary for generating response information corresponding to the question, uses the retrieved memory information as matching memory information and combines it with the matching memory information to generate the response information using a large-scale language model.

[0006] There is provided an electronic device comprising at least one processor and a memory communicatively coupled to the at least one processor, the memory storing instructions executable by the at least one processor, the instructions, when executed by the at least one processor, causing the at least one processor to perform the method set forth above.

[0007] A non-transitory computer-readable storage medium is provided that stores computer instructions for causing a computer to perform the above method.

[0008] A computer program product is provided which includes computer programs / instructions which, when executed by a processor, implement the above method.

[0009] It should be understood that the contents described in this section are not intended to identify key or essential features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will be readily apparent from the following specification. [Brief explanation of the drawings]

[0010] The drawings are for better understanding of the present technical solution and are not intended to limit the present application. [Figure 1] 1 is a flowchart of an embodiment of a human-machine interaction method described in the present disclosure. [Figure 2] FIG. 1 is a schematic diagram of the relationship between a user, a memory bank, and a large-scale language model as described in this disclosure. [Figure 3]3 is a schematic diagram of the configuration of an embodiment 300 of a human-machine interaction device described in the present disclosure. [Figure 4] 4 shows a schematic block diagram of an electronic device 400 in which embodiments of the present disclosure can be implemented. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, exemplary embodiments of the present application will be described based on the drawings. For ease of understanding, various details of the embodiments of the present application are included and should be considered as merely examples. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of brevity, the following description will omit descriptions of well-known functions and structures.

[0012] Furthermore, the term "and / or" in this specification simply describes a relationship between related objects and means that three relationships can exist. For example, A and / or B can mean three situations: A exists alone, A and B exist simultaneously, and B exists alone. Also, the character " / " in this specification generally means that the related objects before and after it are in an "or" relationship.

[0013] 1 is a flowchart of an embodiment of the human-machine interaction method described in the present disclosure. As shown in FIG. 1, the following specific implementation methods are included:

[0014] In step 101, a question input by a user in interaction with a large-scale language model is obtained.

[0015] Step 102 searches for stored information in a memory bank, said stored information being historical stored information relating to said user.

[0016] In step 103, in response to retrieval of the memory information necessary for generating response information corresponding to the question, the retrieved memory information is used as matching memory information, and the response information is generated in combination with the matching memory information using a large-scale language model.

[0017] Current large-scale language models still have some gaps from true artificial general intelligence (AGI), one of the main reasons for this is that they cannot have the same memory capacity as humans, which limits the interaction capabilities between the large-scale language models and users.

[0018] Meanwhile, by adopting the solution described in the above method embodiment, the large-scale language model can combine the retrieved memory information to generate response information corresponding to the question entered by the user, that is, by providing the large-scale language model with memory capability, the large-scale language model can be more efficient and personified in its interaction with the user, and the dialogue effect can be further improved.

[0019] A user-input question is a question that a user inputs into a large-scale language model in the process of interacting with the model. Examples of large-scale language models include chatbots and artificial intelligence (AI) assistants.

[0020] The memory bank can be searched for stored information in response to a question entered by a user. Preferably, the memory bank may include any or all of a long-term memory bank and a short-term memory bank. The long-term memory bank may include a user memory bank that can contain stored information that is user profile information of the user. The short-term memory bank may include a conversation memory bank that can contain stored information that is historical dialogue information between the user and the large-scale language model within a recent predetermined period.

[0021] Preferably, the user profile information may include user profile information generated based on either or all of historical interaction information between the user and the large-scale language model and user information of the user collected from predetermined data sources. Additionally, the user memory bank may include stored information regarding the user's identity settings and / or personalization requests voluntarily stored by the user.

[0022] That is, user profile information can be generated based on historical interaction information between the user and the large-scale language model and / or user information collected from a predetermined data source. The specific type of data source is not limited, and may be, for example, a related product. A user registers a product when using it. Part of the user information entered during registration can be used to generate the user profile information. Alternatively, part of the log information generated in the process of the user's use of the product can also be used as user information to generate the user profile information. Furthermore, there are no limitations on the specific information included in the user profile information, and it may include, for example, the user's name, gender, age, educational background, job, hobbies (preferences), etc. Here, hobbies may include, for example, a love of traveling or gourmet food.

[0023] Furthermore, the user can voluntarily store memorized information in the user memory bank, such as information related to the user's identity settings and / or personalization requirements, for example, "I am a programmer. From now on, when I write code, please use the ** format."

[0024] The conversation memory bank mainly stores historical dialogue information between the user and the large-scale language model within a recent predetermined period, i.e., historical dialogue information of question and answer responses, which may be determined according to actual needs, such as the recent one month.

[0025] Preferably, the long-term memory bank may further comprise a system memory bank that can include stored information of any one or any combination of identity setting information of the large-scale language model, background knowledge information of the large-scale language model, and reference document information of the large-scale language model.

[0026] Here, identity setting information refers to information about the identity of a large-scale language model, such as "You are an AI assistant developed by ** company." The background knowledge base refers to the capabilities of the large-scale language model. Reference documents refer to documents / knowledge that the large-scale language model can refer to when performing operations such as question and answering.

[0027] In traditional approaches, large-scale language models cannot effectively span long-term memory and cannot reference previous interaction information, limiting their application in situations requiring continuous dialogue or long-term memory. Furthermore, due to a lack of memory for a user's past preferences and interactions, they are unable to provide a personalized experience, i.e., there are limitations on personalization. Furthermore, due to a lack of long-term memory, large-scale language models are unable to maintain consistency of topic and context in continuous interactions, which affects the user experience. On the other hand, by adopting the processing method described in this disclosure, it is possible to search stored information in the user memory bank, the system memory bank, and the conversation memory bank, respectively. This allows the large-scale language model to have various memory capabilities, such as long-term user memory capacity, system memory capacity, and short-term conversation memory capacity, thereby effectively resolving the application limitations, personalization limitations, and consistency issues faced by large-scale language models in traditional approaches, and bringing large-scale language models closer to reality on the path to general artificial intelligence.

[0028] Preferably, the stored information in the system memory bank is stored in an unstructured format, and / or the user profile information in the user memory bank is stored in a key-value pair format, and / or the stored information in the conversation memory bank and the user's voluntarily stored stored information in the user memory bank may be stored in a vector format.

[0029] Among these, the stored information in the system memory bank is stored in an unstructured format, providing the context information and knowledge background required for the large-scale language model and helping to optimize the response quality and decision logic of the large-scale language model. User profile information can be stored in a structured key-value pair format, for example, where the key is "gender" and the value is "male." This structured storage method is advantageous for fast and accurate retrieval of stored information, improving the efficiency and accuracy of retrieval of stored information. The stored information in the conversation memory bank and the stored information voluntarily stored by the user in the user memory bank can be encoded and stored in a vector format, but there are no restrictions on how the encoding can be used. Representing stored information using vectorized stored information allows for efficient retrieval and comparison of stored information. In summary, the technical solution described in this disclosure provides high flexibility and versatility in storing stored information and can support the storage and processing of stored information in various formats to meet the needs of various scenarios, etc.

[0030] Preferably, upon receiving an operation instruction from a user for any of the memory banks, memory bank operations including adding new storage information, deleting existing storage information, and modifying existing storage information may be completed in accordance with the operation instruction.

[0031] That is, dynamic management (mainly management of addition, deletion, and modification) of stored information in a memory bank can be realized. It is possible to add new stored information to any memory bank, delete existing stored information, or modify existing stored information. This dynamic management can be realized manually, or in actual applications, this dynamic management can be realized automatically, such as deleting stored information that has not been used for a predetermined period of time. This dynamic management can improve the accuracy and reliability of stored information in a memory bank.

[0032] Preferably, the dialogue information generated between the user and the large-scale language model is stored in a conversation memory bank in real time, and when it is determined that the stored dialogue information matches the extraction conditions, key information is extracted from the stored dialogue information, and the stored dialogue information is replaced with the extracted key information.

[0033] For example, information on questions and responses that arise each time a user and a large-scale language model converse is stored in a conversation memory bank in real time, and every time three pieces of conversation information are stored, the extraction conditions are considered to be met, key information is extracted from those three pieces of conversation information, and those three pieces of conversation information are replaced with the extracted key information.The three pieces of conversation information are then stored again, and the above process is repeated.

[0034] In other words, by optimizing and integrating dialogue information and extracting and storing a relatively small amount of useful information from a large amount of data, it is possible not only to save storage resources but also to improve search speed and accuracy.

[0035] Furthermore, preferably, when it is determined that the memory conversion conditions are met and it is determined that the user profile information in the user memory bank needs to be updated from the stored information in the conversation memory bank, the user profile information may be updated based on the stored information in the conversation memory bank.

[0036] For example, it may be determined that the memory conversion conditions are met at the zero point of each day. Furthermore, it may be determined from the stored information in the conversation memory bank whether the user profile information in the user memory bank needs to be updated, and if necessary, updates such as corrections or additions to the user's preferences may be made.

[0037] The above process enables the conversion of short-term conversational memory into long-term memory, which means that not only can we store short-term conversational memories between a user and a large-scale language model, but we can also archive short-term conversational memories as long-term memories for future interactions, thereby enabling us to continuously learn and update the user's user profile information to make it more comprehensive and accurate.

[0038] As described above, in response to retrieval of the memory information necessary for generating response information to the question, the retrieved memory information can be used as matching memory information, and the response information can be generated in combination with the matching memory information using a large-scale language model.

[0039] There are no limitations on how to determine which memory information is necessary. For example, the necessary memory information can be determined based on the question entered by the user, the historical dialogue information of the current dialogue, etc. Specifically, for each memory bank, one corresponding neural networking model is trained in advance, and then, based on the question entered by the user and the historical dialogue information of the current dialogue, the neural networking model can be used to determine matching memory information from the corresponding memory bank.

[0040] Preferably, in response to the matching stored information being acquired from only one memory bank, the matching stored information may be converted into a predetermined format to obtain a first conversion result, and the response information may be generated in combination with the first conversion result using a large-scale language model. In response to the matching stored information being acquired from at least two different memory banks, the matching stored information acquired from the different memory banks may be integrated, the integrated result may be converted into a predetermined format to obtain a second conversion result, and the response information may be generated in combination with the second conversion result using a large-scale language model.

[0041] For example, if matching memory information is obtained from a user memory bank and a conversation memory bank, the two can be integrated. There are no restrictions on how the information is integrated. The integration process simplifies the memory information input to the large-scale language model, thereby improving the processing efficiency of the large-scale language model. Furthermore, the obtained integration result can be further converted into a format. There are also no restrictions on the specific format of the converted format; for example, it can be a format that is easier for the large-scale language model to recognize and understand. Furthermore, the conversion result can be used as additional input to the large-scale language model to generate the response information. If matching memory information is obtained from one of the memory banks, the matching memory information can be directly converted into a format, and the conversion result can then be used as additional input to the large-scale language model. In other words, by performing processes such as integration and format conversion, it is possible to expect improvements in the processing efficiency of the large-scale language model and the accuracy of the processing results.

[0042] In particular, if matching stored information is not available, the response information can be generated using a large-scale language model as in the past.

[0043] In relation to the above description, Figure 2 is a schematic diagram of the relationship between a user, a memory bank, and a large-scale language model described in the present disclosure. As shown in Figure 2, the memory bank includes a long-term memory bank and a short-term memory bank, the long-term memory bank includes a user memory bank and a system memory bank, and the short-term memory bank includes a conversation memory bank. For a question entered by a user, matching memory information can be obtained by searching the memory information in the memory bank, and then the matching memory information can be combined with the large-scale language model to generate response information to the question. As can be seen from the above, the large-scale language model in the technical solution described in the present disclosure is a large-scale language model with a memory enhancement function.

[0044] The technical solutions described in this disclosure can be further exemplified as follows. User: I'm Akira. What's your name? Large-scale language models: I'm an AI assistant developed by a company. Can I help you? User: I'm planning to travel to Region A (a certain country) next week. Do you have any suggestions? Large-scale language model: Of course. What kind of travel suggestions would you be interested in? Gourmet? User: Yes, I'm interested in gastronomy. Large-scale language model: In region A, we recommend you try traditional pizza and pasta. For example, in region B (one city in region A), there is a very famous restaurant called ***, and in region C (another city in region A), we recommend you visit restaurant ***, which serves unique pasta flavors. User: Sounds good, I'll give it a try. (One month later) User: Hello. Large-scale language model: Hello, Akira, have you traveled to region A? User: Yes, I went. I tried the food you recommended. It was delicious. Thank you. Large-scale Language Model: You're welcome. If you want to know more gourmet food in the future, please feel free to ask me. .....

[0045] In the above example, the large-scale language model introduces itself as "an AI assistant developed by ** Company" based on the memory information stored in the system memory bank, and learns that Ming is a gourmet based on the user profile information stored in the user memory bank. Accordingly, it can ask Ming if he is interested in gourmet food. Furthermore, in subsequent interactions, it can refer to previous interaction information to trace back historical interaction information between Ming and the large-scale language model. For example, if it learns based on the memory information stored in the conversation memory bank that Ming previously wanted to travel to region A, it can ask Ming, "Have you traveled to region A?", thereby providing a more personalized and consistent interaction experience for user Ming.

[0046] Although the above-described method embodiments are expressed as a combination of a series of operations for the sake of simplicity, those skilled in the art will appreciate that some steps according to the present application can be performed in other orders or simultaneously, and therefore the present application is not limited to the order of operations described. Next, those skilled in the art should understand that all of the embodiments described in the specification are preferred embodiments, and related operations and modules are not necessarily required by the present application.

[0047] The above describes the method embodiment, and the following uses the apparatus embodiment to further describe the technical solution described in this disclosure.

[0048] 3 is a schematic diagram illustrating a configuration structure of an embodiment 300 of a human-machine interaction device described in the present disclosure. As shown in FIG. 3, the device includes a question acquisition module 301 that acquires a question input by a user in a dialogue with a large-scale language model, an information processing module 302 that searches for stored information in a memory bank, which is historical stored information related to the user, and a response generation module 303 that, upon retrieval of stored information necessary for generating response information corresponding to the question, uses the retrieved stored information as matching stored information and combines the matching stored information with the large-scale language model to generate the response information.

[0049] By adopting the solution described in the embodiment of the device, the large-scale language model can combine the retrieved memory information to generate response information corresponding to the question entered by the user. That is, by providing the large-scale language model with memory capability, the large-scale language model can be more efficient and personified in its interaction with the user, and the dialogue effect can be further improved.

[0050] The information processing module 302 can search for stored information in the memory bank in response to a question input by a user.

[0051] Preferably, the memory banks may include any or all of a long-term memory bank and a short-term memory bank. The long-term memory bank may include a user memory bank that can contain stored information that is user profile information of a user. The short-term memory bank may include a conversation memory bank that can contain historical dialogue information between a user and a large-scale language model within a recent predetermined period of time.

[0052] Preferably, the user profile information may include user profile information generated based on any or all of the user's historical interaction information with the large-scale language model and / or user information of the user collected from predetermined data sources. Additionally, the user memory bank may include stored information regarding the user's identity settings and / or personalization requests voluntarily stored by the user.

[0053] Preferably, the long-term memory bank may comprise a system memory bank that can contain stored information of any one or any combination of identity setting information for the large language model, background knowledge information for the large language model, and reference document information for the large language model.

[0054] Preferably, the stored information in the system memory bank is stored in an unstructured format, and / or the user profile information in the user memory bank is stored in a key-value pair format, and / or the stored information in the conversation memory bank and the user's voluntarily stored stored information in the user memory bank may be stored in a vector format.

[0055] Furthermore, preferably, the information processing module 302 may, upon receiving an operation instruction from a user for any of the memory banks, complete memory bank operations including adding new storage information, deleting existing storage information, and modifying existing storage information in accordance with the operation instruction.

[0056] Preferably, the information processing module 302 stores the dialogue information generated between the user and the large-scale language model in a conversation memory bank in real time, and, upon determining that the stored dialogue information matches the extraction conditions, extracts key information from the stored dialogue information and replaces the stored dialogue information with the extracted key information.

[0057] Furthermore, preferably, the information processing module 302 may update the user profile information based on the stored information in the conversation memory bank in response to determining that the memory conversion condition is met and that it is necessary to update the user profile information in the user memory bank from the stored information in the conversation memory bank.

[0058] The response generation module 303 can use the stored information retrieved from the memory bank as matching stored information, and generate the response information by combining the matching stored information with the large-scale language model.

[0059] Preferably, the response generation module 303 may, in response to the matching stored information being acquired from only one memory bank, convert the matching stored information into a predetermined format to obtain a first conversion result, and combine the first conversion result with the large-scale language model to generate the response information. In response to the matching stored information being acquired from at least two different memory banks, the response generation module 303 may, in response to the matching stored information being acquired from the different memory banks, integrate the integrated result into a predetermined format to obtain a second conversion result, and further combine the second conversion result with the large-scale language model to generate the response information.

[0060] The specific workflow of the device embodiment shown in FIG. 3 can be referred to the relevant description of the method embodiment above, and will not be described in detail here.

[0061] In summary, the technical solutions described in this disclosure employ a memory mechanism that enables large-scale language models to provide more accurate and customized outputs across time and interactions, further improving interaction effectiveness and making large-scale language models closer to reality on the road to general artificial intelligence.

[0062] The technical solutions described in this application can be applied to the field of artificial intelligence, particularly to fields such as natural language processing, large-scale language models, and deep learning. Artificial intelligence is a discipline that studies how computers simulate human thought processes and intelligent behaviors (e.g., learning, reasoning, thinking, planning, etc.), and includes both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing. Artificial intelligence software technologies mainly include several directions, such as computer vision technology, speech recognition technology, natural language processing technology, machine learning / deep learning, big data processing technology, and knowledge mapping technology.

[0063] The stored information in the embodiments described herein is not targeted to a specific user and does not reflect the personal information of a specific user. Furthermore, the implementer of the technical solution described herein may obtain the stored information in a variety of disclosure and legally compliant ways. The acquisition, storage, application, processing, transmission, provision, and distribution of the personal information of users involved in the technical solution described herein all comply with the provisions of relevant laws and regulations and are not contrary to public order and morals.

[0064] According to embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.

[0065] 4 is a schematic block diagram of an electronic device 400 that may be used to implement embodiments of the present disclosure. The electronic device represents various forms of digital computers, such as laptops, desktop computers, workbenches, servers, blade servers, mainframes, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as PDAs, mobile phones, smartphones, wearable devices, and other similar computing devices. The components, their connections and relationships, and their functions shown herein are merely exemplary and are not intended to limit the implementation of the present disclosure as described and / or claimed herein.

[0066] 4, device 400 includes a computing means 401 that can perform various appropriate operations and processes in accordance with a computer program stored in a read-only memory (ROM) 402 or loaded from a storage means 408 into a random access memory (RAM) 403. The RAM 403 may store various programs and data necessary for the operation of device 400. The computing means 401, ROM 402, and RAM 403 are connected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0067] Several components of device 400 are connected to I / O interface 405, including input means 404, e.g., a keyboard, a mouse, etc., output means 407, e.g., various types of displays, speakers, etc., storage means 408, e.g., a magnetic disk, an optical disk, etc., and communication means 409, e.g., a network card, a modem, a wireless communication transceiver, etc. The communication means 409 enables device 400 to exchange information / data with other devices via, e.g., a computer network of the Internet and / or various telecommunication networks.

[0068] The computing means 401 may be various general-purpose and / or specialized processing components having processing and computing capabilities. Some examples of the computing means 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that execute machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing means 401 executes various methods and processes described above, such as the methods described herein. For example, in some embodiments, the methods described herein may be implemented as a computer software program physically embodied in a machine-readable medium, such as the storage means 408. In some embodiments, some or all of the computer program can be loaded and / or installed into the device 400 via the ROM 402 and / or the communication means 409. When the computer program is loaded into the RAM 403 and executed by the computing means 401, it can perform one or more steps of the methods described herein. Alternatively, in other embodiments, the computing means 401 may be configured in any other suitable manner (eg, via firmware) to perform the methods described in the present disclosure.

[0069] Various embodiments of the systems and techniques described herein may be implemented in digital electronic circuitry systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), field programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include being embodied in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor. The programmable processor may be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and that can transfer data and instructions to the storage system, the at least one input device, and the at least one output device.

[0070] Program code for implementing the methods of the present disclosure can be written using any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus such that, when executed by the processor or controller, the program code performs the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, as a stand-alone package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0071] In the context of this disclosure, a machine-readable medium is a tangible medium that can contain or store a program used by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. More specific examples of machine-readable storage media include one or more line-based electrical connections, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0072] To provide interaction with a user, the systems and techniques described herein may be implemented on a computer that includes a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball) for providing input by the user to the computer. Other types of devices may also be used to provide interaction with a user. For example, feedback provided to the user may be any form of sensing feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input from the user may be received in any form (including sound input, speech input, or tactile input).

[0073] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a client computer having a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include a local area network ("LAN"), a wide area network ("WAN"), and an internetwork.

[0074] The computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server arises by virtue of computer programs running on the corresponding computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server in a blockchain combination.

[0075] It should be understood that steps can be rearranged, added, or deleted using the various types of flows shown above. For example, the steps described in this application can be performed in a parallel order, a sequential order, or can be performed in a different order, and are not limited thereto, as long as the desired results of the technical solution disclosed in this application can be achieved.

[0076] The above specific embodiments do not constitute limitations on the scope of protection of the present application. Those skilled in the art should understand that various modifications, combinations, partial combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application. (Other possible items) (Item 1) Obtaining a question input by a user in interaction with a large-scale language model; retrieving stored information in a memory bank, said stored information being historical stored information relating to said user; In response to retrieving stored information necessary for generating response information corresponding to the question, the retrieved stored information is used as matching stored information, and the response information is generated by combining the searched stored information with the matching stored information using the large-scale language model; A human-machine interaction method including: (Item 2) the memory banks include any or all of a long-term memory bank and a short-term memory bank; the long-term memory bank comprises a user memory bank containing user profile information of the user as stored information; the short-term memory bank comprises a conversation memory bank containing, as stored information, historical dialogue information between the user and the large-scale language model within a recent predetermined period of time; The method according to item 1. (Item 3) the user profile information includes user profile information generated based on any or all of historical interaction information between the user and the large-scale language model and user information of the user collected from a predetermined data source; and / or the user memory bank further includes stored information relating to the user's identity settings and / or personalization requests, voluntarily stored by the user; The method described in item 2. (Item 4) the long-term memory bank further comprises a system memory bank; The system memory bank stores, as stored information, one or any combination of identity setting information of the large-scale language model, background knowledge information of the large-scale language model, and reference document information of the large-scale language model. The method according to item 3. (Item 5) the stored information in the system memory banks is stored in an unstructured format; and / or the user profile information in the user memory bank is stored in a key-value pair format; and / or The stored information in the conversation memory bank and the user's voluntarily stored stored information in the user memory bank are stored in a vector format. The method according to item 4. (Item 6) In response to an operation instruction from the user for any one of the memory banks being acquired, completing a memory bank operation including adding new storage information, deleting existing storage information, and modifying existing storage information based on the operation instruction; Item 5. The method according to item 4, further comprising: (Item 7) storing dialogue information generated between the user and the large-scale language model in the conversation memory bank in real time, extracting key information from the stored dialogue information in response to a determination that the stored dialogue information matches an extraction condition, and replacing the stored dialogue information with the extracted key information; The method according to item 2, further comprising: (Item 8) updating the user profile information based on the stored information in the conversation memory bank in response to determining that the storage conversion condition is met and determining that the user profile information in the user memory bank needs to be updated based on the stored information in the conversation memory bank; Item 4. The method according to item 3, further comprising: (Item 9) generating the response information using the large-scale language model in combination with the matching stored information, In response to the matching stored information being obtained from only one memory bank, converting the matching stored information into a predetermined format to obtain a first conversion result, and combining the first conversion result with the large-scale language model to generate the response information; In response to the matching stored information being acquired from at least two different memory banks, integrating the matching stored information acquired from the different memory banks, converting the integration result into a predetermined format to obtain a second conversion result, and combining the second conversion result with the large-scale language model to generate the response information; 9. The method according to any one of items 1 to 8, comprising: (Item 10) a question acquisition module that acquires a question input by a user in interaction with a large-scale language model; an information processing module for searching stored information in a memory bank, the stored information being historical stored information relating to said user; a response generation module that, in response to retrieval of stored information necessary for generating response information corresponding to the question, uses the retrieved stored information as matching stored information, and generates the response information by combining the searched stored information with the matching stored information and using a large-scale language model; A human-machine interaction device comprising: (Item 11) the memory banks include any or all of a long-term memory bank and a short-term memory bank; the long-term memory bank comprises a user memory bank containing user profile information of the user as stored information; the short-term memory bank comprises a conversation memory bank containing, as stored information, historical dialogue information between the user and the large-scale language model within a recent predetermined period of time; Item 11. The device according to item 10. (Item 12) the user profile information includes user profile information generated based on any or all of historical interaction information between the user and the large-scale language model and user information of the user collected from a predetermined data source; and / or the user memory bank further includes stored information relating to the user's identity settings and / or personalization requests, voluntarily stored by the user; Item 12. The device according to item 11. (Item 13) the long-term memory bank further comprises a system memory bank; The system memory bank stores, as stored information, one or any combination of identity setting information of the large-scale language model, background knowledge information of the large-scale language model, and reference document information of the large-scale language model. Item 13. The device according to item 12. (Item 14) the stored information in the system memory banks is stored in an unstructured format; and / or the user profile information in the user memory bank is stored in a key-value pair format; and / or The stored information in the conversation memory bank and the user's voluntarily stored stored information in the user memory bank are stored in a vector format. Item 14. The device according to item 13. (Item 15) the information processing module further completes a memory bank operation including adding new storage information, deleting existing storage information, and modifying existing storage information in response to an operation instruction from the user for any of the memory banks being acquired, based on the operation instruction; Item 14. The device according to item 13. (Item 16) The information processing module further stores dialogue information generated between the user and the large-scale language model in the conversation memory bank in real time, and, in response to determining that the stored dialogue information matches an extraction condition, extracts key information from the stored dialogue information and replaces the stored dialogue information with the extracted key information. Item 12. The device according to item 11. (Item 17) The information processing module further updates the user profile information based on the stored information in the conversation memory bank in response to determining that the storage conversion condition is met and determining that the user profile information in the user memory bank needs to be updated based on the stored information in the conversation memory bank. Item 13. The device according to item 12. (Item 18) the response generation module, in response to the matching stored information being acquired from only one memory bank, converts the matching stored information into a predetermined format to obtain a first conversion result, and generates the response information in combination with the first conversion result using the large-scale language model; and, in response to the matching stored information being acquired from at least two different memory banks, integrates the matching stored information acquired from the different memory banks, converts the integration result into a predetermined format to obtain a second conversion result, and generates the response information in combination with the second conversion result using the large-scale language model. The device according to any one of items 10 to 17. (Item 19) at least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor performs the method according to any one of items 1 to 9. An electronic device. (Item 20) A non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method according to any one of items 1 to 9. (Item 21) A computer program that, when executed by a processor, implements the method according to any one of items 1 to 9.

Claims

1. A human-machine interaction method, comprising: Obtaining a question input by a user in interaction with a large-scale language model; retrieving stored information in a memory bank, said stored information being historical stored information relating to said user; In response to retrieving stored information necessary for generating response information corresponding to the question, the retrieved stored information is used as matching stored information, and the response information is generated by combining the searched stored information with the matching stored information using the large-scale language model; Including, the memory banks include short-term memory banks; the short-term memory bank comprises a conversation memory bank containing, as stored information, historical dialogue information between the user and the large-scale language model within a recent predetermined period; The human-machine interaction method includes: storing dialogue information generated between the user and the large-scale language model in the conversation memory bank in real time, extracting key information from the stored dialogue information in response to a determination that the stored dialogue information matches an extraction condition, and replacing the stored dialogue information with the extracted key information; The human-machine interaction method further comprises:

2. the memory banks further include a long-term memory bank; 2. The human-machine interaction method of claim 1, wherein the long-term memory bank comprises a user memory bank containing user profile information of the user as stored information.

3. the user profile information includes user profile information generated based on any or all of historical interaction information between the user and the large-scale language model and user information of the user collected from a predetermined data source; and / or the user memory bank further includes stored information relating to the user's identity settings and / or personalization requests, voluntarily stored by the user; 3. The human-machine interaction method according to claim 2.

4. the long-term memory bank further comprises a system memory bank; The system memory bank stores, as stored information, one or any combination of identity setting information of the large-scale language model, background knowledge information of the large-scale language model, and reference document information of the large-scale language model.

4. The human-machine interaction method according to claim 3.

5. the stored information in the system memory banks is stored in an unstructured format; and / or the user profile information in the user memory bank is stored in a key-value pair format; and / or The stored information in the conversation memory bank and the user's voluntarily stored stored information in the user memory bank are stored in a vector format.

5. A human-machine interaction method according to claim 4.

6. In response to an operation instruction from the user for any one of the memory banks being acquired, completing a memory bank operation including adding new storage information, deleting existing storage information, and modifying existing storage information based on the operation instruction; The human-machine interaction method of claim 4, further comprising:

7. updating the user profile information based on the stored information in the conversation memory bank in response to determining that the storage conversion condition is met and determining that the user profile information in the user memory bank needs to be updated based on the stored information in the conversation memory bank; The human-machine interaction method of claim 3, further comprising:

8. generating the response information using the large-scale language model in combination with the matching stored information, In response to the matching stored information being obtained from only one memory bank, converting the matching stored information into a predetermined format to obtain a first conversion result, and combining the first conversion result with the large-scale language model to generate the response information; In response to the matching stored information being acquired from at least two different memory banks, integrating the matching stored information acquired from the different memory banks, converting the integration result into a predetermined format to obtain a second conversion result, and combining the second conversion result with the large-scale language model to generate the response information; 2. The human-machine interaction method of claim 1, comprising:

9. a question acquisition module that acquires a question input by a user in interaction with a large-scale language model; an information processing module for searching stored information in a memory bank, the stored information being historical stored information relating to said user; a response generation module that, in response to retrieval of stored information necessary for generating response information corresponding to the question, uses the retrieved stored information as matching stored information, and generates the response information by combining the searched stored information with the matching stored information and using a large-scale language model; Equipped with the memory banks include short-term memory banks; the short-term memory bank comprises a conversation memory bank containing, as stored information, historical dialogue information between the user and the large-scale language model within a recent predetermined period; The information processing module further stores dialogue information generated between the user and the large-scale language model in the conversation memory bank in real time, and, in response to determining that the stored dialogue information matches an extraction condition, extracts key information from the stored dialogue information and replaces the stored dialogue information with the extracted key information. Human-machine interaction device.

10. the memory banks further include a long-term memory bank; 10. The human-machine interaction device of claim 9, wherein the long-term memory bank comprises a user memory bank containing user profile information of the user as stored information.

11. the user profile information includes user profile information generated based on any or all of historical interaction information between the user and the large-scale language model and user information of the user collected from a predetermined data source; and / or the user memory bank further includes stored information relating to the user's identity settings and / or personalization requests, voluntarily stored by the user; A human-machine interaction device according to claim 10.

12. the long-term memory bank further comprises a system memory bank; The system memory bank stores, as stored information, one or any combination of identity setting information of the large-scale language model, background knowledge information of the large-scale language model, and reference document information of the large-scale language model. A human-machine interaction device according to claim 11.

13. the stored information in the system memory banks is stored in an unstructured format; and / or the user profile information in the user memory bank is stored in a key-value pair format; and / or The stored information in the conversation memory bank and the user's voluntarily stored stored information in the user memory bank are stored in a vector format. A human-machine interaction device according to claim 12.

14. the information processing module further completes a memory bank operation including adding new storage information, deleting existing storage information, and modifying existing storage information in response to an operation instruction from the user for any of the memory banks being acquired, based on the operation instruction; A human-machine interaction device according to claim 12.

15. The information processing module further updates the user profile information based on the stored information in the conversation memory bank in response to determining that the storage conversion condition is met and determining that the user profile information in the user memory bank needs to be updated based on the stored information in the conversation memory bank. A human-machine interaction device according to claim 11.

16. the response generation module, in response to the matching stored information being acquired from only one memory bank, converts the matching stored information into a predetermined format to obtain a first conversion result, and generates the response information in combination with the first conversion result using the large-scale language model; and, in response to the matching stored information being acquired from at least two different memory banks, integrates the matching stored information acquired from the different memory banks, converts the integration result into a predetermined format to obtain a second conversion result, and generates the response information in combination with the second conversion result using the large-scale language model. A human-machine interaction device according to any one of claims 9 to 15.

17. at least one processor; a memory communicatively coupled to the at least one processor; An electronic device, wherein the memory stores instructions executable by the at least one processor, the instructions, when executed by the at least one processor, causing the at least one processor to perform the human-machine interaction method described in any one of claims 1 to 8.

18. A non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the human-machine interaction method according to any one of claims 1 to 8.

19. A computer program which, when executed by a processor, implements the human-machine interaction method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Data processing method for chat and related device

    CN114003702A

  • Human-machine intelligence chatting method with artificial intelligence and device therefor

    JP2017010517A

  • Learning device, information processing device, learning method, information processing method, and program

    WO2021176714A1