Human-computer method and apparatus, electronic device, and storage medium

By registering the plug-in in the large language model and determining the target plug-in related to the dialogue text entered by the user, and generating and entering the second dialogue text to obtain the reply text, the accuracy problem of the large language model when processing new data is solved, and the efficiency of the user's selection of plug-ins and the processing efficiency of the large language model is improved.

WO2025091924A1PCT designated stage expired Publication Date: 2025-05-08BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/099252
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-31
Filing Date
2024-06-14
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

When processing new data outside the training data set, the questions that cannot be answered accurately result in users needing to manually select plug-ins in multiple rounds of conversations, which increases the user's selection cost, and excessively long conversation text will exceed the input content limit of the large language model, resulting in long processing time and reduced output accuracy.

Method used

By determining the target plug-in related to the user-entered dialogue text from multiple plug-ins registered in the large language model, generating a second dialogue text based on the dialogue text and plug-in description text, and inputting it into the large language model to obtain reply text, simplifying the process of user selection of plug-ins and reducing the length of dialogue text.

Benefits of technology

It improves the efficiency of users selecting plug-ins in the dialogue system, reduces the length of dialogue text, and improves the processing efficiency and reply accuracy of large language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024099252_08052025_PF_FP_ABST
    Figure CN2024099252_08052025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a human-computer interaction method and apparatus, an electronic device and a storage medium, relating to the technical field of artificial intelligence, in particular to the technical fields of deep learning, natural language processing and large models. The specific implementation solution comprises: in response to a human-computer interaction request and on the basis of a first dialogue text comprised in the human-computer interaction request, determining from among a plurality of plug-ins registered in a large language model a first target plug-in related to the first dialogue text; obtaining a second dialogue text on the basis of the first dialogue text and a description text of the first target plug-in; and inputting the second dialogue text into the large language model to obtain a reply text.
Need to check novelty before this filing date? Find Prior Art

Description

Human-computer interaction method, device, electronic device, and storage medium

[0001] This application claims priority to Chinese patent application No. 202311433823.9 filed on October 31, 2023, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present disclosure relates to the field of artificial intelligence technology, specifically to the field of deep learning, natural language processing, and large model technology, and especially to a human-computer interaction method, device, electronic device, and storage medium. Background Art

[0003] A conversational system based on a large language model is an intelligent system built using deep learning and natural language processing technologies to simulate human conversation. By learning and training on large amounts of text data, it can capture the nuances of natural language and better understand user input, thereby generating more natural responses.

[0004] Summary of the Invention

[0005] The present disclosure provides a human-computer interaction method, device, electronic device, and storage medium.

[0006] According to one aspect of the present disclosure, a human-computer interaction method is provided, comprising: in response to a human-computer interaction request, based on a first dialogue text included in the human-computer interaction request, determining a first target plug-in related to the first dialogue text from a plurality of plug-ins registered in a large language model; obtaining a second dialogue text based on the first dialogue text and a description text of the first target plug-in; and inputting the second dialogue text into the large language model to obtain a reply text.

[0007] According to another aspect of the present disclosure, a human-computer interaction device is provided, comprising: a determination module for, in response to a human-computer interaction request, determining, based on a first dialogue text included in the human-computer interaction request, a first target plug-in related to the first dialogue text from a plurality of plug-ins registered in a large language model; a first processing module for obtaining a second dialogue text based on the first dialogue text and a description text of the first target plug-in; and a first input module for inputting the second dialogue text into the large language model to obtain a reply text.

[0008] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.

[0009] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method described above.

[0010] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the method described above when executed by a processor.

[0011] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0013] FIG1 schematically shows an exemplary system architecture to which the human-computer interaction method and apparatus according to an embodiment of the present disclosure can be applied.

[0014] FIG2 schematically shows a flow chart of a human-computer interaction method according to an embodiment of the present disclosure.

[0015] FIG3A schematically shows a schematic diagram of a matching process of a first target plug-in according to an embodiment of the present disclosure.

[0016] FIG3B schematically shows a schematic diagram of a matching process of a first target plug-in according to another embodiment of the present disclosure.

[0017] FIG4A schematically shows a schematic diagram of a reply text generation process according to an embodiment of the present disclosure.

[0018] FIG4B schematically shows a schematic diagram of a reply text generation process according to another embodiment of the present disclosure.

[0019] FIG5 schematically shows a flow chart of a human-computer interaction method according to another embodiment of the present disclosure.

[0020] FIG6 schematically shows a block diagram of a human-computer interaction device according to an embodiment of the present disclosure.

[0021] FIG7 shows a schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure. DETAILED DESCRIPTION

[0022] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0023] With the continuous development of artificial intelligence technology, dialogue systems, as a key application area, have been widely used in human-computer interaction, intelligent customer service, intelligent assistants, and other fields. Large language models, as an important component of dialogue systems, play a vital role in these systems.

[0024] The implementation of large language models primarily relies on deep learning and natural language processing technologies. Deep learning uses multi-layer neural networks and backpropagation algorithms to learn and train input text data to produce output. Natural language processing, on the other hand, processes and analyzes text data to process and understand natural language input, thereby enabling the functionality of the dialogue system.

[0025] In actual applications, large language models suffer from a common flaw of artificial neural networks: they are unable to accurately answer questions about new data outside the training dataset. To address this flaw, relevant technical personnel have begun to configure plug-ins for large language models. These plug-ins can serve as the "eyes and ears" of large language models, allowing them to access new, private, or specific information not included in the training data, so that they can better serve users. On the other hand, plug-ins can also combine the large language model's powerful content generation and contextual understanding capabilities to broaden the application areas of large language models and increase the credibility of the generated results. In addition, plug-ins can enable large language models to perform safe and restricted operations on their behalf, improving the practicality of the entire system.

[0026] In the related plug-in management method, users need to manually select the required plug-ins. During the user's conversation, the dialogue system can input the plug-in description of the user-selected plug-in into the large language model in the form of prompt words. In other words, the plug-in description of the user-selected plug-in is inserted into the context of the user's input conversation content. The processed conversation content is then input into the large language model. The large language model can determine which plug-in is required for the current conversation content and then call the corresponding plug-in to complete the relevant function.

[0027] As the number of plugins configured in a large language model increases, for example, there may be tens of thousands of plugins configured. This large number of plugins makes it difficult for users to cost-effectively select the plugins they need. Furthermore, during multiple rounds of conversations with users in a dialogue system based on a large language model, the input conversation content and output reply content of previous conversations are added to the context of subsequent conversations, giving the conversation content of new rounds a longer context. If the plugin descriptions of the multiple plugins selected by the user are inserted into the context of the user's input conversation content, the actual text input into the large language model will be too long, even exceeding the large language model's limit on the number of tokens for input content, resulting in longer processing time for the large language model and reduced accuracy of the output reply content.

[0028] In light of this, embodiments of the present disclosure provide a human-computer interaction method, apparatus, electronic device, and storage medium to at least partially address the aforementioned issues. The human-computer interaction method includes: responding to a human-computer interaction request, determining a first target plug-in associated with the first dialog text from a plurality of plug-ins registered in a large language model based on a first dialog text included in the human-computer interaction request; obtaining a second dialog text based on the first dialog text and a description text of the first target plug-in; and inputting the second dialog text into the large language model to obtain a reply text.

[0029] FIG1 schematically shows an exemplary system architecture to which the human-computer interaction method and apparatus according to an embodiment of the present disclosure can be applied.

[0030] It should be noted that FIG1 is merely an example of a system architecture to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but does not mean that the embodiments of the present disclosure cannot be applied to other devices, systems, environments, or scenarios. For example, in another embodiment, the exemplary system architecture to which the human-computer interaction method and apparatus may be applied may include a terminal device, but the terminal device may implement the human-computer interaction method and apparatus provided by the embodiments of the present disclosure without interacting with a server.

[0031] As shown in FIG1 , a system architecture 100 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 is used as a medium for providing communication links between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links.

[0032] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.

[0033] The client application of the dialogue system may be installed on the terminal devices 101, 102, 103. The dialogue system may be a dialogue system based on a large language model. Users can use the client application on the terminal devices 101, 102, 103 to perform text processing using the large language model.

[0034] The server 105 may be a server or a cloud server that provides various services. The backend application of the dialogue system may be installed on the server 105.

[0035] It should be noted that the human-computer interaction method provided in the embodiment of the present disclosure can generally be executed by the terminal device 101, 102, or 103. Accordingly, the human-computer interaction apparatus provided in the embodiment of the present disclosure can also be provided in the terminal device 101, 102, or 103.

[0036] Alternatively, the human-computer interaction method provided in the embodiment of the present disclosure may also be generally executed by the server 105. Accordingly, the human-computer interaction apparatus provided in the embodiment of the present disclosure may generally be provided in the server 105. The human-computer interaction method provided in the embodiment of the present disclosure may also be performed by a server or server cluster that is different from the server 105 and that is capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the human-computer interaction apparatus provided in the embodiment of the present disclosure may also be provided in a server or server cluster that is different from the server 105 and that is capable of communicating with the terminal devices 101, 102, 103 and / or the server 105.

[0037] It should be understood that the number of terminal devices, networks, and servers in Figure 1 is merely illustrative and any number of terminal devices, networks, and servers may be provided as required.

[0038] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure and application of user personal information involved comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.

[0039] In the technical solution disclosed herein, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.

[0040] FIG2 schematically shows a flow chart of a human-computer interaction method according to an embodiment of the present disclosure.

[0041] As shown in FIG. 2 , the method includes operations S210 to S230 .

[0042] In operation S210 , in response to a human-computer interaction request, based on a first dialogue text included in the human-computer interaction request, a first target plug-in related to the first dialogue text is determined from a plurality of plug-ins registered in a large language model.

[0043] In operation S220, a second dialogue text is obtained based on the first dialogue text and the description text of the first target plug-in.

[0044] In operation S230, the second conversation text is input into the large language model to obtain a reply text.

[0045] According to an embodiment of the present disclosure, a human-computer interaction request can be triggered when a user engages in a conversation using a dialogue system. The dialogue system can be configured with information input interfaces such as text input controls and audio input controls. For example, a user can enter conversation text in a text box included in the text input control. After completing the conversation text input, the user can click a send button included in the text input control to input the conversation text into the dialogue system. The dialogue system can trigger the human-computer interaction request upon receiving the conversation text. For another example, a user can press a conversation button included in an audio input control and input audio information through an audio receiving device configured in an electronic device equipped with the dialogue system. The input audio information can be converted into conversation text through natural language processing. After the user releases the conversation button, the converted conversation text can be input into the dialogue system. The dialogue system can trigger the human-computer interaction request upon receiving the conversation text. The dialogue system described above can be a dialogue system based on a large language model, that is, in the dialogue system, a large language model can be used to process the received conversation text.

[0046] According to an embodiment of the present disclosure, the first dialogue text is the dialogue text received by the dialogue system when the human-computer interaction request is triggered.

[0047] According to embodiments of the present disclosure, developers can pre-register multiple plug-ins for the large language model. After registering a plug-in, the large language model can use the plug-in's description to call it. The plug-in description may include, for example, a description of the plug-in's functionality and usage examples. The functional description may explain the plug-in's purpose. For example, for a weather query plug-in, the functional description may include, for example, "This plug-in outputs weather conditions based on the city and date entered by the user." Plugin usage examples may include, but are not limited to, correct and incorrect usage examples of the plug-in. Each usage example may include at least the form of the user's input text and the form of the large language model's output text. For example, for a weather query plug-in, the usage example may include, "Input: Query the weather in city A on day D of month C of year B; Output: The weather in city A on day D of month C of year B is sunny, with a temperature of E°C, a relative humidity of F%, and a wind force of G." The large language model may call the plug-in to process the conversation text by adjusting the conversation text to the same textual form as the input of the usage example, and then using the plug-in to process the adjusted conversation text to obtain the output text of the plug-in.

[0048] According to an embodiment of the present disclosure, a first target plug-in is determined from multiple plug-ins. For example, the conversation text is matched with the text of the input part of the usage example included in the description text of the plug-in, and the plug-in with the highest matching degree is determined. The plug-in with the highest matching degree is the first target plug-in.

[0049] According to an embodiment of the present disclosure, the second dialogue text can be obtained by concatenating the first dialogue text and the description text of the first target plug-in. Alternatively, the first dialogue text and the description text of the first target plug-in can be filled into the same text template to obtain the second dialogue text. Alternatively, the first dialogue text and the description text of the first target plug-in can be reconstructed using a large language model to obtain the second dialogue text. The second dialogue text can include the text content and feature information of the first dialogue text and the text content and feature information of the description text of the first target plug-in. The second dialogue text can serve as the text actually input into the large language model.

[0050] According to an embodiment of the present disclosure, as an optional implementation, in response to a human-computer interaction request, it is also possible to detect whether the first dialogue text can be implemented using the functions of the large language model itself. If it is determined that the first dialogue text can be implemented using the functions of the large language model itself, the large language model can be directly used to process the first dialogue text to obtain a reply text. If it is determined that the first dialogue text cannot be implemented using the functions of the large language model itself, the text content of the first dialogue text can be used to match the most appropriate first target plug-in, and then the functions of the first target plug-in can be used to generate the reply text.

[0051] According to an embodiment of the present disclosure, when a user inputs a new dialogue text into the dialogue system to trigger a human-computer interaction request, the most relevant first target plug-in can be matched from multiple registered plug-ins based on the text content of the first dialogue text. The second dialogue text obtained by fusing the first dialogue text and the description text of the first target plug-in is input into the large language model to obtain a reply text corresponding to the first dialogue text. By matching plug-ins based on the text content of the first dialogue text, the user no longer needs to actively select the required plug-in during the conversation, thereby effectively improving the user experience. At the same time, by screening the plug-ins, the dialogue system does not need to splice too much plug-in description text in the context of the dialogue text, thereby effectively reducing the length of the text actually input into the large language model and improving the processing efficiency and accuracy of the large language model.

[0052] The method shown in FIG. 2 will be further described below with reference to FIG. 3A to FIG. 3B , FIG. 4A to FIG. 4B and FIG. 5 in combination with specific embodiments.

[0053] According to an embodiment of the present disclosure, in order to satisfy the call of the large language model to the plug-in, the plug-in needs to be registered in the large language model. When registering the plug-in, an index can be created based on the description of the plug-in. Later, when performing human-computer interaction, the plug-in can be matched based on this index.

[0054] According to an embodiment of the present disclosure, taking the third target plug-in as an example, the plug-in registration process may include the following operations:

[0055] In response to the plug-in registration request, index information of the third target plug-in is generated based on the description text of the third target plug-in included in the plug-in registration request; and the index information of the third target plug-in is written into the plug-in information library.

[0056] According to an embodiment of the present disclosure, a plug-in registration request can be actively triggered by a user. Specifically, the user can import the third target plug-in to be registered into the dialogue system through command line instructions, plug-in registration controls on the front-end page of the dialogue system, etc., and trigger a plug-in registration request.

[0057] According to embodiments of the present disclosure, the format of plugin index information is not limited herein. For example, plugin index information may include one or more keywords, which may be extracted from the plugin description text. For another example, plugin index information may be represented as a vector, which may be a feature vector derived from the plugin description text.

[0058] According to an embodiment of the present disclosure, the plug-in information library may be any type of database, or the plug-in information library may also be represented as a data structure in other forms, which is not limited here.

[0059] According to an embodiment of the present disclosure, corresponding to the plug-in registration process, plug-in matching can also be achieved by utilizing the index information of each plug-in in the plug-in information library. Specifically, based on the first dialogue text included in the human-computer interaction request, determining a first target plug-in related to the first dialogue text from multiple plug-ins registered in the large language model can include the following operations:

[0060] Respective index information of multiple plug-ins is obtained from a plug-in information library; the first dialogue text is matched with the respective index information of the multiple plug-ins to obtain multiple matching results; and a first target plug-in is determined from the multiple plug-ins based on the multiple matching results.

[0061] FIG3A schematically shows a schematic diagram of a matching process of a first target plug-in according to an embodiment of the present disclosure.

[0062] As shown in FIG3A , a large language model may be configured with N registered plug-ins 301 , which may be represented as plug-in 1, plug-in 2, ..., plug-in N. A plug-in information library 302 may record N index information 303 corresponding to the N plug-ins 301 , for example, index information 1 corresponding to plug-in 1, index information 2 corresponding to plug-in 2, index information N corresponding to plug-in N, and so on.

[0063] According to an embodiment of the present disclosure, the first conversation text 304 can be used to match N index information 303 to obtain N matching results 305. For example, matching the first conversation text 304 with index information 1 can obtain matching result 1, matching the first conversation text 304 with index information 2 can obtain matching result 2, and so on.

[0064] According to an embodiment of the present disclosure, the N matching results 305 can all be represented as numerical values ​​within a fixed interval, and based on the calculation method of the matching results, the probability of matching represented by the matching results and the numerical value of the matching results have a fixed corresponding relationship. For example, the N matching results can all be between 0 and 1, and the closer the numerical value of the matching result is to 1, the more the matching result can indicate that the first conversation text and the plug-in are matched. Determining the first target plug-in 306 based on the N matching results is to compare the sizes of the N matching results to determine the matching result with the largest numerical value, and the plug-in corresponding to the matching result with the largest numerical value is the first target plug-in 306.

[0065] According to an embodiment of the present disclosure, a specific calculation method of the matching result may be related to a generation method and form of the index information when the plug-in is registered.

[0066] For example, when registering a plug-in, feature extraction can be performed on the plug-in's description text to obtain a feature vector of the description text, and the plug-in's index information can be generated based on the feature vector. That is, the plug-in's index information can include the feature vector of the plug-in's description text. Matching the first conversation text with the index information of multiple plug-ins to obtain multiple matching results can include the following operations:

[0067] Feature extraction is performed on the first dialogue text to obtain a feature vector of the first dialogue text; similarity calculation is performed on the feature vector of the first dialogue text and the feature vectors of the description texts of the multiple plug-ins to obtain multiple similarity calculation results; and multiple matching results are obtained based on the multiple similarity calculation results.

[0068] According to an embodiment of the present disclosure, similarity calculation may be implemented by using any inter-vector similarity calculation method, and any inter-vector similarity calculation method may include but is not limited to a cosine similarity calculation method, a correlation coefficient method, and the like.

[0069] According to an embodiment of the present disclosure, the larger the value of the similarity calculation result is, the higher the degree of similarity between the feature vector of the first conversation text and the feature vector of the description text of the corresponding plug-in is, and the corresponding matching result is also represented as a better match.

[0070] For another example, when registering a plug-in, keyword extraction can be performed on the plug-in's description text to obtain at least one keyword from the description text, and plug-in index information can be generated based on the at least one keyword. That is, the plug-in index information can include at least one keyword related to the plug-in's description text. Matching the first conversation text with the respective index information of multiple plug-ins to obtain multiple matching results can include the following operations:

[0071] Keyword extraction is performed on the first dialogue text to obtain keywords related to the first dialogue text; and the keywords related to the first dialogue text are matched with at least one keyword related to the description texts of the multiple plug-ins to obtain multiple matching results.

[0072] According to an embodiment of the present disclosure, the similarity between the plugin's index information and the first conversation text can be determined based on the number of keyword hits, that is, the matching result can be expressed as the number of keywords hit. For example, keyword extraction of the description text of plugin α can obtain keywords a, b, c, and d. Similarly, keyword extraction of the first conversation text can obtain keywords b, d, and e. Since plugin α and the first conversation text both include keywords b and d, the matching result obtained by matching the first conversation text with plugin α can be expressed as 2.

[0073] According to an embodiment of the present disclosure, by setting index information for a plug-in during plug-in registration, plug-in matching can be achieved quickly based on the index information, thereby effectively improving plug-in matching accuracy and improving the efficiency of plug-in matching operations.

[0074] According to an embodiment of the present disclosure, each plug-in can have its own problem type that can be handled. For example, plug-in 1 can be used to solve mathematical calculation problems, plug-in 2 can be used to solve weather query problems, etc. Based on this, multiple registered plug-ins can be grouped and processed based on the problems they can solve. For example, the weather query plug-in and the humidity query plug-in are both used to solve the problem of how to query the climate. Therefore, the weather query plug-in and the humidity query plug-in can be classified into the same category and the plug-in type of the weather query plug-in and the humidity query plug-in can be determined to be climate plug-ins.

[0075] According to an embodiment of the present disclosure, as an optional implementation, the amount of index information contained in the plug-in information library can be large. At this time, plug-in matching based on index information may still take a long time. Therefore, before plug-in matching based on index information, plug-ins can be initially screened based on their types.

[0076] FIG3B schematically shows a schematic diagram of a matching process of a first target plug-in according to another embodiment of the present disclosure.

[0077] As shown in FIG3B , N registered plug-ins 301 may be configured in the large language model, and N index information 303 corresponding to the N plug-ins 301 may be recorded in the plug-in information library 302 .

[0078] According to an embodiment of the present disclosure, plug-in type information 307 related to the first conversation text 304 can be determined based on the first conversation text 304. Based on the plug-in type information 307, at least one second target plug-in 308 can be determined from the N plug-ins 301. That is, the plug-in type of each of the at least one second target plug-in 308 is consistent with the plug-in type indicated by the plug-in type information 307. After determining the at least one second target plug-in 308, the index information 303 of each of the at least one second target plug-in 308 can be obtained from the plug-in information library 302. The first conversation text 304 is matched with the index information 303 of each of the at least one second target plug-in 308 to obtain at least one matching result 305. Based on the at least one matching result 305, the first target plug-in 306 can be determined from the at least one second target plug-in 308.

[0079] According to an embodiment of the present disclosure, the process of matching at least one second target plug-in and the process of determining the first target plug-in from at least one second target plug-in can be implemented using the method of matching multiple plug-ins and the method of determining the first target plug-in from multiple plug-ins as described above, which will not be repeated here.

[0080] According to an embodiment of the present disclosure, by performing preliminary screening of plug-ins based on plug-in types before performing plug-in matching based on index information, the computing resources consumed by the plug-in matching process can be effectively reduced and processing efficiency can be improved.

[0081] According to an embodiment of the present disclosure, after the first target plug-in is determined, the text actually input into the large language model, ie, the second dialog text, can be obtained based on the first dialog text and the description text of the first target plug-in.

[0082] According to an embodiment of the present disclosure, a second dialogue text can be generated based on the first dialogue text and the description text of the first target plug-in by direct splicing, that is, the description text of the first target plug-in can be spliced ​​in the context of the first dialogue text to obtain the second dialogue text.

[0083] According to an embodiment of the present disclosure, for example, the first dialogue text may be "Please help me calculate what 256 times 4 equals", and the first target plug-in matched based on the first dialogue text may be a mathematical calculation plug-in, and the description text of the mathematical calculation plug-in may be "Mathematical calculation plug-in, which can input an expression consisting of numbers and operators and output the calculation result". The description text of the first target plug-in is spliced ​​into the context of the first dialogue text, and the obtained second dialogue text can be expressed as "Mathematical calculation plug-in, which can input an expression consisting of numbers and operators and output the calculation result. Please help me calculate what 256 times 4 equals".

[0084] According to an embodiment of the present disclosure, a template-based splicing method can also be used to generate a second dialogue text based on the first dialogue text and the description text of the first target plug-in, that is, the first dialogue text and the description text of the first target plug-in can be filled into the first prompt template respectively to obtain the second dialogue text.

[0085] According to an embodiment of the present disclosure, the first prompt template can be a text template set by the user, having one or more replaceable text paragraphs, and the language expression is suitable for input into the large language model. The replaceable text paragraph can be represented as an information slot in the first prompt template. For example, the first prompt template can include two information slots, namely a first text information slot and a plug-in information slot, the first text information slot is suitable for filling in the dialogue text input by the user, and the plug-in information slot is suitable for filling in the description text of the plug-in. The second dialogue text generated based on the first prompt template can be used to guide the large language model to call the plug-in and process the text.

[0086] According to an embodiment of the present disclosure, for example, the first prompt template can be expressed as "You can answer the question by using the following plug-in: [insert text 1]. Given the following question: [insert text 2]". In the first prompt template, "[insert text 1]" represents the plug-in information slot, and "[insert text 2]" represents the first text information slot. The first dialogue text can be "Please help me calculate what 256 times 4 equals", and the first target plug-in matched based on the first dialogue text can be a mathematical calculation plug-in, and the description text of the mathematical calculation plug-in can be "Mathematical calculation plug-in, which can input an expression consisting of numbers and operators and output the calculation result". The first dialogue text can be filled into the first text information slot, and the description text of the first target plug-in can be filled into the plug-in information slot to obtain the second dialogue text. The obtained second dialogue text can be expressed as "You can answer the question by using the following plug-in: Mathematical calculation plug-in, which can input an expression consisting of numbers and operators and output the calculation result. Given the following question: Please help me calculate what 256 times 4 equals".

[0087] According to an embodiment of the present disclosure, after the second dialogue text is generated, the large language model can be used to call the first target plug-in to process the second dialogue text to obtain a reply text.

[0088] According to embodiments of the present disclosure, the output text generated by the first target plug-in after processing the second conversation text can be text with practical or specific meaning. In this case, this output text can be directly used as the reply text of the large language model. Specifically, the second conversation text can be input into the large language model, which then uses the large language model to call the first target plug-in based on the description text of the first target plug-in included in the second conversation text, and then process the first conversation text included in the second conversation text to generate the reply text.

[0089] FIG4A schematically shows a schematic diagram of a reply text generation process according to an embodiment of the present disclosure.

[0090] As shown in Figure 4A, second conversation text 401 can be input into large language model 402. Based on description text 4011 included in second conversation text 401, the large language model can call first target plug-in 403 to process first conversation text 4012 included in second conversation text 401. The processing result obtained by first target plug-in 403 is reply text 404 of large language model 402.

[0091] For example, the second dialogue text may be expressed as "You can answer questions by using the following plug-in: Math Calculation Plug-in, which accepts input of expressions consisting of numbers and operators and outputs calculation results. Given the following question: Please help me calculate what 256 times 4 equals." After the second dialogue text is input into the large language model, the large language model may determine that the plug-in to be called is the math calculation plug-in based on the description text "Math Calculation Plug-in, which accepts input of expressions consisting of numbers and operators and outputs calculation results" contained in the second dialogue text. The math calculation plug-in can then be used to process the first dialogue text "Please help me calculate what 256 times 4 equals" contained in the second dialogue text. After processing, the output text "1024" can be obtained, which can be directly used as the reply text of the large language model.

[0092] According to an embodiment of the present disclosure, as an optional implementation, the processing results output by the first target plug-in can be re-input into the large language model to leverage the large language model's text processing capabilities and output a response text that is closer to the intended expression. For example, the second conversation text can be input into the large language model, which then uses the large language model to invoke the first target plug-in based on the description text of the first target plug-in included in the second conversation text, process the first conversation text included in the second conversation text, and obtain an initial response text; based on the first conversation text and the initial response text, obtain a third conversation text; and then input the third conversation text into the large language model to obtain a response text.

[0093] FIG4B schematically shows a schematic diagram of a reply text generation process according to another embodiment of the present disclosure.

[0094] As shown in Figure 4B, second conversation text 401 can be input into large language model 402. Based on description text 4011 included in second conversation text 401, the large language model can call first target plug-in 403 to process first conversation text 4012 included in second conversation text 401, thereby obtaining initial reply text 405. Initial reply text 405 can be merged with first conversation text 4012 to obtain third conversation text 406. Specifically, first conversation text 4012 and initial reply text 405 can be respectively entered into a second prompt template to obtain third conversation text 406. Third conversation text 406 can then be input into large language model 402 to obtain reply text 404.

[0095] According to an embodiment of the present disclosure, similar to the first prompt template, the second prompt template may also include multiple information slots. For example, the second prompt template may include two information slots, which may be represented as a second text information slot and a third text information slot, respectively. It should be noted that the second prompt template may also include more than two information slots, which is not limited here.

[0096] According to an embodiment of the present disclosure, the first dialogue text 4012 and the initial reply text 405 are respectively filled into the second prompt template to obtain the third dialogue text 406. Specifically, the first dialogue text 4012 can be filled into the second text information slot, and the initial reply text 405 can be filled into the third text information slot to obtain the third dialogue text 406.

[0097] For example, the second prompt template can be expressed as "You can answer the question based on the following information: [insert text 3]. Given the following question: [insert text 4]". Among them, the information slot "[insert text 3]" can represent the third text information slot, and the information slot "[insert text 4]" can represent the second text information slot. The first dialogue text can be expressed as "Please help me calculate what 256 times 4 equals", and the initial reply text can be expressed as "1024". After filling the first dialogue text and the initial reply text into the second prompt template, the third dialogue text obtained can be expressed as "You can answer the question based on the following information: 1024. Given the following question: Please help me calculate what 256 times 4 equals". After inputting the third dialogue text into the large language model, the obtained reply text can be expressed as "The result of 256 times 4 is 1024".

[0098] According to the embodiments of the present disclosure, with the help of the understanding ability of the large language model, the corresponding plug-in can be called based on the description text of the plug-in to handle tasks that the large language model could not handle before. This can effectively improve the usability and universality of the large language model, and there is no need to retrain the large language model when facing different tasks, thereby reducing the cost of using the large language model.

[0099] According to an embodiment of the present disclosure, as an optional implementation, the user can manually select a plug-in to be used by the large language model for processing conversation text. Specifically, the user can specify that the description text of certain plug-ins should be added to the context of the conversation text. The text actually input into the large language model can include the first conversation text, the description text of the plug-in selected by the conversation system, and the description text of the plug-in manually selected by the user.

[0100] According to an embodiment of the present disclosure, the operation of manual selection by the user can be used to change the selection status mark of the plug-in. For example, the user can manually select at least one fourth target plug-in. Accordingly, multiple plug-ins can include at least one fourth target plug-in, and at least one fourth target plug-in can be marked as selected.

[0101] FIG5 schematically shows a flow chart of a human-computer interaction method according to another embodiment of the present disclosure.

[0102] As shown in FIG. 5 , the method includes operations S510 to S530 .

[0103] In operation S510 , in response to a human-computer interaction request, based on a first dialogue text included in the human-computer interaction request, a first target plug-in related to the first dialogue text is determined from a plurality of plug-ins registered in a large language model.

[0104] In operation S520 , a fourth dialogue text is obtained based on the first dialogue text, the description text of the first target plug-in, and the description text of each of at least one fourth target plug-in.

[0105] In operation S530, the fourth conversation text is input into the large language model to obtain a reply text.

[0106] According to an embodiment of the present disclosure, the method for determining the first target plug-in from multiple plug-ins can refer to the matching process of the first target plug-in described above, and will not be repeated here.

[0107] According to an embodiment of the present disclosure, a method for obtaining a fourth dialogue text based on the first dialogue text, the description text of the first target plug-in and the description text of at least one fourth target plug-in can refer to the method for generating the second dialogue text as described above, replacing the description text of the first target plug-in with the description text of the first target plug-in and the description text of at least one fourth target plug-in, and then replacing the second dialogue text with the fourth dialogue text. No further details will be given here.

[0108] According to an embodiment of the present disclosure, during the reply text generation process, at least one matching result obtained by matching the first conversation text with at least one fourth target plug-in can be further determined. Based on this at least one matching result, it can be determined whether the plug-in that actually processes the first conversation text is the first target plug-in autonomously selected by the conversation system or the plug-in selected by the user. Specifically, the fourth conversation text can be input into a large language model, and the large language model can be used to determine a fourth target plug-in from the first target plug-in and at least one fourth target plug-in based on the fourth conversation text. Furthermore, the large language model can be used to invoke the fourth target plug-in based on the description text of the fourth target plug-in included in the fourth conversation text, and process the first conversation text included in the fourth conversation text to generate the reply text. The process of generating the reply text using the fourth conversation text can refer to the reply text generation process described above and will not be further described here. The method used when matching the first conversation text with at least one fourth target plug-in can differ from the method used when matching the first conversation text with multiple plug-ins. Furthermore, a correction factor can be added to the matching result of the fourth target plug-in. This correction factor can be a value greater than 1, so that the large language model can maximize the use of the user-selected plug-in for text processing.

[0109] According to an embodiment of the present disclosure, by adding the description text of the plug-in automatically selected by the dialogue system and the description text of the plug-in selected by the user to the context of the dialogue text, the user's usage experience can be improved while ensuring the tendency of the large language model when processing text, so that the reply text output by the large language model is more in line with the user's original intention, thereby improving the usability of the dialogue system.

[0110] FIG6 schematically shows a block diagram of a human-computer interaction device according to an embodiment of the present disclosure.

[0111] As shown in FIG. 6 , the human-computer interaction device 600 includes a determination module 610 , a first processing module 620 and a first input module 630 .

[0112] A determination module 610 is configured to determine, in response to a human-computer interaction request and based on a first dialogue text included in the human-computer interaction request, a first target plug-in associated with the first dialogue text from a plurality of plug-ins registered in the large language model;

[0113] A first processing module 620 is configured to obtain a second dialogue text based on the first dialogue text and the description text of the first target plug-in; and

[0114] The first input module 630 is used to input the second dialogue text into the large language model to obtain a reply text.

[0115] According to an embodiment of the present disclosure, the determination module 610 includes a first determination unit, a second determination unit, and a third determination unit.

[0116] The first determining unit is configured to obtain index information of each of the plurality of plug-ins from the plug-in information library.

[0117] The second determining unit is configured to match the first conversation text with the respective index information of the plurality of plug-ins to obtain a plurality of matching results.

[0118] The third determining unit is configured to determine a first target plug-in from the plurality of plug-ins based on the plurality of matching results.

[0119] According to an embodiment of the present disclosure, the index information of the plug-in includes a feature vector of the description text of the plug-in.

[0120] According to an embodiment of the present disclosure, the second determining unit includes a first determining subunit, a second determining subunit, and a third determining subunit.

[0121] The first determining subunit is configured to extract features from the first dialogue text to obtain a feature vector of the first dialogue text.

[0122] The second determining subunit is configured to perform similarity calculations on the feature vector of the first dialogue text and the feature vectors of the description texts of the plurality of plug-ins to obtain a plurality of similarity calculation results.

[0123] The third determining subunit is configured to obtain multiple matching results based on multiple similarity calculation results.

[0124] According to an embodiment of the present disclosure, the index information of the plug-in includes at least one keyword related to the description text of the plug-in.

[0125] According to an embodiment of the present disclosure, the second determining unit includes a fourth determining subunit and a fifth determining subunit.

[0126] The fourth determining subunit is configured to extract keywords from the first dialogue text to obtain keywords related to the first dialogue text.

[0127] The fifth determining subunit is configured to match the keyword related to the first conversation text with at least one keyword related to the description texts of the plurality of plug-ins to obtain a plurality of matching results.

[0128] According to an embodiment of the present disclosure, the determination module 610 includes a fourth determination unit, a fifth determination unit, a sixth determination unit, a seventh determination unit, and an eighth determination unit.

[0129] The fourth determining unit is configured to determine plug-in type information related to the first dialogue text based on the first dialogue text.

[0130] The fifth determining unit is configured to determine at least one second target plug-in from the plurality of plug-ins based on the plug-in type information.

[0131] The sixth determining unit is configured to obtain index information of at least one second target plug-in from the plug-in information library.

[0132] The seventh determining unit is configured to match the first conversation text with the respective index information of at least one second target plug-in to obtain at least one matching result.

[0133] An eighth determining unit is configured to determine a first target plug-in from at least one second target plug-in based on the at least one matching result.

[0134] According to an embodiment of the present disclosure, the human-computer interaction device 600 further includes a generating module and a writing module.

[0135] The generating module is configured to generate index information of the third target plug-in in response to the plug-in registration request and based on the description text of the third target plug-in included in the plug-in registration request.

[0136] The writing module is used to write the index information of the third target plug-in into the plug-in information library.

[0137] According to an embodiment of the present disclosure, the first processing module 620 includes a first processing unit.

[0138] The first processing unit is configured to splice the description text of the first target plug-in into the context of the first dialogue text to obtain a second dialogue text.

[0139] According to an embodiment of the present disclosure, the first processing module 620 includes a second processing unit.

[0140] The second processing unit is configured to fill the first dialogue text and the description text of the first target plug-in into the first prompt template respectively to obtain a second dialogue text.

[0141] According to an embodiment of the present disclosure, the first prompt template includes a first text information slot and a plug-in information slot.

[0142] According to an embodiment of the present disclosure, the second processing unit includes a processing sub-unit.

[0143] The processing subunit is used to fill the first dialogue text into the first text information slot and fill the description text of the first target plug-in into the plug-in information slot to obtain the second dialogue text.

[0144] According to an embodiment of the present disclosure, the first input module 630 includes a first input unit.

[0145] The first input unit is used to input the second dialogue text into the large language model, so as to use the large language model to call the first target plug-in based on the description text of the first target plug-in included in the second dialogue text, and process the first dialogue text included in the second dialogue text to obtain a reply text.

[0146] According to an embodiment of the present disclosure, the first input module 630 includes a second input unit, a third input unit, and a fourth input unit.

[0147] The second input unit is used to input the second dialogue text into the large language model, so as to use the large language model to call the first target plug-in based on the description text of the first target plug-in included in the second dialogue text, process the first dialogue text included in the second dialogue text, and obtain an initial reply text.

[0148] The third input unit is used to fill the first dialogue text and the initial reply text into the second prompt template respectively to obtain a third dialogue text.

[0149] The fourth input unit is used to input the third dialogue text into the large language model to obtain a reply text.

[0150] According to an embodiment of the present disclosure, the second prompt template includes a second text information slot and a third text information slot.

[0151] According to an embodiment of the present disclosure, the third input unit includes an input sub-unit.

[0152] The input subunit is used to fill the first dialogue text into the second text information slot and fill the initial reply text into the third text information slot to obtain the third dialogue text.

[0153] According to an embodiment of the present disclosure, the plurality of plug-ins include at least one fourth target plug-in, wherein the at least one fourth target plug-in is marked as a selected state.

[0154] According to an embodiment of the present disclosure, the human-computer interaction device 600 further includes a second processing module and a second input module.

[0155] The second processing module is configured to obtain a fourth dialogue text based on the first dialogue text, the description text of the first target plug-in, and the description text of at least one fourth target plug-in.

[0156] The second input module is used to input the fourth dialogue text into the large language model to obtain a reply text.

[0157] According to an embodiment of the present disclosure, the second input module includes a fifth input unit and a sixth input unit.

[0158] The fifth input unit is configured to input the fourth dialogue text into the large language model, so as to determine a fourth target plug-in from the first target plug-in and the at least one fourth target plug-in based on the fourth dialogue text using the large language model.

[0159] The sixth input unit is configured to use the large language model to call the fourth target plug-in based on the description text of the fourth target plug-in included in the fourth dialogue text, and process the first dialogue text included in the fourth dialogue text to obtain a reply text.

[0160] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0161] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.

[0162] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method described above.

[0163] According to an embodiment of the present disclosure, a computer program product includes a computer program, and when the computer program is executed by a processor, the computer program implements the method described above.

[0164] FIG7 shows a schematic block diagram of an example electronic device that can be used to implement an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0165] As shown in Figure 7, device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of device 700 can also be stored in RAM 703. Computing unit 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to bus 704.

[0166] Various components in device 700 are connected to an input / output (I / O) interface 705, including an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0167] The computing unit 701 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above, such as the human-computer interaction method. For example, in some embodiments, the human-computer interaction method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the human-computer interaction method described above can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to perform the human-computer interaction method by any other appropriate means (e.g., by means of firmware).

[0168] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0169] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0170] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0171] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0172] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0173] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0174] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0175] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A human-computer interaction method, comprising: In response to a human-computer interaction request, based on a first dialogue text included in the human-computer interaction request, determining a first target plug-in related to the first dialogue text from a plurality of plug-ins registered in a large language model; Obtaining a second dialogue text based on the first dialogue text and the description text of the first target plug-in; as well as The second dialogue text is input into the large language model to obtain a reply text.

2. The method according to claim 1, wherein: The determining, based on the first dialogue text included in the human-computer interaction request, a first target plug-in related to the first dialogue text from a plurality of plug-ins registered in the large language model comprises: Obtaining index information of each of the plurality of plug-ins from a plug-in information repository; Matching the first conversation text with the respective index information of the multiple plug-ins to obtain multiple matching results; and Based on the multiple matching results, the first target plug-in is determined from the multiple plug-ins.

3. The method according to claim 2, wherein: The index information of the plug-in includes a feature vector of a description text of the plug-in; The first conversation text is matched with the respective index information of the plurality of plug-ins to obtain a plurality of matching results, including: Performing feature extraction on the first dialogue text to obtain a feature vector of the first dialogue text; Calculating similarities between the feature vector of the first conversation text and the feature vectors of the description texts of the plurality of plug-ins, respectively, to obtain a plurality of similarity calculation results; and Based on the multiple similarity calculation results, the multiple matching results are obtained.

4. The method according to claim 2, wherein: The index information of the plug-in includes at least one keyword related to the description text of the plug-in; The first conversation text is matched with the respective index information of the plurality of plug-ins to obtain a plurality of matching results, including: Extracting keywords from the first conversation text to obtain keywords related to the first conversation text; and The keywords related to the first conversation text are matched respectively with at least one keyword related to the description texts of the multiple plug-ins to obtain the multiple matching results.

5. The method according to claim 1, wherein: The determining, based on the first dialogue text included in the human-computer interaction request, a first target plug-in related to the first dialogue text from a plurality of plug-ins registered in the large language model comprises: Based on the first dialogue text, determining plug-in type information related to the first dialogue text; Based on the plug-in type information, determining at least one second target plug-in from the plurality of plug-ins; Acquire the index information of each of the at least one second target plug-in from the plug-in information library; Matching the first conversation text with the respective index information of the at least one second target plug-in to obtain at least one matching result; and Based on the at least one matching result, the first target plug-in is determined from the at least one second target plug-in.

6. The method according to any one of claims 2 to 5, further comprising: In response to the plug-in registration request, generating index information of the third target plug-in based on the description text of the third target plug-in included in the plug-in registration request; as well as The index information of the third target plug-in is written into the plug-in information library.

7. The method according to claim 1, wherein: The obtaining of the second dialogue text based on the first dialogue text and the description text of the first target plug-in includes: The description text of the first target plug-in is spliced ​​into the context of the first dialogue text to obtain the second dialogue text.

8. The method according to claim 1, wherein: The obtaining of the second dialogue text based on the first dialogue text and the description text of the first target plug-in includes: The first dialogue text and the description text of the first target plug-in are respectively filled into the first prompt template to obtain the second dialogue text.

9. The method according to claim 8, wherein: The first prompt template includes a first text information slot and a plug-in information slot; The step of filling the first dialogue text and the description text of the first target plug-in into the first prompt template to obtain the second dialogue text includes: The first dialogue text is filled into the first text information slot, and the description text of the first target plug-in is filled into the plug-in information slot to obtain the second dialogue text.

10. The method according to claim 1, wherein: The step of inputting the second dialogue text into the large language model to obtain a reply text includes: The second dialogue text is input into the large language model, so as to use the large language model to call the first target plug-in based on the description text of the first target plug-in included in the second dialogue text, and the first dialogue text included in the second dialogue text is processed to obtain the reply text.

11. The method according to claim 1, wherein: The step of inputting the second dialogue text into the large language model to obtain a reply text includes: Inputting the second dialogue text into the large language model, using the large language model to call the first target plug-in based on the description text of the first target plug-in included in the second dialogue text, and processing the first dialogue text included in the second dialogue text to obtain an initial reply text; Filling the first dialogue text and the initial reply text into the second prompt template respectively to obtain a third dialogue text; and The third dialogue text is input into the large language model to obtain the reply text.

12. The method according to claim 11, wherein: The second prompt template includes a second text information slot and a third text information slot; The step of filling the first dialogue text and the initial reply text into the second prompt template to obtain a third dialogue text includes: The first dialogue text is filled into the second text information slot, and the initial reply text is filled into the third text information slot to obtain the third dialogue text.

13. The method according to claim 1, wherein: The plurality of plug-ins include at least one fourth target plug-in, wherein the at least one fourth target plug-in is marked as a selected state; The method further comprises: Obtaining a fourth dialogue text based on the first dialogue text, the description text of the first target plug-in, and the description text of each of the at least one fourth target plug-in; and The fourth dialogue text is input into the large language model to obtain the reply text.

14. The method according to claim 13, wherein: The step of inputting the fourth dialogue text into the large language model to obtain the reply text comprises: Inputting the fourth dialogue text into the large language model to determine a fourth target plug-in from the first target plug-in and the at least one fourth target plug-in based on the fourth dialogue text by using the large language model; and Using the large language model to call the fourth target plug-in based on the description text of the fourth target plug-in included in the fourth dialogue text, and processing the first dialogue text included in the fourth dialogue text, The reply text is obtained.

15. A human-computer interaction device, comprising: A determination module, configured to determine, in response to a human-computer interaction request, a first target plug-in related to the first dialogue text from a plurality of plug-ins registered in a large language model based on a first dialogue text included in the human-computer interaction request; A first processing module, configured to obtain a second dialogue text based on the first dialogue text and a description text of the first target plug-in; as well as The first input module is used to input the second dialogue text into the large language model to obtain a reply text.

16. The device according to claim 15, wherein: The determining module comprises: A first determining unit, configured to obtain index information of each of the plurality of plug-ins from a plug-in information library; A second determining unit, configured to match the first conversation text with the respective index information of the plurality of plug-ins to obtain a plurality of matching results; and A third determining unit is configured to determine the first target plug-in from the multiple plug-ins based on the multiple matching results.

17. The device according to claim 16, wherein: The index information of the plug-in includes a feature vector of a description text of the plug-in; Wherein, the second determining unit includes: A first determining subunit is used to extract features from the first dialogue text to obtain a feature vector of the first dialogue text; a second determining subunit, configured to perform similarity calculations on the feature vector of the first conversation text and the feature vectors of the description texts of the plurality of plug-ins to obtain a plurality of similarity calculation results; and The third determining subunit is configured to obtain the plurality of matching results based on the plurality of similarity calculation results.

18. The device according to claim 16, wherein: The index information of the plug-in includes at least one keyword related to the description text of the plug-in; Wherein, the second determining unit includes: a fourth determining subunit, configured to extract keywords from the first dialogue text to obtain keywords related to the first dialogue text; and The fifth determining subunit is used to match the keywords related to the first dialogue text with at least one keyword related to the description texts of the multiple plug-ins respectively, to obtain the multiple matching results.

19. The device according to claim 15, wherein: The determining module comprises: The fourth determining unit is used to determine, based on the first dialogue text, an insert text related to the first dialogue text. Document type information; a fifth determining unit, configured to determine at least one second target plug-in from the plurality of plug-ins based on the plug-in type information; A sixth determining unit, configured to obtain index information of each of the at least one second target plug-in from a plug-in information repository; a seventh determining unit, configured to match the first conversation text with the respective index information of the at least one second target plug-in to obtain at least one matching result; and An eighth determining unit is configured to determine the first target plug-in from the at least one second target plug-in based on the at least one matching result.

20. The device according to any one of claims 16 to 19, further comprising: A generating module, configured to generate, in response to a plug-in registration request, index information of the third target plug-in based on a description text of the third target plug-in included in the plug-in registration request; as well as A writing module is used to write the index information of the third target plug-in into the plug-in information library.

21. The device according to claim 15, wherein: The first processing module comprises: The first processing unit is used to splice the description text of the first target plug-in into the context of the first dialogue text to obtain the second dialogue text.

22. The device according to claim 15, wherein: The first processing module comprises: The second processing unit is used to fill the first dialogue text and the description text of the first target plug-in into the first prompt template respectively to obtain the second dialogue text.

23. The device according to claim 22, wherein: The first prompt template includes a first text information slot and a plug-in information slot; Wherein, the second processing unit includes: The processing subunit is used to fill the first dialogue text into the first text information slot, and fill the description text of the first target plug-in into the plug-in information slot to obtain the second dialogue text.

24. The device according to claim 15, wherein: The first input module comprises: The first input unit is used to input the second dialogue text into the large language model, so as to use the large language model to call the first target plug-in based on the description text of the first target plug-in included in the second dialogue text, and process the first dialogue text included in the second dialogue text to obtain the reply text.

25. The device according to claim 15, wherein: The first input module comprises: The second input unit is used to input the second dialogue text into the large language model to use the large language model The language model calls the first target plug-in based on the description text of the first target plug-in included in the second dialogue text, processes the first dialogue text included in the second dialogue text, and obtains an initial reply text; A third input unit, used to fill the first dialogue text and the initial reply text into the second prompt template respectively to obtain a third dialogue text; and The fourth input unit is used to input the third dialogue text into the large language model to obtain the reply text.

26. The device according to claim 25, wherein The second prompt template includes a second text information slot and a third text information slot; Wherein, the third input unit includes: The input subunit is used to fill the first dialogue text into the second text information slot, and fill the initial reply text into the third text information slot to obtain the third dialogue text.

27. The device according to claim 15, wherein: The plurality of plug-ins include at least one fourth target plug-in, wherein the at least one fourth target plug-in is marked as a selected state; The device also includes: A second processing module, configured to obtain a fourth dialogue text based on the first dialogue text, the description text of the first target plug-in, and the description text of the at least one fourth target plug-in; and The second input module is used to input the fourth dialogue text into the large language model to obtain the reply text.

28. The device according to claim 27, wherein The second input module comprises: a fifth input unit, configured to input the fourth dialogue text into the large language model, so as to determine a fourth target plug-in from the first target plug-in and the at least one fourth target plug-in based on the fourth dialogue text by using the large language model; and The sixth input unit is used to call the fourth target plug-in based on the description text of the fourth target plug-in included in the fourth dialogue text by using the large language model, and process the first dialogue text included in the fourth dialogue text to obtain the reply text.

29. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any one of claims 1 to 14. Law.

30. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 14.

31. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Language model training method, device, electronic equipment and readable storage medium

    CN111859982A

  • Human-computer interaction method, device and system

    CN116483980A

  • Control method and control device of intelligent dialogue system and electronic equipment

    CN116521893A

  • Human-computer interaction method and device, electronic equipment and storage medium

    CN117332068A

  • Processing natural language text with context-specific linguistic model

    US20160350280A1