Voice interaction method and electronic equipment
By using the large language model LLM to determine user intentions and filter plug-ins, the problem of cumbersome and inefficient plug-ins selection in the existing technology is solved, and more efficient and accurate user interaction services are achieved.
Patent Information
- Application Number
- CN202311550888.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-11-17
AI Technical Summary
The prior art has shortcomings in improving the efficiency of human-computer interaction with users in large language models, especially in terms of cumbersome operations and inefficient in plug-in selection and memory resource utilization.
By receiving text input information, the user's intent text is determined using the large language model LLM, and the intent vector is calculated. Based on the intent vector, the matching plug-in description vector is filtered out from the vector database, and then the corresponding plug-in metadata information is obtained from the text database. Then, the target plug-in is called through LLM to generate target text information that matches the user's intent.
It realizes more accurate and efficient plug-in selection, reduces the cumbersomeness of user manual operations, improves the system's interaction efficiency, and optimizes the utilization of memory resources.
Smart Images

Figure CN120067284A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminals, and in particular, to a voice interaction method and an electronic device. Background Art
[0002] Artificial intelligence (AI) is a branch of computer science, which is a technology that studies the design principles and implementation methods of various intelligent machines, enabling electronic devices to have the capabilities of perception, reasoning, and decision-making. Since the release of ChatGPT developed based on AI technology, the capabilities and future potential of large foundation models (e.g., large language models (LLMs)) have received extensive attention from all walks of life. The question-and-answer system based on the large language model can interact with users and help users handle daily affairs, such as answering users' questions, querying hotel rooms according to users' needs, and so on. Therefore, how to improve the efficiency of the human-computer interaction between the large language model and users has become an urgent problem to be solved currently. Summary of the Invention
[0003] This application provides a voice interaction method and an electronic device, which realizes accurately and efficiently selecting the plug-ins corresponding to the user intention, so that the electronic device can provide corresponding services more efficiently and accurately based on the user intention.
[0004] In a first aspect, this application provides a voice interaction method, including: receiving text input information; determining user intention text from the text input information through a large language model LLM; calculating an intention vector corresponding to the user intention text; where the intention vector is used to represent the user intention; based on the intention vector, determining M plug-in description vectors matching the intention vector from a vector database; where N plug-in description vectors are included in the vector database, and M is less than N; based on the M plug-in description vectors, determining M plug-in metadata information corresponding to the M plug-in description vectors from a text database; where N plug-in metadata information is included in the text database; through the LLM, determining K target plug-ins corresponding to the user intention based on the user intention text and the M plug-in metadata information; where K is less than M, and K, M, and N are positive integers; calling the K target plug-ins to generate target text information matching the user intention.
[0005] In a possible implementation manner, the determining M plug-in description vectors matching the intention vector from the vector database based on the intention vector includes: determining M plug-in description vectors whose similarity to the intention vector is greater than or equal to a first threshold.
[0006] In a possible implementation, determining the M plugin metadata information corresponding to the M plugin description vectors from the text database includes: determining the M plugin metadata information corresponding to the M plugin description vectors based on a first mapping relationship; wherein, the first mapping relationship includes the mapping relationship between each plugin description vector and each plugin metadata information.
[0007] In a possible implementation, determining the K target plugins corresponding to the user intention by the LLM based on the user intention text and the M plugin metadata information includes: filling the user intention text and the plugin description information in the M plugin metadata information into a first prompt template to obtain a first prompt. Determining the K target plugins corresponding to the user intention by the LLM based on the first prompt.
[0008] In a possible implementation, calling the K target plugins to generate target text information matching the user intention includes: filling the text input information and the request parameters in the M plugin metadata information into a second prompt template to obtain a second prompt; extracting input parameters from the text input information by the LLM based on the second prompt; calling the K target plugins based on the input parameters to generate target text information matching the user intention.
[0009] In a possible implementation, determining the user intention text from the text input information by the large language model LLM includes: filling the text input information into a third prompt template to obtain a third prompt; determining the user intention text from the text input information by the LLM based on the third prompt.
[0010] In a possible implementation, the LLM is ChatGPT-3, ChatGPT-4, BERT or XLNet.
[0011] In a possible implementation, before receiving the text input information, the method further includes: receiving the input voice information and converting the voice information into the text input information; after receiving the text input information, the method further includes: displaying the text input information; after calling the K target plugins to generate target text information matching the user intention, the method further includes: displaying the target text information.
[0012] In a possible implementation, the plug-in metadata information includes one or more of the following: plug-in description information, the uniform resource locator (URL) of the plug-in, the request parameters of the plug-in, and the response parameters of the plug-in.
[0013] In a possible implementation, the K target plug-ins include one or more of the following: weather query plug-in, map plug-in, train and high-speed rail ticket booking plug-in, movie ticket booking plug-in, and hotel query plug-in.
[0014] In a second aspect, the present application provides an electronic device, including: one or more processors and one or more memories; the one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer-executable programs. When the one or more processors execute the computer-executable programs, the electronic device is caused to execute the method in any one of the possible implementations in the first aspect described above.
[0015] In a third aspect, the present application provides a chip system, including a processing circuit and an interface circuit. The interface circuit is used to receive code instructions and transmit them to the processing circuit, and the processing circuit is used to run the code instructions so that the chip system executes the method in any one of the possible implementations in the first aspect described above.
[0016] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer-executable program. When the computer-executable program runs on an electronic device, the electronic device is caused to execute the method in any one of the possible implementations in the first aspect described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figures 1A - 1D A set of schematic diagrams of user interfaces provided by an embodiment of the present application;
[0018] Figure 2 A schematic diagram of a voice interaction process provided by an embodiment of the present application;
[0019] Figure 3A A software architecture applied to an electronic device provided by an embodiment of the present application;
[0020] Figure 3B A schematic diagram of data interaction provided by an embodiment of the present application;
[0021] Figure 4 A specific process schematic diagram of a voice interaction method provided by an embodiment of the present application;
[0022] Figure 5A A schematic diagram of a process for user intention recognition provided by an embodiment of the present application;
[0023] Figure 5B Schematic diagram of an example for user intention recognition provided by an embodiment of the present application;
[0024] Figure 5C Schematic diagram of an example for plug-in query provided by an embodiment of the present application;
[0025] Figure 5D Schematic diagram of an example for plug-in query provided by an embodiment of the present application;
[0026] Figure 5E Schematic diagram of another example for plug-in query provided by an embodiment of the present application;
[0027] Figure 5F Schematic diagram of a method for plug-in invocation provided by an embodiment of the present application;
[0028] Figure 5G Schematic diagram of a method for plug-in registration provided by an embodiment of the present application;
[0029] Figure 6 Schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0030] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. As used in the specification and the appended claims of the present application, the singular forms "a", "an", "the", "above-mentioned", "said", and "this" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term " / and / " used in the present application refers to any or all possible combinations including one or more of the listed items. In the embodiments of the present application, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.
[0031] To better describe the technical solutions provided by the embodiments of the present application, some terms related to the embodiments of the present application are first explained:
[0032] 1. Large Language Model (LLM): Refers to a type of neural network model based on deep learning, used for learning and generating natural language texts, and can be applied to tasks such as language understanding, text classification, machine translation, and dialogue generation. The large language model in the embodiments of this application can be ChatGPT-3, ChatGPT-4, BERT, XLNet, etc. The above large language models can have hundreds of millions or billions of parameters and are capable of handling longer, more complex, and more abstract natural language tasks. Due to their huge model scale and strong generalization ability, LLMs have important application values in the field of natural language processing.
[0033] 2. Voice Assistant: Refers to an intelligent voice interaction system based on artificial intelligence technology, which can interact with users through voice to execute user instructions and help users complete various daily tasks, such as querying the weather, playing music, sending text messages, setting alarms, etc. The voice assistant can be built into electronic devices such as smartphones, smart speakers, and smart watches. Its working principle is to recognize the voice information input by the user as text, and then use natural language processing technology to analyze the text, understand the user's intention, and provide corresponding services according to the user's needs.
[0034] 3. Prompt: Refers to the text information used to describe a task, usually including the text input information converted from voice information and the relevant description information of the task that the large language model needs to execute. The prompt can guide the large language model to output the expected result.
[0035] 4. Plugin: Refers to an extension program of an application. Without changing the source code of the existing program, the plugin can add new functions to the application. The plugin can interact with the corresponding application to provide specific functions to the module that calls the plugin (such as LLM, etc.).
[0036] 5. Plugin Metadata Information: A collection of plugin-related information, which can include one or more of the following: plugin description information, the universal resource locator (URL) of the plugin, the request parameters of the plugin, the response parameters of the plugin, etc. Among them, the request parameters of the plugin are used to indicate the type of input parameters input to the plugin, and the response parameters (also called response information) of the plugin are used to indicate the type of target data information returned by the plugin.
[0037] 6. Plugin Description Information: Can be used to indicate the functions and properties of the plugin, etc.
[0038] 7. Intent Vector: It can be a set of floating-point numbers used to represent the corresponding user intent. Exemplarily, if the user intent is "query the weather tomorrow", the corresponding intent vector can be [0.792, -0.177, -0.107, 0.109, -0.542, …].
[0039] 8. Plugin Description Vector: It can be a set of floating-point numbers used to represent the corresponding plugin description information. Exemplarily, if the plugin description information is "weather query api, used for querying weather conditions, can query meteorology, temperature, humidity, perceived temperature, pressure, etc. at a specified time and location", the corresponding plugin description vector can be [-0.010028, -0.0229……].
[0040] In some application scenarios, an electronic device (which can be referred to as electronic device 100) can integrate an LLM in an internal software module. When a user interacts with the electronic device by voice, the electronic device can receive the voice information input by the user and convert the voice information into corresponding text input information. Then, the electronic device can determine the user intent (such as querying weather information, querying hotel information, etc.) from the text input information through the LLM. The electronic device can execute the task corresponding to the user intent (such as displaying weather information, displaying hotel information, etc.) through the LLM.
[0041] Figures 1A - 1D It is a user interface for a group of electronic devices provided by an embodiment of the present application to apply an LLM and interact with a user by voice.
[0042] Refer to Figures 1A - 1B , the electronic device 100 can receive an operation by the user to start the voice assistant. In response to this operation to start the voice assistant, the electronic device 100 can start the voice assistant and interact with the user by voice. This voice assistant can integrate the aforementioned LLM. Exemplarily, as Figure 1A shown, the electronic device 100 can receive a touch operation by the user pressing a physical button (such as the side power button as Figure 1A shown). In response to this touch operation, the electronic device 100 can start the voice assistant and display the user interface 1000. Or, exemplarily, as Figure 1B shown, the electronic device 100 can receive a voice wake-up word input by the user (such as "Hello, YOYO" as Figure 1B shown). In response to this voice wake-up word, the electronic device 100 can start the voice assistant and display the user interface 1000. Among them, the user interface 1000 can display a prompt icon 1000A of the voice assistant, and this prompt icon 1000A is used to prompt the user that the current electronic device 100 can perform the function of interacting with the user by voice.
[0043] Refer toFigure 1C When the electronic device 100 starts the voice assistant and displays the user interface 1000, the electronic device 100 can conduct voice interaction with the user through the voice assistant. For example, Figure 1C as shown, the electronic device 100 can receive the voice information "Query the weather in Shenzhen tomorrow" input by the user and convert the voice information into the corresponding text input information. The electronic device 100 can display the text input information on the user interface 1000.
[0044] Referring to Figure 1D , the electronic device 100 can determine from the text input information "Query the weather in Shenzhen tomorrow" that the user's intention is "Query the weather in Shenzhen tomorrow" through the voice assistant. Then, the electronic device 100 can provide the weather query service to the user through the voice assistant and display the query result on the user interface 1000. As Figure 1D shown, the electronic device 100 can display the query result (which can also be called the target text information) "Shenzhen, sunny tomorrow, 26 degrees" on the user interface 1000.
[0045] In the embodiments of the present application, Figures 1A - 1D the user interface shown is only used to exemplarily explain the present application and does not impose any limitation on the present application.
[0046] Figure 2 is a voice interaction process provided by the embodiments of the present application based on LLM.
[0047] As Figure 2 shown, the voice interaction process may include:
[0048] S201: The electronic device 100 pre-sets callable application plug-ins (which can be simply referred to as plug-ins).
[0049] In the embodiments of the present application, the electronic device 100 can receive the tick operation of the user for one or more plug-ins. In response to this operation, the electronic device 100 sets one or more pre-ticked plug-ins by the user as callable plug-ins; alternatively, the electronic device 100 can also pre-set callable plug-ins through the system.
[0050] S202: The electronic device 100 receives the voice information input by the user and converts the voice information into the corresponding text input information. Among them, the text input information may include the user's intention.
[0051] In the embodiments of the present application, when the electronic device 100 starts the voice assistant to conduct voice interaction with the user, the electronic device 100 can receive the voice information input by the user through the microphone. For example, as Figure 1CAs shown, the electronic device 100 can receive the voice information "Query the weather in Shenzhen tomorrow" input by the user through the microphone. Then, the electronic device 100 can convert the received voice information into corresponding text input information. Among them, the voice assistant can integrate the LLM, and the description of the LLM can refer to the foregoing description.
[0052] S203: The electronic device 100 constructs a prompt based on the text input information, the plugin description information of the callable plugins, etc.
[0053] In the embodiments of the present application, the prompt may include: text input information, plugin description information of callable plugins, other instructions, etc.
[0054] Exemplarily, if the electronic device 100 has pre-set 10 callable plugins, the electronic device 100 can pre-store the plugin description information of these 10 plugins. When constructing the prompt, the plugin description information of these 10 plugins is included in the prompt. That is to say, the prompt includes the plugin description information of all callable plugins.
[0055] S204: The electronic device 100 determines whether to call a plugin based on the prompt through the LLM.
[0056] S205: When it is determined that a plugin needs to be called, the electronic device 100 calls the plugin through the LLM, generates and outputs target text information that matches the user's intention.
[0057] In the embodiments of the present application, different user intentions may correspond to different plugins. For example, if the user intention is to query the weather, the plugin corresponding to this user intention is a weather application plugin; if the user intention is to query an address, the plugin corresponding to this user intention is a map application plugin.
[0058] Specifically, when the electronic device 100 determines that a plugin needs to be called, the electronic device 100 can call the plugin corresponding to the user intention in the text input information based on the prompt. The plugin can interact with the corresponding third-party application (for example, the weather application plugin interacts with the weather application, the map application plugin interacts with the map application, etc.), and obtains target data information that matches the user intention. Then, the plugin can return the target data information to the LLM. Then, the LLM can convert and output the target data information into corresponding target text information, and the target text information matches the user intention. Exemplarily, the above target text information can be displayed on the user interface of the electronic device 100, such as Figure 1D As shown, the "Shenzhen, sunny tomorrow, 26 degrees" displayed on the user interface 1000 is the target text information that matches the user intention.
[0059] S206: When it is determined that no plugin needs to be invoked, the electronic device 100 outputs target text information that matches the user's intention through the LLM.
[0060] In the embodiment of the present application, when the electronic device 100 determines that no plugin needs to be invoked, that is, the electronic device 100 does not need to generate target text information that matches the user's intention based on the target data information returned by the plugin. Therefore, the electronic device 100 can directly output target text information that matches the user's intention through the LLM.
[0061] However, Figure 2 The voice interaction process shown above has the following problems: 1). Since the electronic device 100 can integrate a large number of plugins, in the voice interaction process, the user needs to manually check the plugins that can be invoked or the system presets the plugins that can be invoked. Then, the electronic device 100 adds the plugin description information of the selected plugin to the prompt. It can be seen that the selection of plugins in this voice interaction process is not flexible enough and the operation is extremely cumbersome. 2). Due to the input length limitation of the LLM, if the text input information is too long, the plugin description information that can be input into the LLM will be relatively short, which may lead to the problem of incomplete input plugin description information. 3). Since different LLMs have different capabilities and the characteristics of third-party applications and their plugins in different business fields are different, the prompt also needs to be adjusted according to the type of the LLM and the type of the plugin. In this way, it will lead to the problems of waste of memory resources of the electronic device 100 and low voice interaction efficiency.
[0062] Therefore, the present application provides a voice interaction method, which can be applied to the electronic device 100. The method may include: The electronic device 100 can receive text input information. Wherein, the text input information may be text information converted from the voice information input by the user. Then, the electronic device 100 can determine the user intention text from the text input information through the large language model LLM. The electronic device 100 calculates the intention vector corresponding to the user intention text. The electronic device 100 determines M plugin description vectors that match the intention vector from the vector database based on the intention vector. Wherein, the vector database may include N plugin description vectors, and M is less than N. The electronic device 100 determines M plugin metadata information corresponding to the M plugin description vectors from the text database based on the M plugin description vectors. Wherein, the text database may include N plugin metadata information. The electronic device 100 determines K target plugins corresponding to the user intention text based on the user intention text and the plugin description information in the M plugin metadata information through the LLM. Wherein, K is less than or equal to M, and K, M, and N are positive integers. Next, the electronic device 100 can invoke the K target plugins through the LLM to generate target text information that matches the user's intention.
[0063] Among them, the user intention text represents the user intention in text form, while the intention vector represents the user intention in the form of a set of floating-point numbers.
[0064] When implementing the voice interaction method provided in this application, the electronic device 100 can screen out the callable plugins through the intention vector and the plugin description vector. Compared with the Figure 2 way that the system shown above pre-sets the callable plugins or the user pre-checks the callable plugins, it is more convenient, the selection of plugins is more accurate, and the efficiency is relatively high. At the same time, the electronic device 100 can also avoid inputting the plugin description information of irrelevant plugins, so that the electronic device 100 can provide corresponding services more efficiently and accurately based on the user intention.
[0065] Figure 3A This is a software architecture applied to the electronic device 100 provided by an embodiment of this application.
[0066] As Figure 3A shown, this software architecture (which can also be called the LLM plugin integration system) can be set in the application layer of the electronic device 100. The LLM plugin integration system can include: a large language model (LLM), a prompt processing module, a vector calculation module, a data storage module, a plugin management module, and a plugin selection module. Among them:
[0067] The large language model LLM can be used for:
[0068] 1). Receive the voice information input by the user, convert the voice information into the corresponding text input information, and determine the user intention text according to the text input information.
[0069] 2). It can also be used to determine the target plugin corresponding to the user intention (for example, the aforementioned K target plugins) according to the user intention text and the plugin description information in the plugin metadata information. The target plugin is an extension program of the target application and can interact with the target application.
[0070] 3). It can also be used to extract the input parameters that need to be input to the target plugin when calling the target plugin.
[0071] 4). It can also be used to receive the target data information returned after the target plugin and the target application interact, and convert the target data information into the target text information that matches the user intention, etc. The specific implementation method and other functions can refer to the subsequent embodiments.
[0072] The prompt processing module can be used for:
[0073] 1). Store prompt templates of different types. Among them, the prompt templates can include one or more of the following: intent recognition prompt templates, plug-in fine-tuning prompt templates, input parameter extraction prompt templates, and so on. In this way, plug-in call tasks in different business fields can directly use the above prompt templates without targeted modification, improving the efficiency of plug-in calls.
[0074] 2). Generate a complete prompt based on each type of prompt template. This complete prompt can be input into the large language model LLM to obtain a specified output result. Among them, the complete prompt can include one or more of the following: intent recognition prompt, plug-in fine-tuning prompt, input parameter extraction prompt, and so on. Among them, the intent recognition prompt can be used to determine the user intent text, the plug-in fine-tuning prompt is used to determine the plug-in corresponding to the user intent, and the input parameter extraction prompt is used to extract the input parameters input to the plug-in from the text input information. The specific implementation method and other functions can refer to the subsequent embodiments.
[0075] The Embedding Generator can be used for:
[0076] 1). Generate corresponding semantic vectors according to the original text information.
[0077] In the embodiments of the present application, the Embedding Generator can generate corresponding intent vectors according to the user intent text and generate corresponding plug-in description vectors according to the plug-in description information. The specific implementation method can refer to the subsequent embodiments.
[0078] The data storage module can include:
[0079] 1). A vector database. Among them, the vector database can be used to store semantic vectors. In the embodiments of the present application, the vector database can store one or more plug-in description vectors, etc.
[0080] 2). A text database. Among them, the text database can be used to store the text information corresponding to the semantic vectors and other text information. In the embodiments of the present application, the text database can store one or more plug-in metadata information, etc. Among them, the number of plug-in metadata information in the text database is the same as the number of plug-in description vectors in the vector database.
[0081] The plug-in management module can be used for:
[0082] 1). When a new plugin is added to the electronic device 100, the vector calculation module is called to calculate the corresponding plugin description vector according to the plugin description information of the new plugin, and store the plugin description vector in the vector database. At the same time, the plugin management module stores the plugin metadata information of the new plugin in the text database.
[0083] 2). Establish and store the mapping relationship between the plugin description vector and the plugin metadata information. When a certain plugin description vector and a certain plugin metadata information both correspond to the same plugin, a mapping relationship can be established between the plugin description vector and the plugin metadata information. For example, if the weather application plugin corresponds to the plugin description vector A1 and the plugin metadata information A2, the plugin management module can establish and store the mapping relationship between the plugin description vector A1 and the plugin metadata information A2.
[0084] 3). Store and manage the plugins.
[0085] 4). Invoke the plugin. In some embodiments, the plugin management module may include a plugin execution proxy module, and the plugin management module can invoke the plugin through the plugin execution proxy module.
[0086] The plugin selection module can be used for:
[0087] 1). Query one or more plugin description vectors that match the specified intent vector, and determine the plugin metadata information corresponding to the plugin description vector that matches the intent vector according to the mapping relationship between the plugin description vector and the plugin metadata information in the plugin management module. That is to say, the plugin selection module can be used to query the plugin metadata information that matches the specified intent vector. The specific implementation method can refer to the subsequent embodiments.
[0088] In a possible implementation manner, one or more modules in the LLM plugin integration system can be set in the cloud server that establishes a communication connection with the electronic device 100. For example, the plugin management module, the plugin selection module, the vector calculation module, the data storage module, the large language model (LLM), and the prompt processing module in the LLM plugin integration system can all be set in the cloud server. The electronic device 100 can encapsulate the interfaces corresponding to each module, and the electronic device 100 can call each module in the LLM plugin integration system on the cloud server through each interface. In the specific implementation manner, which modules in the LLM plugin integration system are set in the cloud server and which modules are set in the application layer of the electronic device 100 are not limited in this application.
[0089] Figure 3B This is a data stream interaction provided by the embodiments of this application.
[0090] Such as Figure 3BAs shown, when the electronic device 100 interacts with the user via voice and receives the voice information input by the user, the electronic device 100 can convert the voice information into the corresponding text input information. The plugin management module can receive the text input information and pass the text input information to the prompt processing module.
[0091] Based on the text input information, the prompt processing module can return the intent recognition prompt information to the plugin management module.
[0092] Then, the plugin management module can pass the intent recognition prompt information to the LLM. Based on its semantic understanding ability, the LLM determines the user intent text based on the intent recognition prompt information and returns the user intent text to the plugin management module.
[0093] The plugin management module can pass the user intent text to the vector calculation module. The vector calculation module can calculate the intent vector corresponding to the user intent text and return the intent vector to the plugin management module.
[0094] The plugin management module can pass the intent vector to the plugin selection module. The plugin selection module filters out M plugin description vectors that match the intent vector from the vector database. Then, according to the mapping relationship between the plugin description vectors and the plugin metadata information, the plugin selection module determines M plugin metadata information corresponding to the M plugin description vectors from the text database. The plugin selection module returns the M plugin metadata information to the plugin management module.
[0095] The plugin management module passes the user intent text and the plugin description information in the M plugin metadata information to the prompt processing module. The prompt processing module generates the plugin fine-tuning prompt information based on the user intent text and the M plugin description data information, and then returns the plugin fine-tuning prompt information to the plugin management module.
[0096] The plugin management module passes the plugin fine-tuning prompt information to the LLM. Based on its semantic understanding ability, the LLM determines the target plugin corresponding to the user intent based on the plugin fine-tuning prompt information and returns the target plugin name to the plugin management module.
[0097] The plugin management module passes the text input information and the request parameters of the target plugin to the prompt processing module. The prompt processing module generates the input parameter extraction prompt information based on the text input information and the request parameters of the target plugin, and returns the input parameter extraction prompt information to the plugin management module.
[0098] The plugin management module passes the input parameter extraction prompt information to the LLM. Based on its semantic understanding ability, the LLM extracts the input parameter from the text input information according to the input parameter extraction prompt information and returns the input parameter to the plugin management module.
[0099] Based on the input parameter, the plugin management module calls the target plugin through the plugin execution proxy module to obtain the target data information. Then, the plugin management module passes the target data information and the text input information to the LLM. The LLM generates the target text information that matches the user's intention according to the target data information and the text input information and returns the target text information to the plugin management module.
[0100] The plugin management module can output the target text information to the user.
[0101] Combined with Figure 3A the embodiments shown, Figure 4 This is the specific process of a voice interaction method provided by an embodiment of the present application.
[0102] As Figure 4 shown, the specific process of the voice interaction method may include:
[0103] S401: The electronic device 100 receives the voice information input by the user through the microphone.
[0104] In the embodiment of the present application, when the electronic device 100 starts the voice assistant to interact with the user by voice, the electronic device 100 can receive the voice information input by the user through the microphone. For example, as Figure 1C shown, the electronic device 100 can receive the voice information "Query the weather in Shenzhen tomorrow" input by the user through the microphone. Among them, the voice assistant can integrate the LLM, and the description of the LLM can refer to the foregoing description.
[0105] Among them, for the implementation manner of the electronic device 100 to start the voice assistant, reference can be made to the description in the foregoing Figures 1A - 1B shown embodiments, which will not be elaborated here.
[0106] S402: The electronic device 100 converts the voice information into the corresponding text input information.
[0107] In some embodiments, in addition to converting the received voice information into the corresponding text input information, the electronic device 100 can also receive the touch operation of the user on the virtual keyboard / physical keyboard to obtain the text input information. That is to say, for the acquisition method of the text input information, the present application does not limit this.
[0108] S403: The electronic device 100 generates an intent recognition prompt message based on the intent recognition prompt message template and the text input information.
[0109] In the embodiments of the present application, the electronic device 100 can fill the text input information into the intent recognition prompt message template through a prompt message processing module to obtain the intent recognition prompt message.
[0110] S404: The electronic device 100 determines the user intent text through the LLM based on the intent recognition prompt message.
[0111] In the embodiments of the present application, when the electronic device 100 generates an intent recognition prompt message through the prompt message processing module, the electronic device 100 can input the intent recognition prompt message into the LLM. After receiving the intent recognition prompt message, the LLM can determine the user intent text according to its semantic understanding ability.
[0112] S405: The electronic device 100 calculates the intent vector corresponding to the user intent text.
[0113] In the embodiments of the present application, the electronic device 100 can input the determined user intent text into a vector calculation module. The vector calculation module can calculate the intent vector corresponding to the user intent text. For the description of the intent vector, reference can be made to the description in the foregoing embodiments.
[0114] S406: The electronic device 100 determines M plugin description vectors that match the intent vector.
[0115] In the embodiments of the present application, the electronic device 100 can input the intent vector to a plugin selection module, and then the plugin selection module determines M plugin description vectors from the vector database whose similarity to the intent vector is greater than or equal to a threshold 1. Specifically, in one implementation, the electronic device 100 can determine M plugin description vectors from the vector database whose distance difference from the intent vector is less than or equal to a threshold 2. Among them, the vector database may include N plugin description vectors, and M is less than N.
[0116] S407: The electronic device 100 determines M plugin metadata information corresponding to the M plugin description vectors according to the mapping relationship between the plugin description vectors and the plugin metadata information.
[0117] In the embodiments of the present application, the plugin management module can pre-store the mapping relationship between the plugin description vectors and the plugin metadata information. The plugin selection module can determine M plugin metadata information corresponding to the M plugin description vectors from the text database according to the above mapping relationship in the plugin management module. Among them, the text database may include N plugin metadata information.
[0118] S408: The electronic device 100 determines K target plugins corresponding to the user intention through the LLM based on the user intention text and M plugin metadata information.
[0119] In the embodiments of the present application, the electronic device may input the user intention text and the plugin description information in the M plugin metadata information to a prompt processing module. Then, the prompt processing module fills the user intention text and the M plugin description information into the plugin fine-tuning prompt information template to obtain the plugin fine-tuning prompt information. Next, the prompt processing module may pass the plugin fine-tuning prompt information to the LLM. The LMM determines K target plugins corresponding to the user intention based on the plugin fine-tuning prompt information according to its semantic understanding ability. Among them, K can take values of 1, 2, 3, etc., and the specific value is not limited. K is less than or equal to M, and K, M, and N are positive integers.
[0120] S409: The electronic device 100 invokes K target plugins to generate target text information matching the user intention.
[0121] In the embodiments of the present application, the electronic device 100 may pass the text input information and the request parameters of the K target plugins to the prompt processing module. The prompt processing module may fill the text input information and the request parameters of the K target plugins into the input parameter extraction prompt information template to obtain the input parameter extraction prompt information. Then, the prompt processing module may pass the input parameter extraction prompt information to the LLM. The LLM extracts the input parameters corresponding to each of the K target plugins from the text input information based on the input parameter extraction prompt information according to its semantic understanding ability, and sends the above input parameters to the plugin management module. The plugin management module may invoke the K target plugins based on the input parameters, receive the target data information returned by the K target plugins. The plugin management module may convert the target data information into target text information matching the user intention through the LLM and output the target text information to the user.
[0122] Further, Figure 4 The specific implementation methods of the steps in the illustrated embodiments are described in detail.
[0123] As Figure 5A shown, the specific implementation of steps S403 to S404 may be as follows:
[0124] a). The plugin management module (Plugin Manager) may receive the text input information, and then input the text input information into the prompt processing module.
[0125] b). The prompt information processing module can fill the text input information into the intent recognition prompt information template to obtain the intent recognition prompt information.
[0126] Specifically, the intent recognition prompt information template can be pre-stored in the prompt information processing module. The intent recognition prompt information template can include: a text input information filling area, an intent recognition task description, etc.
[0127] Exemplarily, the intent recognition prompt information template can specifically be: "Expression: ${input}. Please identify the intent in the above expression. Note that only directly answer the intent, and the intent expression should be as short as possible. Multiple intents are separated by commas." Among them, "${input}" is the text input information filling area, and "Please identify the intent in the above expression. Note that only directly answer the intent, and the intent expression should be as short as possible. Multiple intents are separated by commas" is the intent recognition task description.
[0128] The prompt information processing module can fill the text input information into the text input information filling area in the intent recognition prompt information template to obtain the intent recognition prompt information.
[0129] Exemplarily, if the text input information is "Hello gpt, please help me check the weather in Nanjing today", when the prompt information processing module receives the text input information, it can fill the text input information into the text input information filling area in the intent recognition prompt information template, and the obtained intent recognition prompt information is "Expression: Hello gpt, please help me check the weather in Nanjing today. Please identify the intent in the above expression. Note that only directly answer the intent, and the intent expression should be as short as possible. Multiple intents are separated by commas."
[0130] c). The prompt information processing module inputs the generated intent recognition prompt information into the LLM through the plugin management module (PluginManager).
[0131] d). The LLM determines the user intent text based on the intent recognition prompt information according to its semantic understanding ability.
[0132] e). The LLM returns the user intent text to the plugin management module (Plugin Manager).
[0133] Among them, the LLM can input the user intent text to the plugin management module (PluginManager) in the form of a list, and this list can be called the intent list. Not limited to the list, the LLM can also input the user intent text to the plugin management module (Plugin Manager) in other forms, and this application does not limit this.
[0134] Exemplarily, if the text input information is "Hello, GPT. Please help me check the weather in Nanjing today.", the specific code implementation can be as follows:
[0135] input = "Hello, GPT. Please help me check the weather in Nanjing today."
[0136] print = ("Original input:", input, "\n")
[0137] intents_list = intents_recognition(input)
[0138] print("Intent list", intents_list, "\n")
[0139] That is to say, the text input information, namely the original input: Hello, GPT. Please help me check the weather in Nanjing today, generates the intent recognition prompt information: "Expression: Hello, GPT. Please help me check the weather in Nanjing today. Please identify the intent in the above expression. Note that only directly answer the intent, and the intent expression should be as short as possible. Multiple intents are separated by commas." The user intent text determined by the LLM is "Check the weather in Nanjing". The text input information is not limited to the above example. In the specific implementation, it can also be other text input information, and this application does not limit it.
[0140] Another exemplarily, as Figure 5B shown, if the text input information is "I am in Xinjiekou and want to have spicy hot pot for about 20 yuan. Please help me check the surrounding spicy hot pot takeaways. Additionally, help me check what hotels are there within 1 kilometer around Nanjing Glory with prices between 300 yuan and 500 yuan." Then the intent recognition prompt information generated by the prompt information processing module based on this text input information is: "Expression: I am in Xinjiekou and want to have spicy hot pot for about 20 yuan. Please help me check the surrounding spicy hot pot takeaways. Additionally, help me check what hotels are there within 1 kilometer around Nanjing Glory with prices between 300 yuan and 500 yuan. Please identify the intent in the above expression. Note that only directly answer the intent, and the intent expression should be as short as possible. Multiple intents are separated by commas." The user intent text determined by the LLM can be: "Check spicy hot pot takeaways, check hotels within 1 kilometer around Nanjing Glory with prices between 300 yuan and 500 yuan."
[0141] As Figure 5C shown, the specific implementation of steps S406 to S408 can be as follows:
[0142] Step S406:
[0143] a). The Plugin Manager inputs the user intention text into the vector calculation module.
[0144] b). The vector calculation module calculates the intention vector corresponding to the user intention text.
[0145] c). The vector calculation module inputs the intention vector to the plugin selection module through the plugin management module.
[0146] Among them, the intention vector can be a set of floating-point numbers, used to represent the corresponding user intention.
[0147] d). When receiving the intention vector, the plugin selection module can determine M plugin description vectors matching the intention vector from the vector database ( Figure 5C not shown in the figure).
[0148] In one implementation, when receiving the intention vector, the plugin selection module can calculate the distance difference between each plugin description vector in the vector database and the intention vector. Then, the plugin selection module can determine M plugin description vectors from the vector database whose distance difference from the intention vector is less than or equal to threshold 2, and the M plugin description vectors are the plugin description vectors matching the intention vector. Among them, threshold 2 can be 0.1, 0.3, 0.2, etc., or other values, and the present application does not limit this.
[0149] Step S407:
[0150] e). The plugin selection module can determine M plugin metadata information corresponding to the above M plugin description vectors from the text database ( Figure 5C not shown in the figure) according to the mapping relationship between the plugin description vectors and the plugin metadata information in the plugin management module.
[0151] In the embodiments of the present application, the text database may store N pieces of plugin metadata information, the vector database may store N plugin description vectors, and the plugin management module may store the mapping relationship between each plugin description vector and the plugin metadata information.
[0152] Exemplarily, the mapping relationship between each plugin description vector and the plugin metadata information may be as shown in Table 1:
[0153] Table 1
[0154] Plugin description vector Plugin metadata information Plugin description vector 1 Plugin metadata information 1 Plugin description vector 2 Plugin metadata information 2 Plugin description vector 3 Plugin metadata information 3 …… ……
[0155] As shown in Table 1, plugin description vector 1 can correspond to plugin metadata information 1, plugin description vector 2 can correspond to plugin metadata information 2, plugin description vector 3 can correspond to plugin metadata information 3, and so on. Table 1 is only used for exemplary explanation of this application and does not impose any limitation on this application.
[0156] Exemplarily, the code implementation for querying plugin metadata information can be:
[0157] i = 0;
[0158] plugin_info_dict = plugin_search(intents_list[i]);
[0159] For example, when the user intent is "query the weather in Nanjing", the retrieved plugin metadata information can be:
[0160] Plugin description information: Weather query API, used for querying weather conditions, and can query meteorology, temperature, humidity, felt temperature, pressure, etc. at a specified time and location
[0161] Plugin URL: http: / / ${IP}:${port} / aicloud / connector / v1 / fulfillment / weather / query
[0162] Plugin parameter information: City
[0163] Plugin request parameter: city
[0164] Plugin response information: weatherid, temperature, humidity, pressure, updatetime, realfeel, mobileLink
[0165] Step S408:
[0166] f). The plugin selection module sends the determined M pieces of plugin metadata information to the plugin management module.
[0167] g). The plugin management module inputs the plugin description information and the user intent text in the received M pieces of plugin metadata information to the prompt information processing module.
[0168] Among them, the plugin management module can input the plugin description information to the prompt information processing module in the form of a list. Without limitation to the list, the plugin management module can also input the plugin description information to the prompt information processing module in other forms, and this application does not impose any limitation on this.
[0169] h). The prompt information processing module fills the M plugin description information and the user intention text into the plugin fine-ranking prompt information template to generate the plugin fine-ranking prompt information.
[0170] Among them, the prompt information processing module can pre-store the plugin fine-ranking prompt information template. The plugin fine-ranking prompt information template can include: the user intention text filling area, the plugin description information filling area, the plugin query task description, etc.
[0171] Exemplarily, the plugin fine-ranking prompt information template can specifically be: "Intention list: ${intents_list}. For these intents, the system can provide some auxiliary plugin apis, and the names and capabilities are described as follows: ${plugin_description_list}. Please return the plugin apis that should be used for each intent in turn. Note: Only return the plugin api names, and multiple names are separated by commas. Use null to represent the intent plugin that does not meet the requirements." Among them, "${intents_list}" is the user intention text filling area, "${plugin_description_list}" is the plugin description information filling area, and "Please return the plugin apis that should be used for each intent in turn. Note: Only return the plugin api names, and multiple names are separated by commas. Use null to represent the intent plugin that does not meet the requirements." is the plugin query task description.
[0172] The prompt information processing module can fill the user intention text into the user intention text filling area in the plugin fine-ranking prompt information template, and fill the plugin description information into the plugin description information filling area to obtain the plugin fine-ranking prompt information.
[0173] Exemplarily, if the user intention text is "Query the weather in Nanjing", "Order takeout", the plug-in description information is "Weather query API, used for querying weather conditions, and can query meteorological, temperature, humidity, perceived temperature, pressure, etc. information of a specified time and location", "Train ticket and high-speed rail ticket query API, used for querying ticket information such as train, high-speed rail, and railway flight numbers, departure times or dates, and provides railway operation information such as stations, departure times, arrival times, and stop times". When the prompt information processing module receives the user intention text and the plug-in description information, the prompt information processing module can fill the user intention text into the user intention text filling area of the plug-in fine-rank prompt information template, and fill the plug-in description information into the plug-in description information filling area of the plug-in fine-rank prompt information template. The obtained plug-in fine-rank prompt information is "Intention list: Query the weather in Nanjing, Order takeout. For these intentions, the system can provide some auxiliary plug-in APIs, the names and capabilities are described as follows: Weather query API, used for querying weather conditions, and can query meteorological, temperature, humidity, perceived temperature, pressure, etc. information of a specified time and location. Train ticket and high-speed rail ticket query API, used for querying ticket information such as train, high-speed rail, and railway flight numbers, departure times or dates, and provides railway operation information such as stations, departure times, arrival times, and stop times. Please return the plug-in APIs that should be used for each intention in turn. Note: Only return the plug-in API names, and multiple names are separated by commas. Use null to represent the intention plug-in that does not meet the requirements."
[0174] i). The prompt information processing module inputs the generated plug-in fine-rank prompt information to the LLM through the plug-in management module.
[0175] j). The LLM determines K target plug-ins corresponding to the user intention based on the plug-in fine-rank prompt information according to the semantic understanding ability.
[0176] Among them, since there is a corresponding relationship between the user intention and the intention vector, the K target plug-ins corresponding to the user intention are also the plug-ins corresponding to the intention vector of the user intention.
[0177] Exemplarily, such as Figure 5DAs shown, if the plug-in fine-tuning prompt information is "Intent list: Query the weather in Nanjing, order takeout. For these intents, the system can provide some auxiliary plug-in APIs, the names and capabilities are described as follows: Weather query API, used to query weather conditions, and can query meteorological, temperature, humidity, perceived temperature, pressure, etc. information at a specified time and location. Train ticket and high-speed rail ticket query API, used to query ticket information such as train, high-speed rail, and railway flight numbers, departure times, or dates, and provides railway operation information such as stations, departure times, arrival times, and stop times. Please return the plug-in APIs that should be used for each intent in sequence. Note: Only return the names of the plug-in APIs, separated by commas for multiple names, and use null to represent the intent plug-in that does not meet the requirements.", then the target plug-ins determined by the LLM can be "Weather query API, null".
[0178] Also, by way of example, as Figure 5E shown, if the plug-in fine-tuning prompt information is "Intent list: Query the weather in Nanjing, order high-speed rail. For these intents, the system can provide some auxiliary plug-in APIs, the names and capabilities are described as follows: Weather query API, used to query weather conditions, and can query meteorological, temperature, humidity, perceived temperature, pressure, etc. information at a specified time and location. Train ticket and high-speed rail ticket query API, used to query ticket information such as train, high-speed rail, and railway flight numbers, departure times, or dates, and provides railway operation information such as stations, departure times, arrival times, and stop times. Please return the plug-in APIs that should be used for each intent in sequence. Note: Only return the names of the plug-in APIs, separated by commas for multiple names, and use null to represent the intent plug-in that does not meet the requirements.", then the target plug-ins determined by the LLM can be "Weather query API, Train ticket and high-speed rail ticket query API".
[0179] k). The LLM returns the names of the K determined target plug-ins to the plug-in management module.
[0180] Among them, the LLM can return the names of the K target plug-ins to the plug-in management module in the form of a list. Not limited to a list, the LLM can also return the names of the K target plug-ins to the plug-in management module in other forms, and this application does not limit this.
[0181] As Figure 5F shown, the specific implementation of step S409 can be as follows:
[0182] a). The plug-in management module can input the text input information and the request parameters in the plug-in metadata information of the K target plug-ins (which can be abbreviated as the request parameters of the K target plug-ins) to the prompt information processing module.
[0183] b). The prompt information processing module can fill the text input information and the request parameters of K target plugins into the input parameter extraction prompt information template to obtain the input parameter extraction prompt information.
[0184] Specifically, the prompt information processing module can pre-store the input parameter extraction prompt information template. The input parameter extraction prompt information template can include one or more of the following: text input information filling area, request parameter filling area, input parameter extraction task description, etc.
[0185] Exemplarily, the input parameter extraction prompt information template can specifically be: "Expression: ${input}. Please extract the following information from the above expression: ${keys}. The returned results are separated by commas, and the information not extracted is represented by null." Among them, "${input}" can be called the text input information filling area, "${keys}" can be called the request parameter filling area, and "Please extract the following information from the above expression: ${keys}. The returned results are separated by commas, and the information not extracted is represented by null" can be called the input parameter extraction task description.
[0186] The prompt information processing module can fill the text input information into the text input information filling area in the input parameter extraction prompt information template, and fill the request parameters of K target plugins into the request parameter filling area in the input parameter extraction prompt information template to obtain the input parameter extraction prompt information.
[0187] Exemplarily, if the text input information is "Hello gpt, please help me check the weather in Nanjing today", and the request parameter of the target plugin is "city", then the obtained input parameter extraction prompt information can be "Expression: Hello gpt, please help me check the weather in Nanjing today. Please extract the following information from the above expression: city. The returned results are separated by commas, and the information not extracted is represented by null."
[0188] c). The prompt information processing module can input the input parameter extraction prompt information to the LLM through the plugin management module.
[0189] d). The LLM extracts the input parameters from the text input information based on the input parameter extraction prompt information according to its semantic understanding ability.
[0190] Exemplarily, the code implementation of this step can be as follows:
[0191] request_body = plugin_params_extraction(input, plugin_info_dict['params'], plugin_info_dict['request'])
[0192] Exemplarily, if the input parameter extraction prompt information is "Expression: Hello gpt, please help me query the weather in Nanjing today. Please extract the following information from the above expression: city. The return result is separated by commas, and the information not extracted is represented by null.", then the input parameter extracted by the LLM from the text input information "Hello gpt, please help me query the weather in Nanjing today" is "Nanjing".
[0193] e). The LLM returns the input parameter to the plugin management module.
[0194] f). The plugin management module inputs the plugin metadata information of the K target plugins and the input parameter to the plugin execution proxy module, and calls the K target plugins through the plugin execution proxy module.
[0195] Among them, when the plugin execution proxy module calls the K target plugins, the plugin execution proxy module can input the input parameter to the target plugin.
[0196] g). The plugin management module receives the target data information returned by the K target plugins.
[0197] Exemplarily, the code implementation for calling the target plugin to obtain the target data information can be:
[0198] plugin_call(plugin_info_dict['url'], request_body).
[0199] Among them, the plugin execution proxy module can call the target plugin based on the url of the target plugin.
[0200] Exemplarily, if the text input information is "Hello gpt, please help me query the weather in Nanjing today", the extracted input parameter is "Nanjing", and the user intention is "query the weather in Nanjing", then the called plugin is the weather query api. Then the plugin execution proxy module calls the weather query api based on the following input parameter and url:
[0201] Url of the weather query api: http: / / ${IP}:${port} / aicloud / connector / v1 / fulfillment / weather / query
[0202] Body: {"city": "Nanjing"}
[0203] Then the target data information returned by the weather query api can be:[[]]
[0204] Plugin call result: {'weatherid': 'Cloudy', 'temperature': '28.0', 'humidity': '74', 'pressure': '996.0', 'updatetime': '1688464469000','realfeel': '32.0','mobileLink': 'https: / / weather.com / zh / cn'}
[0205] h). The plugin management module inputs the target data information and the text input information to the LLM.
[0206] In the embodiment of the present application, since the target data information returned by the target plugin is the response body of the interface, generally in JSON structure, it is necessary to use the LLM to rewrite it into a colloquial expression that conforms to the human expression way (that is, the target text information). The plugin management module inputs the text input information and the target data information into the LLM together, which can enhance the effect of the LLM to transcribe the target data information into the target text information.
[0207] i). Based on the target data information and the text input information, the LLM converts the target data information into the target text information that matches the user's intention.
[0208] j). The LLM outputs the target text information to the user through the plugin management module.
[0209] Exemplarily, as Figure 1D shown, the electronic device 100 can display the target text information in the user interface. Another exemplarily, the electronic device 100 can play the target text information by voice. That is to say, the present application does not limit the way the electronic device 100 outputs the target text information.
[0210] In the embodiment of the present application, as Figure 5G shown, the implementation manner of the electronic device 100 to add a specified plugin can be as follows:
[0211] a). The plugin management module obtains the plugin metadata information of the specified plugin.
[0212] Among them, the description of the plugin metadata information can refer to the foregoing embodiments and will not be elaborated here.
[0213] Specifically, when the electronic device 100 adds a specified plugin, the electronic device 100 can construct the plugin description information, the url of the specified plugin, the request parameters of the specified plugin, and the response parameters of the specified plugin of the obtained specified plugin into the plugin metadata information.
[0214] b). The plugin management module inputs the plugin description information of the specified plugin to the vector calculation module.
[0215] c). The vector calculation module can calculate the plug-in description vector corresponding to the plug-in description information.
[0216] d). The vector calculation module can return the plug-in description vector to the plug-in management module.
[0217] e). The plug-in management module sends the plug-in metadata information and the plug-in description vector of the specified plug-in to the data storage module for storage.
[0218] Among them, the plug-in metadata information can be stored in the text database in the data storage module, and the plug-in description vector can be stored in the vector database in the data storage module.
[0219] f). The plug-in management module establishes and stores the mapping relationship between the plug-in metadata information and the plug-in description vector of the specified plug-in.
[0220] In some examples, the plug-in management module can also store the newly added specified plug-in.
[0221] In the embodiments of the present application, the threshold 1 can be referred to as the first threshold. The mapping relationship between each plug-in description vector and each plug-in metadata information can be included in the first mapping relationship. The plug-in refined ranking prompt information template can be referred to as the first prompt information prompt template, and the plug-in refined ranking prompt information can be referred to as the first prompt information prompt. The input parameter extraction prompt information template can be referred to as the second prompt information prompt template, and the input parameter extraction prompt information can be referred to as the second prompt information prompt. The intent recognition prompt information template can be referred to as the third prompt information prompt template, and the intent recognition prompt information can be referred to as the third prompt information prompt. The K target plug-ins can include one or more of the following: weather query plug-in, map plug-in, train and high-speed rail ticket booking plug-in, movie ticket ordering plug-in, hotel query plug-in, etc.
[0222] Figure 6 This is the hardware structure of an electronic device 100 provided by the embodiments of the present application.
[0223] As Figure 6 shown, the electronic device 100 may include a processor 601, a memory 602, a wireless communication module 603 (optional), a display screen 604, a camera 605, an audio module 606 (optional), and a microphone 607 (optional). The processor 601, the memory 602, the wireless communication module 603 (optional), the display screen 604, the camera 605, the audio module 606 (optional), and the microphone 607 (optional) can be connected through a bus.
[0224] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may further include more or fewer components than those Figure 6 shown, or combine certain components, or split certain components, or have different component arrangements. Figure 6 The components shown can be implemented in hardware, software, or a combination of software and hardware.
[0225] The processor 601 may include one or more processor units. For example, the processor 601 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors. The controller can generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
[0226] A memory may also be provided in the processor 601 for storing instructions and data. In some embodiments, the memory in the processor 601 is a cache memory. This memory can save the instructions or data that the processor 601 has just used or recycled. If the processor 601 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 601, and thus improves the efficiency of the system.
[0227] In some embodiments, the processor 601 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a USB interface, etc.
[0228] The memory 602 is coupled to the processor 601 and is used to store various software programs and / or multiple sets of instructions. In a specific implementation, the memory 602 may include a volatile memory, such as a random access memory (RAM); it may also include a non-volatile memory, such as a ROM, a flash memory, a hard disk drive (HDD), or a solid state drive (SSD); the memory 602 may also include a combination of the above types of memories. The memory 602 may also store some program codes so that the processor 601 can call the program codes stored in the memory 602 to implement the implementation method of the embodiments of the present application in the electronic device 100. The memory 602 may store an operating system, such as an embedded operating system like uCOS, VxWorks, RTLinux, etc.
[0229] The wireless communication module 603 can provide solutions for wireless communications applied to the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. The wireless communication module 603 can be one or more devices integrating at least one communication processing module. The wireless communication module 603 receives electromagnetic waves via an antenna, performs frequency modulation and filtering processing on the electromagnetic wave signals, and sends the processed signals to the processor 601. The wireless communication module 603 can also receive signals to be sent from the processor 601, perform frequency modulation and amplification on them, and convert them into electromagnetic waves through the antenna for radiation. In some embodiments, the electronic device 100 can also detect or scan devices near the electronic device 100 by transmitting signals through the Bluetooth module ( Figure 6 not shown) and the WLAN module ( Figure 6 not shown) in the wireless communication module 603, and establish a wireless communication connection with the nearby devices to transmit data. Among them, the Bluetooth module can provide solutions for one or more Bluetooth communications including classic Bluetooth (basic rate / enhanced data rate, BR / EDR) or Bluetooth low energy (BLE), and the WLAN module can provide solutions for one or more WLAN communications including Wi-Fi direct, Wi-Fi LAN, or Wi-Fi softAP.
[0230] The display screen 604 can be used to display images, videos, etc. The display screen 604 may include a display panel. The display panel may adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 604, where N is a positive integer greater than 1.
[0231] The camera 605 is used to capture static images or videos. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element may be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard format such as RGB or YUV. In some embodiments, the electronic device 100 may include one or N cameras 605, where N is a positive integer greater than 1.
[0232] The audio module 606 can be used to convert digital audio information into an analog audio signal for output, or to convert an analog audio input into a digital audio signal. The audio module 606 can also be used to encode and decode audio signals. In some embodiments, the audio module 606 may also be disposed in the processor 601, or some functional modules of the audio module 606 may be disposed in the processor 601.
[0233] The microphone 607, also known as a "microphone" or "transmitter", can be used to collect sound signals in the environment surrounding the electronic device, convert the sound signals into electrical signals, and then process the electrical signals through a series of operations, such as analog-to-digital conversion, to obtain audio signals in digital form that can be processed by the processor 601 of the electronic device. When making a call or sending a voice message, the user can speak close to the microphone 607 with their mouth to input the sound signal into the microphone 607. The electronic device 100 can be provided with at least one microphone 607. In some other embodiments, the electronic device 100 can be provided with two microphones 607, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the electronic device 100 can also be provided with three, four or more microphones 607 to collect sound signals, reduce noise, identify the sound source, and implement functions such as directional recording.
[0234] The electronic device 100 may further include a sensor module ( Figure 6 not shown in the figure). The sensor module may include multiple sensor components, such as a touch sensor ( Figure 6 not shown in the figure), etc. The touch sensor may also be referred to as a "touch control device". The touch sensor may be disposed on the display screen 604, and together with the display screen 604, it forms a touch screen, also known as a "touch control screen". The touch sensor can be used to detect touch operations acting thereon or nearby.
[0235] It should be noted that Figure 6 the electronic device 100 shown in the figure is only used to exemplarily explain the hardware structure of the electronic device provided in this application and does not constitute a specific limitation to this application.
[0236] In the above embodiments, depending on the context, the term "when..." can be interpreted to mean "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if detecting (the stated condition or event)" can be interpreted to mean "if determining...", "in response to determining...", "when detecting (the stated condition or event)", or "in response to detecting (the stated condition or event)".
[0237] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, or a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive), etc.
[0238] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware with a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it may include the processes of the above method embodiments. The foregoing storage medium includes various media that can store program codes, such as ROM or random access memory RAM, magnetic disks, or optical discs.
Claims
1. A voice interaction method, characterized in that, comprising: receiving text input information; determining user intent text from the text input information through a large language model LLM; calculating an intent vector corresponding to the user intent text; wherein, the intent vector is used to represent the user intent; based on the intent vector, determining M plugin description vectors matching the intent vector from a vector database; wherein, the vector database includes N plugin description vectors, and M is less than N; based on the M plugin description vectors, determining M plugin metadata information corresponding to the M plugin description vectors from a text database; wherein, the text database includes N plugin metadata information; through the LLM, determining K target plugins corresponding to the user intent based on the user intent text and the M plugin metadata information; wherein, K is less than M, and K, M, and N are positive integers; invoking the K target plugins to generate target text information matching the user intent.
2. The method according to claim 1, characterized in that, the determining M plugin description vectors matching the intent vector from the vector database based on the intent vector includes: determining M plugin description vectors whose similarity with the intent vector is greater than or equal to a first threshold.
3. The method according to claim 1 or 2, characterized in that, the determining M plugin metadata information corresponding to the M plugin description vectors from the text database based on the M plugin description vectors includes: determining the M plugin metadata information corresponding to the M plugin description vectors based on a first mapping relationship; wherein, the first mapping relationship includes the mapping relationship between each plugin description vector and each plugin metadata information.
4. The method according to claim 1, characterized in that, the determining K target plugins corresponding to the user intent through the LLM based on the user intent text and the M plugin metadata information includes: filling the user intent text and the plugin description information in the M plugin metadata information into a first prompt template to obtain a first prompt; determining K target plugins corresponding to the user intent through the LLM based on the first prompt.
5. The method according to claim 1, characterized in that, the invoking the K target plugins to generate target text information matching the user intent includes: filling the text input information and the request parameters in the M plugin metadata information into a second prompt template to obtain a second prompt; extracting input parameters from the text input information through the LLM based on the second prompt; invoking the K target plugins based on the input parameters to generate target text information matching the user intent.
6. The method according to claim 1, characterized in that, The determination of the user intention text from the text input information by the large language model LLM includes: Filling the text input information into a third prompt template to obtain a third prompt; Determining the user intention text from the text input information by the LLM based on the third prompt.
7. The method according to claim 1, wherein, the LLM is ChatGPT-3, ChatGPT-4, BERT or XLNet.
8. The method according to claim 1, wherein, before receiving the text input information, the method further includes: Receiving input voice information and converting the voice information into the text input information; after receiving the text input information, the method further includes: Displaying the text input information; after calling the K target plugins to generate target text information matching the user intention, the method further includes: Displaying the target text information.
9. The method according to claim 1, wherein, the plugin metadata information includes one or more of the following: plugin description information, the uniform resource locator URL of the plugin, the request parameters of the plugin, and the response parameters of the plugin.
10. The method according to claim 1, wherein, the K target plugins include one or more of the following: weather query plugin, map plugin, train and high-speed rail ticket booking plugin, movie ticket ordering plugin, hotel query plugin.
11. An electronic device, wherein, it includes: one or more processors and one or more memories; the one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer-executable programs. When the one or more processors execute the computer-executable programs, the electronic device executes the method according to any one of claims 1-10.
12. A chip system, wherein, it includes a processing circuit and an interface circuit. The interface circuit is used to receive code instructions and transmit them to the processing circuit, and the processing circuit is used to run the code instructions so that the chip system executes the method according to any one of claims 1-10.
13. A computer-readable storage medium, wherein, it stores a computer-executable program. When the computer-executable program runs on an electronic device, the electronic device executes the method according to any one of claims 1-10.
Citation Information
Patent Citations
Metadata search method and device, electronic equipment and storage medium
CN110442614A
Generative large language model training method and model-based man-machine voice interaction method
CN116127045A
Human-computer interaction method, device and system
CN116483980A
Creation of language models for speech recognition
US10943583B1
Supporting compilation and extensibility on unified graph-based intent models
US20200274772A1