Voice interaction methods and electronic devices
By using a large language model to calculate user intent vectors in voice interaction, and filtering plugin description vectors and metadata information, the problems of inflexible plugin selection and resource waste are solved, resulting in a more efficient and accurate voice interaction service.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HONOR DEVICE CO LTD
- Filing Date
- 2023-11-17
- Publication Date
- 2026-04-21
AI Technical Summary
In existing technologies, plugin selection during voice interaction is not flexible enough, the operation is cumbersome, the input length limit of LLM results in incomplete plugin description information, and the differences between different LLM and plugin types lead to wasted memory resources, resulting in low voice interaction efficiency.
The user intent vector is calculated using a large language model (LLM), and matching plugin description vectors and metadata information are filtered out. The target plugin is then called to generate the target text information. Plugin selection is performed using both vector and text databases to avoid manual selection and resource waste.
It enables more flexible, accurate, and efficient plugin selection, improves voice interaction efficiency, reduces resource waste, and ensures the accuracy of plugin selection and the efficiency of services.
Smart Images

Figure CN120067284B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminals, and more particularly to a voice interaction method and electronic device. Background Technology
[0002] Artificial intelligence (AI), a branch of computer science, studies the design principles and implementation methods of various intelligent machines, enabling electronic devices to possess perception, reasoning, and decision-making capabilities. Since the release of ChatGPT, a system developed based on AI technology, the capabilities and future potential of large-scale foundational models (e.g., large language models (LLMs)) have attracted widespread attention. Question-answering systems based on large language models can interact with users, helping them handle daily tasks, such as answering user questions and searching for hotel rooms based on user needs. Therefore, improving the efficiency of human-computer interaction between large language models and users has become an urgent problem to be solved. Summary of the Invention
[0003] This application provides a voice interaction method and an electronic device that enables accurate and efficient selection of the plugin corresponding to the user's intent, allowing the electronic device to provide corresponding services more efficiently and accurately based on the user's intent.
[0004] In a first aspect, this application provides a voice interaction method, comprising: receiving text input information; determining user intent text from the text input information using a Large Language Model (LLM); calculating an intent vector corresponding to the user intent text; wherein the intent vector is used to represent the user intent; determining M plugin description vectors matching the intent vector from a vector database based on the intent vector; wherein the vector database includes N plugin description vectors, and M is less than N; determining M plugin metadata information corresponding to the M plugin description vectors from a text database based on the M plugin description vectors; wherein the text database includes N plugin metadata information; determining K target plugins corresponding to the user intent based on the user intent text and the M plugin metadata information using the LLM; wherein K is less than M, and K, M, and N are positive integers; and calling the K target plugins to generate target text information matching the user intent.
[0005] In one possible implementation, determining M plugin description vectors that match the intent vector from a vector database based on the intent vector includes: determining M plugin description vectors whose similarity to the intent vector is greater than or equal to a first threshold.
[0006] In one possible implementation, determining the M plugin metadata information corresponding to the M plugin description vectors from the text database based on the M plugin description vectors includes: determining the M plugin metadata information corresponding to the M plugin description vectors based on a first mapping relationship; wherein, the first mapping relationship includes the mapping relationship between each plugin description vector and each plugin metadata information.
[0007] In one possible implementation, determining the K target plugins corresponding to the user intent based on the user intent text and the M plugin metadata information via the LLM includes: filling the plugin description information from the user intent text and the M plugin metadata information into a first prompt template to obtain a first prompt. The K target plugins corresponding to the user intent are then determined via the LLM based on the first prompt.
[0008] In one possible implementation, the step of invoking the K target plugins to generate target text information matching the user intent includes: filling the text input information and request parameters from the metadata information of the M plugins into a second prompt template to obtain a second prompt; extracting input parameters from the text input information based on the second prompt using the LLM; and invoking the K target plugins based on the input parameters to generate target text information matching the user intent.
[0009] In one possible implementation, determining the user intent text from the text input information using a Large Language Model (LLM) includes: filling the text input information into a third prompt template to obtain a third prompt; and determining the user intent text from the text input information using the LLM based on the third prompt.
[0010] In one possible implementation, the LLM is ChatGPT-3, ChatGPT-4, BERT, or XLNet.
[0011] In one possible implementation, before receiving the text input information, the method further includes: receiving input voice information and converting the voice information into the text input information; after receiving the text input information, the method further includes: displaying the text input information; after invoking the K target plugins to generate target text information matching the user intent, the method further includes: displaying the target text information.
[0012] In one possible implementation, the plugin metadata information includes one or more of the following: plugin description information, plugin's Uniform Resource Locator (URL), plugin's request parameters, and plugin's response parameters.
[0013] In one possible implementation, the K target plugins include one or more of the following: weather query plugin, map plugin, train / high-speed rail ticket booking plugin, movie ticket booking plugin, and hotel query plugin.
[0014] In a second aspect, this application provides an electronic device, including: one or more processors and one or more memories; the one or more memories are coupled to the one or more processors, and the one or more memories are used to store a computer-executable program, which, when executed by the one or more processors, causes the electronic device to perform a method as described in any of the possible implementations of the first aspect above.
[0015] Thirdly, this application provides a chip system including a processing circuit and an interface circuit, wherein the interface circuit is used to receive code instructions and transmit them to the processing circuit, and the processing circuit is used to execute the code instructions to cause the chip system to perform the method as described in any of the possible implementations of the first aspect above.
[0016] Fourthly, this application provides a computer-readable storage medium storing a computer-executable program that, when run on an electronic device, causes the electronic device to perform the method as described in any of the possible implementations of the first aspect above. Attached Figure Description
[0017] Figures 1A-1D A set of user interface diagrams provided for embodiments of this application;
[0018] Figure 2 A schematic diagram of a voice interaction process provided in an embodiment of this application;
[0019] Figure 3A A software architecture for use in electronic devices is provided as an embodiment of this application;
[0020] Figure 3B A data interaction diagram provided for an embodiment of this application;
[0021] Figure 4 This is a schematic diagram illustrating a specific process of a voice interaction method provided in an embodiment of this application;
[0022] Figure 5A A schematic diagram of a user intent recognition process provided in an embodiment of this application;
[0023] Figure 5B A schematic diagram illustrating an example of user intent recognition provided in an embodiment of this application;
[0024] Figure 5C This is a schematic diagram of a plug-in query provided in an embodiment of this application;
[0025] Figure 5D A schematic diagram illustrating a plug-in query example provided in an embodiment of this application;
[0026] Figure 5E A schematic diagram illustrating another example of plug-in query provided in this application embodiment;
[0027] Figure 5F A schematic diagram illustrating a plugin invocation method provided in an embodiment of this application;
[0028] Figure 5G A schematic diagram illustrating a plug-in registration method provided in an embodiment of this application;
[0029] Figure 6 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0030] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to including one or more of the listed prominent features, any or all possible combinations. In the embodiments of this application, the terms “first” and “second” are used for descriptive purposes only and should not be construed as implying relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as “first” or “second” may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, “a plurality” means two or more.
[0031] To better describe the technical solutions provided in the embodiments of this application, some terms involved in the embodiments of this application will be explained first:
[0032] 1. Large Language Model (LLM): This refers to a class of deep learning-based neural network models used for learning and generating natural language text. It can be used for tasks such as language understanding, text classification, machine translation, and dialogue generation. In this application, the large language model can be ChatGPT-3, ChatGPT-4, BERT, XLNet, etc. These large language models can have hundreds of millions or billions of parameters and are capable of handling longer, more complex, and more abstract natural language tasks. Due to their massive model scale and powerful generalization ability, LLMs have significant application value in the field of natural language processing.
[0033] 2. Voice Assistant: This refers to an intelligent voice interaction system based on artificial intelligence technology. It can interact with users through voice to execute user commands and help users complete various daily tasks, such as checking the weather, playing music, sending text messages, and setting alarms. Voice assistants can be built into electronic devices such as smartphones, smart speakers, and smartwatches. Their working principle is to recognize the user's voice input as text, then use natural language processing technology to analyze the text, understand the user's intent, and provide corresponding services based on the user's needs.
[0034] 3. Prompt: This refers to the textual information used to describe a task. It typically includes the text input information converted from speech information and relevant descriptive information about the task that the large language model needs to perform. The prompt can guide the large language model to output the expected result.
[0035] 4. Plugin: This refers to an application extension. Plugins can add new functionality to an application without changing the existing application source code. Plugins can interact with the corresponding application to provide specific functions to the modules (such as LLMs) that call the plugin.
[0036] 5. Plugin Metadata Information: A collection of plugin-related information, which may include one or more of the following: plugin description information, plugin's universal resource locator (URL), plugin's request parameters, plugin's response parameters, etc. Among these, the plugin's request parameters indicate the type of input parameters given to the plugin, and the plugin's response parameters (also called response information) indicate the type of target data information returned by the plugin.
[0037] 6. Plugin description information: This can be used to indicate the plugin's functions and attributes.
[0038] 7. Intent Vector: This can be a set of floating-point numbers used to represent the corresponding user intent. For example, if the user intent is "check tomorrow's weather", the corresponding intent vector can be [0.792, -0.177, -0.107, 0.109, -0.542, ...].
[0039] 8. Plugin Description Vector: This can be a set of floating-point numbers used to represent the corresponding plugin description information. For example, if the plugin description information is "Weather query API, used for querying weather conditions, can query meteorological, temperature, humidity, perceived temperature, pressure and other information at a specified time and location", then the corresponding plugin description vector can be [-0.010028, -0.0229...].
[0040] In some application scenarios, electronic devices (which can be referred to as electronic device 100) can integrate LLM (Limited Language Management) into their internal software modules. When a user interacts with the electronic device via voice, the device can receive the user's voice input and convert it into corresponding text input. The electronic device can then use the LLM to determine the user's intent from the text input (e.g., querying weather information, querying hotel information, etc.). The electronic device can then use the LLM to execute the task corresponding to the user's intent (e.g., displaying weather information, displaying hotel information, etc.).
[0041] Figures 1A-1D This application provides a set of user interfaces for electronic devices to use LLM (Limited Language Management) and interact with users via voice.
[0042] refer to Figure 1A-Figure 1B The electronic device 100 can receive a user's command to activate a voice assistant. In response to this command, the electronic device 100 can activate the voice assistant and interact with the user via voice. This voice assistant can integrate the aforementioned LLM (Local Management Model). For example, such as... Figure 1A As shown, the electronic device 100 can receive the user's press of a physical button (such as...). Figure 1A The touch operation of the side power button (as shown) allows the electronic device 100 to activate a voice assistant and display a user interface 1000 in response to the touch operation. Alternatively, as exemplified, such as Figure 1B As shown, the electronic device 100 can receive a voice wake-up word input by the user (such as...). Figure 1B As shown in the image, “Hello, YOYO”, in response to this voice wake-up word, the electronic device 100 can activate the voice assistant and display the user interface 1000. The user interface 1000 can display a prompt icon 1000A for the voice assistant, which indicates to the user that the electronic device 100 can perform functions for voice interaction with the user.
[0043] refer to Figure 1C When electronic device 100 activates the voice assistant and displays user interface 1000, electronic device 100 can interact with the user via voice through the voice assistant. For example... Figure 1C As shown, the electronic device 100 can receive the user's voice input "Check the weather in Shenzhen tomorrow" and convert the voice input into corresponding text input. The electronic device 100 can then display the text input on the user interface 1000.
[0044] refer to Figure 1D Electronic device 100 can determine the user's intent as "check tomorrow's weather in Shenzhen" from the text input "check tomorrow's weather in Shenzhen" through a voice assistant. Then, electronic device 100 can provide the user with a weather query service via the voice assistant and display the query results on the user interface 1000. For example... Figure 1D As shown, the electronic device 100 can display the query result (also known as the target text information) "Shenzhen, tomorrow's weather is sunny, 26 degrees Celsius" on the user interface 1000.
[0045] In the embodiments of this application, Figures 1A-1D The user interface shown is merely for illustrative purposes and does not constitute any limitation on this application.
[0046] Figure 2 This is an LLM-based voice interaction process provided in an embodiment of this application.
[0047] like Figure 2 As shown, the voice interaction process may include:
[0048] S201: Electronic device 100 is pre-configured with callable application plugins (which may be referred to as plugins).
[0049] In this embodiment of the application, the electronic device 100 can receive a user's selection operation for one or more plugins. In response to the operation, the electronic device 100 sets one or more plugins pre-selected by the user as callable plugins; or, the electronic device 100 can also pre-set callable plugins through the system.
[0050] S202: Electronic device 100 receives voice information input by the user and converts the voice information into corresponding text input information. The text input information may include the user's intent.
[0051] In this embodiment of the application, when the electronic device 100 activates a voice assistant and interacts with the user via voice, the electronic device 100 can receive the user's voice input through a microphone, for example, Figure 1CAs shown, electronic device 100 can receive the user's voice input, "Check the weather in Shenzhen tomorrow," via a microphone. Then, electronic device 100 can convert the received voice information into corresponding text input. This voice assistant can integrate an LLM (Local Language Management) module; a description of LLM can be found in the preceding explanation.
[0052] S203: Electronic device 100 constructs a prompt message based on text input information and plugin description information of callable plugins.
[0053] In this embodiment of the application, the prompt information may include: text input information, plugin description information for plugins that can be invoked, and other instructions.
[0054] For example, if electronic device 100 has 10 pre-configured callable plugins, it can pre-store plugin description information for these 10 plugins. When constructing a prompt message, the prompt message includes the plugin description information for these 10 plugins. That is, the prompt message includes the plugin description information for all callable plugins.
[0055] S204: Electronic device 100 determines whether a plugin needs to be invoked based on prompt information via LLM.
[0056] S205: When it is determined that a plugin needs to be invoked, the electronic device 100 invokes the plugin through LLM to generate and output target text information that matches the user's intent.
[0057] In this embodiment of the application, different user intents can correspond to different plugins. For example, if the user intent is to query the weather, the plugin corresponding to this user intent is a weather application plugin; if the user intent is to query an address, the plugin corresponding to this user intent is a map application plugin.
[0058] Specifically, when electronic device 100 determines that a plugin needs to be invoked, it can invoke the plugin corresponding to the user's intent in the text input information based on the prompt information. This plugin can interact with the corresponding third-party application (e.g., a weather application plugin interacts with a weather application, a map application plugin interacts with a map application, etc.) to obtain target data information matching the user's intent. Then, the plugin can return this target data information to the LLM (Local Management Module). The LLM can then convert and output the target data information as corresponding target text information that matches the user's intent. For example, the aforementioned target text information can be displayed on the user interface of electronic device 100, such as... Figure 1D As shown, the text "Shenzhen, sunny tomorrow, 26 degrees Celsius" displayed on the user interface 1000 is the target text information that matches the user's intent.
[0059] S206: When it is determined that no plugin needs to be invoked, the electronic device 100 outputs target text information that matches the user's intent via LLM.
[0060] In this embodiment, the electronic device 100 determines that it does not need to call the plugin, that is, the electronic device 100 does not need to generate target text information matching the user's intent based on the target data information returned by the plugin. Therefore, the electronic device 100 can directly output target text information matching the user's intent through LLM.
[0061] However, Figure 2 The voice interaction process shown has the following problems: 1) Since the electronic device 100 can integrate a large number of plugins, during voice interaction, the user needs to manually select the plugins that can be called or the system needs to preset the plugins that can be called. Then, the electronic device 100 adds the plugin description information of the selected plugins to the prompt. It can be seen that this voice interaction process is not flexible enough in terms of plugin selection and the operation is extremely cumbersome. 2) Due to the input length limitation of LLM, if the text input information is too long, the plugin description information that can be input to the LLM will be relatively short, which may lead to the problem of incomplete input plugin description information. 3) Since different LLMs have different capabilities, and the characteristics of third-party applications and their plugins in different business areas are different, the prompt also needs to be adjusted according to the type of LLM and the type of plugin. This will lead to the waste of memory resources of the electronic device 100 and low voice interaction efficiency.
[0062] Therefore, this application provides a voice interaction method applicable to an electronic device 100. This method includes: the electronic device 100 receiving text input information, which may be text information converted from user-input voice information. Then, the electronic device 100 uses a Large Language Model (LLM) to determine the user's intent text from the text input information. The electronic device 100 calculates the intent vector corresponding to the user's intent text. Based on the intent vector, the electronic device 100 determines M plugin description vectors matching the intent vector from a vector database. The vector database may include N plugin description vectors, where M is less than N. Based on the M plugin description vectors, the electronic device 100 determines M plugin metadata information corresponding to the M plugin description vectors from a text database. The text database may include N plugin metadata information. The electronic device 100 uses LLM to determine K target plugins corresponding to the user's intent text based on the user's intent text and the plugin description information in the M plugin metadata information. Here, K is less than or equal to M, and K, M, and N are positive integers. Next, the electronic device 100 uses LLM to call the K target plugins to generate target text information matching the user's intent.
[0063] Among them, user intent text represents user intent in text form, while intent vector represents user intent in the form of a set of floating-point numbers.
[0064] By implementing the voice interaction method provided in this application, the electronic device 100 can filter out callable plugins through intent vectors and plugin description vectors, compared to the aforementioned Figure 2 The system's pre-defined or user-selected plugins approach is more convenient, allows for more accurate plugin selection, and is more efficient. Furthermore, the electronic device 100 can avoid inputting irrelevant plugin descriptions, enabling it to provide services more efficiently and accurately based on user intent.
[0065] Figure 3A This application provides a software architecture for use in an electronic device 100.
[0066] like Figure 3A As shown, this software architecture (also known as an LLM plugin integration system) can be set up in the application layer of the electronic device 100. This LLM plugin integration system may include: a Large Language Model (LLM), a prompt processing module, a vector calculation module, a data storage module, a plugin management module, and a plugin selection module. Among them:
[0067] Large Language Model (LLM) can be used for:
[0068] 1) Upon receiving voice input from the user, convert the voice input into corresponding text input and determine the user's intent text based on the text input.
[0069] 2) It can also be used to determine the target plugin corresponding to the user intent based on the user intent text and the plugin description information in the plugin metadata information (e.g., the aforementioned K target plugins). The target plugin is an extension of the target application and can interact with the target application.
[0070] 3) It can also be used to extract the input parameters that need to be given to the target plugin when calling the target plugin.
[0071] 4) It can also be used to receive target data information returned after interaction between the target plugin and the target application, and convert the target data information into target text information that matches the user's intent, etc. Specific implementation methods and other functions can be found in subsequent embodiments.
[0072] The prompt handling module can be used for:
[0073] 1) Store different types of prompt templates. These templates can include one or more of the following: intent recognition prompt templates, plugin ranking prompt templates, input parameter extraction prompt templates, etc. This allows plugin invocation tasks from different business domains to directly use these prompt templates without requiring specific modifications, thus improving plugin invocation efficiency.
[0074] 2) Generate complete prompts based on various types of prompt templates. These complete prompts can be input into a Large Language Model (LLM) to obtain specified output results. The complete prompts may include one or more of the following: intent recognition prompts, plugin ranking prompts, input parameter extraction prompts, etc. Intent recognition prompts are used to determine the user's intent text; plugin ranking prompts are used to determine the plugin corresponding to the user's intent; and input parameter extraction prompts are used to extract the input parameters to be input to the plugin from the text input information. Specific implementation methods and other functions can be found in subsequent embodiments.
[0075] The Embedding Generator module can be used for:
[0076] 1) Generate corresponding semantic vectors based on the original text information.
[0077] In this embodiment, the vector calculation module can generate a corresponding intent vector based on the user's intent text and a corresponding plugin description vector based on the plugin description information. Specific implementation details can be found in subsequent embodiments.
[0078] The data storage module may include:
[0079] 1) Vector Database. The vector database can be used to store semantic vectors. In this embodiment, the vector database can store one or more plugin description vectors, etc.
[0080] 2) Text Database. The text database can be used to store text information corresponding to semantic vectors and other text information. In this embodiment, the text database can store one or more plugin metadata information, etc. The number of plugin metadata information entries in the text database is the same as the number of plugin description vectors in the vector database.
[0081] The plugin management module can be used for:
[0082] 1) When an electronic device 100 adds a new plugin, the vector calculation module calculates the corresponding plugin description vector based on the plugin description information of the new plugin and stores the plugin description vector in the vector database. Simultaneously, the plugin management module stores the plugin metadata information of the new plugin in the text database.
[0083] 2) Establish and store the mapping relationship between plugin description vectors and plugin metadata information. When a plugin description vector and a piece of plugin metadata information both correspond to the same plugin, a mapping relationship can be established between the plugin description vector and the plugin metadata information. For example, if a weather application plugin corresponds to plugin description vector A1 and plugin metadata information A2, then the plugin management module can establish and store the mapping relationship between plugin description vector A1 and plugin metadata information A2.
[0084] 3) Store and manage plugins.
[0085] 4) Invoking plugins. In some embodiments, the plugin management module may include a plugin execution proxy module, which can invoke plugins.
[0086] The plugin selection module can be used for:
[0087] 1) Query one or more plugin description vectors that match a specified intent vector, and determine the plugin metadata information corresponding to the plugin description vector that matches the intent vector based on the mapping relationship between plugin description vectors and plugin metadata information in the plugin management module. In other words, the plugin selection module can be used to query plugin metadata information that matches a specified intent vector. For specific implementation details, please refer to subsequent embodiments.
[0088] In one possible implementation, one or more modules of the LLM plugin integration system can be located on a cloud server that establishes a communication connection with the electronic device 100. For example, the plugin management module, plugin selection module, vector calculation module, data storage module, Large Language Model (LLM), and prompt processing module of the LLM plugin integration system can all be located on the cloud server. The electronic device 100 can encapsulate the interfaces corresponding to each module, and the electronic device 100 can call the modules of the LLM plugin integration system on the cloud server through these interfaces. In a specific implementation, this application does not limit which modules of the LLM plugin integration system are located on the cloud server and which are located in the application layer of the electronic device 100.
[0089] Figure 3B This application provides a data stream interaction method for its embodiments.
[0090] like Figure 3BAs shown, when the electronic device 100 interacts with the user via voice and receives voice input from the user, the electronic device 100 can convert the voice input into corresponding text input. The plug-in management module can receive the text input and pass it to the prompt processing module.
[0091] Based on this text input information, the prompt processing module can return intent-recognition prompt information to the plug-in management module.
[0092] Then, the plugin management module can pass this intent recognition prompt information to the LLM. Based on its semantic understanding capabilities, the LLM determines the user intent text based on the intent recognition prompt information and returns the user intent text to the plugin management module.
[0093] The plugin management module can pass the user intent text to the vector calculation module. The vector calculation module can calculate the intent vector corresponding to the user intent text and return the intent vector to the plugin management module.
[0094] The plugin management module can pass the intent vector to the plugin selection module. The plugin selection module filters M plugin description vectors that match the intent vector from the vector database. Then, based on the mapping relationship between the plugin description vectors and plugin metadata information, it determines the M plugin metadata information corresponding to the M plugin description vectors from the text database. The plugin selection module returns these M plugin metadata information to the plugin management module.
[0095] The plugin management module passes the user intent text and plugin description information from M plugin metadata to the prompt processing module. Based on the user intent text and the M plugin description information, the prompt processing module generates a refined plugin ranking prompt and then returns this refined prompt to the plugin management module.
[0096] The plugin management module passes plugin ranking suggestions to the LLM. Based on its semantic understanding capabilities, the LLM determines the target plugin corresponding to the user's intent based on the plugin ranking suggestions and returns the name of the target plugin to the plugin management module.
[0097] The plugin management module passes the text input information and the target plugin's request parameters to the prompt processing module. Based on this text input information and the target plugin's request parameters, the prompt processing module generates an input parameter extraction prompt and returns this prompt to the plugin management module.
[0098] The plugin management module extracts the input parameter hints and passes them to the LLM. Based on its semantic understanding capabilities, the LLM extracts the input parameters from the text input based on the input parameter hints and returns the input parameters to the plugin management module.
[0099] Based on this input parameter, the plugin management module calls the target plugin through the plugin execution proxy module to obtain the target data information. Then, the plugin management module passes the target data information and text input information to the LLM. The LLM generates target text information that matches the user's intent based on the target data information and text input information, and returns the target text information to the plugin management module.
[0100] The plugin management module can output the target text information to the user.
[0101] Combination Figure 3A The embodiment shown, Figure 4 The following is a detailed flowchart of a voice interaction method provided in an embodiment of this application.
[0102] like Figure 4 As shown, the specific process of this voice interaction method may include:
[0103] S401: Electronic device 100 receives voice information input by the user via a microphone.
[0104] In this embodiment of the application, when the electronic device 100 activates a voice assistant and interacts with the user via voice, the electronic device 100 can receive the user's voice input through a microphone, for example, Figure 1C As shown, the electronic device 100 can receive the user's voice input, "Check the weather in Shenzhen tomorrow," via a microphone. This voice assistant can integrate an LLM (Local Management Model), and a description of LLM can be found in the preceding explanation.
[0105] Regarding the implementation method of launching the voice assistant on electronic device 100, please refer to the aforementioned... Figure 1A-Figure 1B The descriptions in the illustrated embodiments will not be repeated here.
[0106] S402: Electronic device 100 converts voice information into corresponding text input information.
[0107] In some embodiments, in addition to converting received voice information into corresponding text input information, the electronic device 100 can also receive touch operations performed by the user on a virtual keyboard / physical keyboard to obtain text input information. That is to say, this application does not limit the method of obtaining text input information.
[0108] S403: Electronic device 100 generates intent recognition prompt information based on intent recognition prompt information template and text input information.
[0109] In this embodiment of the application, the electronic device 100 can fill the text input information into the intent recognition prompt information template through the prompt information processing module to obtain the intent recognition prompt information.
[0110] S404: Electronic device 100 determines the user's intent text based on intent recognition prompt information through LLM.
[0111] In this embodiment of the application, when the electronic device 100 generates intent recognition prompt information through the prompt information processing module, the electronic device 100 can input the intent recognition prompt information into the LLM. After receiving the intent recognition prompt information, the LLM can determine the user's intent text based on its semantic understanding capability.
[0112] S405: Electronic device 100 calculates the intent vector corresponding to the user's intent text.
[0113] In this embodiment, the electronic device 100 can input the determined user intent text into the vector calculation module. The vector calculation module can calculate the intent vector corresponding to the user intent text. For a description of the intent vector, please refer to the description in the foregoing embodiments.
[0114] S406: Electronic device 100 determines M plug-in description vectors that match the intent vector.
[0115] In this embodiment, the electronic device 100 can input an intent vector to a plugin selection module, and then the plugin selection module can determine M plugin description vectors from a vector database whose similarity to the intent vector is greater than or equal to a threshold of 1. Specifically, in one implementation, the electronic device 100 can determine M plugin description vectors from the vector database whose distance difference to the intent vector is less than or equal to a threshold of 2. The vector database may include N plugin description vectors, where M is less than N.
[0116] S407: Electronic device 100 determines the M plugin metadata information corresponding to the M plugin description vectors based on the mapping relationship between the plugin description vectors and the plugin metadata information.
[0117] In this embodiment, the plugin management module can pre-store the mapping relationship between plugin description vectors and plugin metadata information. The plugin selection module can determine the M plugin metadata information corresponding to the M plugin description vectors from the text database based on the mapping relationship in the plugin management module. The text database may include N pieces of plugin metadata information.
[0118] S408: Electronic device 100 determines K target plugins corresponding to the user intent based on the user intent text and M plugin metadata information through LLM.
[0119] In this embodiment, the electronic device can input the user intent text and plugin description information from M plugin metadata information into the prompt processing module. Then, the prompt processing module fills the user intent text and the M plugin description information into the plugin ranking prompt template to obtain the plugin ranking prompt information. Next, the prompt processing module can pass this plugin ranking prompt information to the LLM. Based on its semantic understanding capabilities, the LMM determines K target plugins corresponding to the user intent based on the plugin ranking prompt information. Here, K can take values of 1, 2, 3, etc., and the specific value is not limited. K is less than or equal to M, and K, M, and N are positive integers.
[0120] S409: Electronic device 100 calls K target plugins to generate target text information that matches the user's intent.
[0121] In this embodiment, the electronic device 100 can pass text input information and request parameters for K target plugins to a prompt processing module. This prompt processing module can fill the text input information and the request parameters of the K target plugins into an input parameter extraction prompt template to obtain input parameter extraction prompt information. Then, the prompt processing module can pass this input parameter extraction prompt information to an LLM (Limited Language Management Module). The LLM, based on its semantic understanding capabilities, extracts the input parameters corresponding to each of the K target plugins from the text input information and sends these input parameters to a plugin management module. The plugin management module can then invoke the K target plugins based on these input parameters and receive the target data information returned by the K target plugins. The plugin management module can then use the LLM to convert the target data information into target text information that matches the user's intent and output this target text information to the user.
[0122] Furthermore, regarding Figure 4 The specific implementation methods of each step in the illustrated embodiment are explained in detail.
[0123] like Figure 5A As shown, the specific implementation of steps S403 to S404 can be as follows:
[0124] a) The Plugin Manager module can receive text input information and then input the text input information into the prompt information processing module.
[0125] b) The prompt information processing module can fill the text input information into the intent recognition prompt information template to obtain the intent recognition prompt information.
[0126] Specifically, the prompt information processing module can pre-store an intent recognition prompt information template. This intent recognition prompt information template may include: a text input information filling area, an intent recognition task description, etc.
[0127] For example, the intent recognition prompt template could be: "Statement: ${input}. Please identify the intent in the above statement. Note that you only need to answer the intent directly, and the intent expression should be as brief as possible. Multiple intents should be separated by commas." Here, "${input}" is the text input information filling area, and "Please identify the intent in the above statement. Note that you only need to answer the intent directly, and the intent expression should be as brief as possible. Multiple intents should be separated by commas" is the intent recognition task description.
[0128] The prompt information processing module can fill the text input information into the text input information filling area of the intent recognition prompt information template to obtain the intent recognition prompt information.
[0129] For example, if the text input is "Hello GPT, please help me check the weather in Nanjing today," when the prompt information processing module receives this text input, it can fill the text input area in the intent recognition prompt information template with this text input. The resulting intent recognition prompt information will be: "Statement: Hello GPT, please help me check the weather in Nanjing today. Please identify the intent in the above statement. Note that you only need to directly answer the intent, and the intent expression should be as brief as possible. Multiple intents should be separated by commas."
[0130] c) The prompt information processing module inputs the generated intent recognition prompt information into the LLM through the plugin management module (PluginManager).
[0131] d) LLM determines the user's intent text based on intent recognition prompts and semantic understanding capabilities.
[0132] e). LLM returns the user intent text to the Plugin Manager.
[0133] In this application, the LLM can input user intent text into the Plugin Manager in the form of a list, which can be called an intent list. However, the LLM is not limited to lists; it can also input user intent text into the Plugin Manager in other forms, and this application does not impose any restrictions on this.
[0134] For example, if the text input is "Hello GPT, please help me check the weather in Nanjing today", the specific code implementation can be as follows:
[0135] input = "Hello GPT, could you please check the weather in Nanjing today?"
[0136] print = ("Original input:", input, "\n")
[0137] intents_list=intents_recognition(input)
[0138] print(“intents_list”, intents_list, “\n”)
[0139] In other words, the original text input information is: "Hello GPT, please help me check the weather in Nanjing today." The generated intent recognition prompt is: "Statement: Hello GPT, please help me check the weather in Nanjing today. Please identify the intent in the above statement. Note that you only need to directly answer the intent, and the intent expression should be as concise as possible. Multiple intents should be separated by commas." The user intent text determined by LLM is "Check Nanjing weather." This is not limited to the text input information in the above example; other text input information can also be used in specific implementations, and this application does not impose any restrictions on this.
[0140] For example, such as Figure 5B As shown, if the text input is "I'm in Xinjiekou and want to eat spicy hot pot for around 20 yuan. Please help me find spicy hot pot delivery options nearby. Also, please help me find hotels within 1 kilometer of Nanjing Glory that cost between 300 and 500 yuan," then the intent recognition prompt generated by the prompt information processing module based on this text input will be: "Statement: I'm in Xinjiekou and want to eat spicy hot pot for around 20 yuan. Please help me find spicy hot pot delivery options nearby. Also, please help me find hotels within 1 kilometer of Nanjing Glory that cost between 300 and 500 yuan. Please identify the intent in the above statement. Note that you only need to directly answer the intent, and the intent expression should be as concise as possible. Multiple intents should be separated by commas." The user intent text determined by LLM can be: "Search for spicy hot pot delivery options, search for hotels within 1 kilometer of Nanjing Glory that cost between 300 and 500 yuan."
[0141] like Figure 5C As shown, the specific implementation of steps S406 to S408 can be as follows:
[0142] Step S406:
[0143] a) The Plugin Manager module inputs user intent text into the vector calculation module.
[0144] b) The vector calculation module calculates the intent vector corresponding to the user's intent text.
[0145] c) The vector calculation module inputs the intent vector to the plugin selection module through the plugin management module.
[0146] The intent vector can be a set of floating-point numbers used to represent the corresponding user intent.
[0147] d) When an intent vector is received, the plug-in selection module can select from the vector database ( Figure 5C (not shown in the image) Identify M plugin description vectors that match the intent vector.
[0148] In one implementation, when an intent vector is received, the plugin selection module can calculate the distance difference between each plugin description vector in the vector database and the intent vector. Then, the plugin selection module can determine M plugin description vectors from the vector database whose distance difference with the intent vector is less than or equal to a threshold of 2. These M plugin description vectors are the plugin description vectors that match the intent vector. The threshold of 2 can be 0.1, 0.3, 0.2, etc., or other values; this application does not impose any restrictions on this.
[0149] Step S407:
[0150] e). The plugin selection module can select from a text database based on the mapping relationship between plugin description vectors and plugin metadata information in the plugin management module. Figure 5C The M plugin metadata information corresponding to the above M plugin description vectors (not shown in the figure) is determined.
[0151] In this embodiment of the application, the text database can store N plugin metadata information, the vector database can store N plugin description vectors, and the plugin management module can store the mapping relationship between each plugin description vector and the plugin metadata information.
[0152] For example, the mapping relationship between each plugin description vector and plugin metadata information can be shown in Table 1:
[0153] Table 1
[0154] Plugin description vector Plugin metadata information Plugin Description Vector 1 Plugin Metadata Information 1 Plugin Description Vector 2 Plugin metadata information 2 Plugin Description Vector 3 Plugin metadata information 3 …… ……
[0155] As shown in Table 1, plugin description vector 1 can correspond to plugin metadata information 1, plugin description vector 2 can correspond to plugin metadata information 2, plugin description vector 3 can correspond to plugin metadata information 3, and so on. Table 1 is only used as an example to explain this application and does not constitute any limitation on this application.
[0156] For example, the code implementation for querying plugin metadata information could be:
[0157] i = 0;
[0158] plugin_info_dict=plugin_search(intents_list[i]);
[0159] For example, when the user's intent is "to check the weather in Nanjing", the retrieved plugin metadata information could be:
[0160] Plugin Description: Weather Query API, used to query weather conditions. It can query information such as weather, temperature, humidity, perceived temperature, and pressure for a specified time and location.
[0161] Plug-in url: http: / / ${IP}:${port} / aicloud / connector / v1 / fulfillment / weather / query
[0162] Plugin parameter information: City
[0163] Plugin request parameters: city
[0164] Plugin response information: weatherid, temperature, humidity, pressure, updatetime, realfeel, mobileLink
[0165] Step S408:
[0166] f) The plugin selection module sends the metadata information of the selected M plugins to the plugin management module.
[0167] g) The plugin management module will send the plugin description information and user intent text input from the M received plugin metadata information to the prompt information processing module.
[0168] The plugin management module can input plugin description information to the prompt information processing module in a list format. However, it is not limited to a list; the plugin management module can also input plugin description information to the prompt information processing module in other formats, and this application does not impose any restrictions on this.
[0169] h) The prompt information processing module fills the M plugin description information and user intent text into the plugin fine-ranking prompt information template to generate plugin fine-ranking prompt information.
[0170] The prompt information processing module can pre-store a plugin ranking prompt information template. This template may include: a user intent text fill area, a plugin description information fill area, and a plugin query task description, etc.
[0171] For example, the plugin ranking prompt template could be: "Intents list: ${intents_list}. For these intents, the system can provide some auxiliary plugin APIs, with names and capability descriptions as follows: ${plugin_description_list}. Please return the plugin APIs that should be used for each intent in sequence. Note: Only the plugin API names need to be returned, separated by commas. Intents without matching plugins should be represented by null." Here, "${intents_list}" is the user intent text population area, "${plugin_description_list}" is the plugin description information population area, and "Please return the plugin APIs that should be used for each intent in sequence. Note: Only the plugin API names need to be returned, separated by commas. Intents without matching plugins should be represented by null." is the plugin query task description.
[0172] The prompt message processing module can fill the user intent text into the user intent text filling area of the plugin fine-mapping prompt message template, and fill the plugin description information into the plugin description information filling area to obtain the plugin fine-mapping prompt message.
[0173] For example, if the user intent text is "Check Nanjing weather" or "Order takeout," and the plugin description information is "Weather query API, used for querying weather conditions, can query meteorological, temperature, humidity, perceived temperature, pressure, etc., at a specified time and location," or "Train ticket and high-speed rail ticket query API, used for querying train, high-speed rail, and railway ticket information such as train number, departure time, or date, providing railway operation information such as station, departure time, arrival time, and layover time," when the prompt information processing module receives the user intent text and plugin description information, the prompt information processing module can fill the user intent text into the user intent text filling area of the plugin layout prompt information template, and fill the plugin description information into the plugin description information of the plugin layout prompt information template. In the information-filled area, the resulting plugin ranking prompts are: "Intent List: Query Nanjing weather, order takeout. For these intents, the system provides several auxiliary plugin APIs with the following names and capabilities: Weather query API, used for querying weather conditions, including weather, temperature, humidity, perceived temperature, and pressure at a specified time and location. Train ticket and high-speed rail ticket query API, used for querying train, high-speed rail, and railway ticket information such as train number, departure time, or date, providing railway operation information such as station, departure time, arrival time, and layover time. Please return the plugin APIs that should be used for each intent in sequence. Note: Only the plugin API name needs to be returned; multiple names should be separated by commas. Use null to indicate if no intent plugin meets the requirements."
[0174] i) The prompt message processing module will input the generated plugin refinement prompt messages into the LLM through the plugin management module.
[0175] j) LLM uses plugin ranking prompts to determine the K target plugins corresponding to the user's intent based on semantic understanding capabilities.
[0176] Since there is a correspondence between user intent and intent vector, the K target plugins corresponding to the user intent are also the plugins corresponding to the intent vector of the user intent.
[0177] For example, such as Figure 5DAs shown, if the plugin ranking prompt message is "Intent List: Query Nanjing weather, order takeout. For these intents, the system can provide some auxiliary plugin APIs, whose names and capabilities are described as follows: Weather query API, used for querying weather conditions, can query information such as weather, temperature, humidity, perceived temperature, and pressure at a specified time and location. Train ticket and high-speed rail ticket query API, used for querying ticket information such as train and high-speed rail numbers, departure times, or dates, providing railway operation information such as station, departure time, arrival time, and stay time. Please return the plugin APIs that should be used for each intent in sequence. Note: Only the plugin API name needs to be returned, multiple names are separated by commas, and no intent plugin meets the requirements is represented by null.", then the target plugin determined by LLM can be "Weather query API, null".
[0178] For example, such as Figure 5E As shown, if the plugin ranking prompt message is "Intent List: Query Nanjing weather, book high-speed rail. For these intents, the system can provide some auxiliary plugin APIs, whose names and capabilities are described as follows: Weather query API, used for querying weather conditions, can query information such as weather, temperature, humidity, perceived temperature, and pressure at a specified time and location. Train ticket and high-speed rail ticket query API, used for querying ticketing information such as train and high-speed rail numbers, departure times, or dates, providing railway operation information such as station, departure time, arrival time, and layover time. Please return the plugin APIs that should be used for each intent in sequence. Note: Only the plugin API name needs to be returned, multiple names are separated by commas, and no intent plugin meets the requirements is represented by null.", then the target plugins determined by LLM can be "Weather query API, Train ticket and high-speed rail ticket query API".
[0179] k).LLM returns the names of the K identified target plugins to the plugin management module.
[0180] The LLM can return the names of K target plugins to the plugin management module in the form of a list. However, it is not limited to a list; the LLM can also return the names of the K target plugins to the plugin management module in other forms, and this application does not impose any restrictions on this.
[0181] like Figure 5F As shown, the specific implementation of step S409 can be as follows:
[0182] a) The plugin management module can input text input information and request parameters from the plugin metadata information of K target plugins (which can be simply referred to as the request parameters of K target plugins) to the prompt information processing module.
[0183] b) The prompt information processing module can fill the input parameter extraction prompt information template with the text input information and the request parameters of K target plugins to obtain the input parameter extraction prompt information.
[0184] Specifically, the prompt information processing module can pre-store input parameter extraction prompt information templates. These templates may include one or more of the following: a text input information fill area, a request parameter fill area, and an input parameter extraction task description.
[0185] For example, the input parameter extraction prompt template could be: "Statement: ${input}. Please extract the following information from the above statement: ${keys}. Return results are comma-separated, and unextracted information is represented by null." Here, "${input}" can be referred to as the text input information filling area, "${keys}" can be referred to as the request parameter filling area, and "Please extract the following information from the above statement: ${keys}. Return results are comma-separated, and unextracted information is represented by null" can be referred to as the input parameter extraction task description.
[0186] The prompt information processing module can fill the text input information into the text input information filling area of the input parameter extraction prompt information template, and fill the request parameters of K target plugins into the request parameter filling area of the input parameter extraction prompt information template to obtain the input parameter extraction prompt information.
[0187] For example, if the text input is "Hello gpt, please help me check the weather in Nanjing today," and the target plugin's request parameter is "city," then the input parameter extraction prompt could be: "Statement: Hello gpt, please help me check the weather in Nanjing today. Please extract the following information from the above statement: city. Return results are comma-separated, and unextracted information is represented by null."
[0188] c) The prompt message processing module can extract prompt messages from input parameters and input them to the LLM through the plug-in management module.
[0189] d) LLM extracts prompt information based on input parameters, extracting input parameters from text input information based on semantic understanding capabilities.
[0190] For example, the code implementation of this step can be as follows:
[0191] request_body=plugin_params_extraction(input, plugin_info_dict['params'], plugin_info_dict['request'])
[0192] For example, if the input parameter extraction prompt is "Statement: Hello GPT, please help me check the weather in Nanjing today. Please extract the following information from the above statement: city. Return results are separated by commas, and information not extracted is represented by null.", then the input parameter extracted by LLM from the text input information "Hello GPT, please help me check the weather in Nanjing today" is "Nanjing".
[0193] e). LLM returns the input parameters to the plugin management module.
[0194] f) The plugin management module inputs the plugin metadata information and input parameters of the K target plugins to the plugin execution agent module, which then calls the K target plugins.
[0195] Specifically, when the plugin execution proxy module calls K target plugins, the plugin execution proxy module can input the input parameters to the target plugins.
[0196] g) The plugin management module receives target data information returned by K target plugins.
[0197] For example, the code implementation for calling the target plugin to obtain target data information can be:
[0198] plugin_call(plugin_info_dict['url'], request_body).
[0199] The plugin execution proxy module can invoke the target plugin based on its URL.
[0200] For example, if the text input is "Hello GPT, please help me check the weather in Nanjing today," and the extracted input parameter is "Nanjing," and the user intent is "check the weather in Nanjing," then the plugin called is the weather query API. The plugin's execution proxy module will then call this weather query API based on the following input parameters and URL:
[0201] The URL for the weather query API is: http: / / ${IP}:${port} / aicloud / connector / v1 / fulfillment / weather / query
[0202] body: {"city": "Nanjing"}
[0203] The target data information returned by the weather query API can be:
[0204] Plugin call result: {'weatherid': 'Cloudy', 'temperature': '28.0', 'humidity': '74', 'pressure': '996.0', 'updatetime': '1688464469000', 'realfeel': '32.0', 'mobileLink': 'https: / / weather.com / zh / cn'}
[0205] h) The plugin management module inputs the target data information and text input information into the LLM.
[0206] In this embodiment, since the target data information returned by the target plugin is the response body of the interface, which is generally in JSON structure, it needs to be rewritten by LLM into a conversational expression that conforms to human expression (i.e., target text information). The plugin management module inputs the text input information and the target data information together into the LLM, which can enhance the effect of the LLM in converting the target data information into target text information.
[0207] i) LLM converts target data information and text input information into target text information that matches the user's intent.
[0208] j) LLM outputs target text information to the user through the plugin management module.
[0209] For example, such as Figure 1D As shown, the electronic device 100 can display target text information in a user interface. Alternatively, the electronic device 100 can also verbally display the target text information. That is to say, this application does not limit the method by which the electronic device 100 outputs the target text information.
[0210] In the embodiments of this application, such as Figure 5G As shown, the implementation method for adding a specified plugin to electronic device 100 can be as follows:
[0211] a) The plugin management module obtains the plugin metadata information for the specified plugin.
[0212] The explanation of plugin metadata information can be found in the aforementioned embodiments and will not be repeated here.
[0213] Specifically, when the electronic device 100 adds a specified plugin, the electronic device 100 can construct plugin metadata information from the obtained plugin description information, the URL of the specified plugin, the request parameters of the specified plugin, and the response parameters of the specified plugin.
[0214] b) The plugin management module inputs the plugin description information of the specified plugin into the vector calculation module.
[0215] c) The vector calculation module can calculate the plugin description vector corresponding to the plugin description information.
[0216] d) The vector calculation module can return the plugin description vector to the plugin management module.
[0217] e) The plugin management module sends the plugin metadata information and plugin description vector of the specified plugin to the data storage module for storage.
[0218] The plugin metadata information can be stored in a text database in the data storage module, and the plugin description vector can be stored in a vector database in the data storage module.
[0219] f) The plugin management module establishes and stores the mapping relationship between the plugin metadata information and the plugin description vector of the specified plugin.
[0220] In some examples, the plugin management module can also store newly added specified plugins.
[0221] In this embodiment, threshold 1 can be referred to as the first threshold. The mapping relationship between each plugin description vector and each plugin metadata information can be included in the first mapping relationship. The plugin fine-tuning prompt information template can be referred to as the first prompt information template, and the plugin fine-tuning prompt information can be referred to as the first prompt information prompt. The input parameter extraction prompt information template can be referred to as the second prompt information template, and the input parameter extraction prompt information can be referred to as the second prompt information prompt. The intent recognition prompt information template can be referred to as the third prompt information template, and the intent recognition prompt information can be referred to as the third prompt information prompt. The K target plugins can include one or more of the following: weather query plugin, map plugin, train / high-speed rail ticket booking plugin, movie ticket booking plugin, hotel query plugin, etc.
[0222] Figure 6 The present application provides a hardware structure for an electronic device 100.
[0223] like Figure 6 As shown, the electronic device 100 may include a processor 601, a memory 602, a wireless communication module 603 (optional), a display screen 604, a camera 605, an audio module 606 (optional), and a microphone 607 (optional). The processor 601, memory 602, wireless communication module 603 (optional), display screen 604, camera 605, audio module 606 (optional), and microphone 607 (optional) may be connected via a bus.
[0224] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may further include... Figure 6 This may involve more or fewer components, or combining certain components, or splitting certain components, or different component arrangements. Figure 6 The components shown can be implemented in hardware, software, or a combination of both.
[0225] Processor 601 may include one or more processor units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.
[0226] The processor 601 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 601 is a cache memory. This memory can store instructions or data that the processor 601 has just used or that are used repeatedly. If the processor 601 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 601, and thus improves the efficiency of the system.
[0227] In some embodiments, the processor 601 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a USB interface, etc.
[0228] The memory 602 is coupled to the processor 601 and is used to store various software programs and / or multiple sets of instructions. In specific implementations, the memory 602 may include volatile memory, such as random access memory (RAM); it may also include non-volatile memory, such as ROM, flash memory, hard disk drive (HDD), or solid-state drive (SSD); the memory 602 may also include combinations of the above types of memory. The memory 602 may also store some program code so that the processor 601 can call the program code stored in the memory 602 to implement the implementation method of the present application embodiment in the electronic device 100. The memory 602 may store an operating system, such as uCOS, VxWorks, RTLinux, or other embedded operating systems.
[0229] The wireless communication module 603 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 603 can be one or more devices integrating at least one communication processing module. The wireless communication module 603 receives electromagnetic waves via an antenna, modulates and filters the electromagnetic wave signals, and sends the processed signal to the processor 601. The wireless communication module 603 can also receive signals to be transmitted from the processor 601, modulate and amplify them, and convert them into electromagnetic waves for radiation via the antenna. In some embodiments, the electronic device 100 can also communicate via the Bluetooth module in the wireless communication module 603 (… Figure 6 (not shown), WLAN module ( Figure 6 (Not shown) The device transmits signals to detect or scan devices near electronic device 100 and establishes wireless communication connections with those devices to transmit data. The Bluetooth module can provide solutions for one or more Bluetooth communication methods, including basic rate / enhanced data rate (BR / EDR) or Bluetooth Low Energy (BLE), and the WLAN module can provide solutions for one or more WLAN communication methods, including Wi-Fi direct, Wi-Fi LAN, or Wi-Fi softAP.
[0230] The display screen 604 can be used to display images, videos, etc. The display screen 604 may include a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a minimized LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 604, where N is a positive integer greater than 1.
[0231] Camera 605 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, electronic device 100 may include one or N cameras 605, where N is a positive integer greater than 1.
[0232] The audio module 606 can be used to convert digital audio information into analog audio signal output, and can also be used to convert analog audio input into digital audio signal. The audio module 606 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 606 can also be disposed in the processor 601, or some functional modules of the audio module 606 can be disposed in the processor 601.
[0233] Microphone 607, also known as a "microphone" or "voice transducer," is used to collect sound signals from the environment surrounding the electronic device. This sound signal is then converted into an electrical signal, which undergoes a series of processing steps, such as analog-to-digital conversion, to obtain a digital audio signal that can be processed by the processor 601 of the electronic device. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to the microphone 607, inputting the sound signal into the microphone 607. The electronic device 100 may have at least one microphone 607. In some embodiments, the electronic device 100 may have two microphones 607, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, the electronic device 100 may have three, four, or more microphones 607, enabling sound signal collection, noise reduction, sound source identification, and directional recording, among other functions.
[0234] Electronic device 100 may also include a sensor module ( Figure 6 (Not shown in the image). The sensor module may include multiple sensor elements, such as a touch sensor (…). Figure 6 (Not shown in the image). A touch sensor can also be called a "touch device". A touch sensor can be placed on the display screen 604, and the touch sensor and the display screen 604 together form a touch screen, also called a "touchscreen". The touch sensor can be used to detect touch operations applied to or near it.
[0235] It should be noted that, Figure 6 The electronic device 100 shown is merely an illustrative explanation of the hardware structure of the electronic device provided in this application and does not constitute a specific limitation on this application.
[0236] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as meaning "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is interpreted as meaning "if determining...", "in response to determining...", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".
[0237] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0238] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A voice interaction method, characterized in that, include: Receive text input information; The user intent text is determined from the text input information using a Large Language Model (LLM). Calculate the intent vector corresponding to the user intent text; wherein the intent vector is used to represent the user intent; Based on the intent vector, M plugin description vectors matching the intent vector are determined from the vector database; wherein the vector database includes N plugin description vectors, and M is less than N; wherein the plugin description vectors are used to indicate the functions and attributes of the plugin; Based on the M plugin description vectors, M plugin metadata information corresponding to the M plugin description vectors are determined from the text database; wherein, the text database includes N plugin metadata information; wherein, the plugin metadata information includes one or more of the following: plugin description information, plugin's Uniform Resource Locator, plugin's request parameters, and plugin's response parameters; Using the LLM, K target plugins corresponding to the user intent are determined based on the user intent text and the metadata information of the M plugins; wherein, K is less than M, and K, M and N are positive integers; The K target plugins are invoked to generate target text information that matches the user's intent.
2. The method according to claim 1, characterized in that, The step of determining M plugin description vectors that match the intent vector from the vector database based on the intent vector includes: Identify M plugin description vectors whose similarity to the intent vector is greater than or equal to a first threshold.
3. The method according to claim 1 or 2, characterized in that, The step of determining the M plugin metadata information corresponding to the M plugin description vectors from the text database based on the M plugin description vectors includes: Based on the first mapping relationship, the M plugin metadata information corresponding to the M plugin description vectors are determined; wherein, the first mapping relationship includes the mapping relationship between each plugin description vector and each plugin metadata information.
4. The method according to claim 1, characterized in that, The step of determining the K target plugins corresponding to the user intent based on the user intent text and the metadata information of the M plugins through the LLM includes: The user intent text and the plugin description information from the M plugin metadata information are filled into the first prompt template to obtain the first prompt. The LLM determines K target plugins corresponding to the user's intent based on the first prompt message.
5. The method according to claim 1, characterized in that, The step of invoking the K target plugins to generate target text information matching the user intent includes: The text input information and the request parameters from the M plugin metadata information are filled into the second prompt template to obtain the second prompt. The LLM extracts input parameters from the text input information based on the second prompt message; Based on the input parameters, the K target plugins are invoked to generate target text information that matches the user's intent.
6. The method according to claim 1, characterized in that, The step of determining the user intent text from the text input information using a Large Language Model (LLM) includes: The text input information is filled into the third prompt template to obtain the third prompt. The LLM determines the user intent text from the text input information based on the third prompt information.
7. The method according to claim 1, characterized in that, The LLM is ChatGPT-3, ChatGPT-4, BERT, or XLNet.
8. The method according to claim 1, characterized in that, Before receiving the text input information, the method further includes: Upon receiving the input voice information, the voice information is converted into the text input information; After receiving the text input information, the method further includes: Display the text input information; After invoking the K target plugins to generate target text information matching the user intent, the method further includes: The target text information is displayed.
9. The method according to claim 1, characterized in that, The plugin metadata information includes one or more of the following: plugin description information, plugin's Uniform Resource Locator (URL), plugin's request parameters, and plugin's response parameters.
10. The method according to claim 1, characterized in that, The K target plugins include one or more of the following: weather query plugin, map plugin, train / high-speed rail ticket booking plugin, movie ticket booking plugin, and hotel query plugin.
11. An electronic device, characterized in that, include: One or more processors and one or more memories; the one or more memories are coupled to the one or more processors, the one or more memories being used to store a computer-executable program, which, when executed by the one or more processors, causes the electronic device to perform the method as described in any one of claims 1-10.
12. A chip system, characterized in that, It includes a processing circuit and an interface circuit, the interface circuit being used to receive code instructions and transmit them to the processing circuit, the processing circuit being used to execute the code instructions to cause the chip system to perform the method as described in any one of claims 1-10.
13. A computer-readable storage medium, characterized in that, The device contains a computer-executable program that, when run on an electronic device, causes the electronic device to perform the method as described in any one of claims 1-10.
Citation Information
Patent Citations
Generative large language model training method and model-based man-machine voice interaction method
CN116127045A