An information interaction method and device, electronic equipment and medium
By using a large language model to perform semantic parsing of speech information, combined with the server-side command database and preset command library, the problem of interactive all-in-one machines being unable to recognize complex voice commands has been solved, enabling accurate responses to complex voice inputs and improving the intelligence of information interaction.
Patent Information
- Application Number
- CN202310944323.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-28
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2043-07-28
AI Technical Summary
Existing interactive kiosks can only recognize simple, fixed voice inputs and cannot recognize complex voice commands, resulting in an inability to accurately respond to user voice inputs.
By using a large language model to perform semantic parsing of speech information, and combining it with the server-side command database and preset command library, the system can parse and respond to complex speech commands.
It achieves accurate parsing and response to complex voice input, improves the intelligence level of information interaction, and meets a wider range of information interaction needs.
Smart Images

Figure CN116894078B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information processing technology, and in particular to an information interaction method, apparatus, electronic device, and medium. Background Technology
[0002] With the rapid development of smart all-in-one machines, their use is becoming increasingly widespread. Interactive all-in-one machines are high-tech products that integrate the functions of high-definition televisions, tablets, and interactive whiteboards, such as conference all-in-one machines and educational all-in-one machines.
[0003] In existing technologies, when implementing human-computer interaction through interactive all-in-one machines, relevant instructions are usually searched and matched based on the voice input from the client. If the instruction is successfully matched, it is executed; if the instruction fails to match, it is not executed, and a message is displayed indicating that the user's voice input cannot be recognized.
[0004] However, existing interactive all-in-one machines can only recognize fixed, simple voice inputs. For example, they can only recognize "open screen recording" but not "I want to record the screen," even though these two sentences convey essentially the same meaning. Because interactive all-in-one machines can only recognize fixed, simple voice inputs, when the voice input exceeds the fixed recognition range, the interactive all-in-one machine cannot recognize the voice input and therefore cannot match the corresponding command for an accurate response. Summary of the Invention
[0005] This invention provides an information interaction method, device, electronic device, and medium that can perform semantic parsing of voice input information based on a large language model and provide accurate and effective responses based on the semantic parsing results, making information interaction more intelligent and better meeting information interaction needs.
[0006] According to one aspect of the present invention, an information interaction method is provided, the method comprising:
[0007] Acquire the voice information to be queried received through the target input interface, and convert the voice information to be queried into text information to be queried;
[0008] The query text information is semantically parsed using a pre-deployed large language model on the server to obtain target response information, so that the server can execute a target operation based on the target response information; wherein, the target operation is to determine the target response instruction corresponding to the target response information from the instruction database pre-deployed on the server, and to determine the identification information associated with the target response instruction;
[0009] The system receives the identification information associated with the target response instruction sent by the server, and determines the target feedback instruction from a preset instruction library based on the identification information associated with the target response instruction; wherein the identification information associated with the target response instruction is the same as that associated with the target feedback instruction.
[0010] Execute the target feedback instruction and return the execution result of the target feedback instruction to the target input interface.
[0011] According to another aspect of the present invention, an information interaction device is provided, comprising:
[0012] The text information determination module is used to acquire the voice information to be queried received through the target input interface and convert the voice information to be queried into text information to be queried.
[0013] The target operation execution module is used to perform semantic parsing on the text information to be queried using a large language model pre-deployed on the server to obtain target response information, so that the server can execute a target operation based on the target response information; wherein, the target operation is to determine the target response instruction corresponding to the target response information from the instruction database pre-deployed on the server, and to determine the identification information associated with the target response instruction;
[0014] The feedback instruction determination module is used to receive the identification information associated with the target response instruction sent by the server, and determine the target feedback instruction from a preset instruction library based on the identification information associated with the target response instruction; wherein the identification information associated with the target response instruction is the same as that associated with the target feedback instruction.
[0015] The feedback instruction execution module is used to execute the target feedback instruction and return the execution result of the target feedback instruction to the target input interface.
[0016] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0017] At least one processor; and
[0018] A memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the information interaction method described in any embodiment of the present invention.
[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the information interaction method described in any embodiment of the present invention.
[0021] The technical solution of this invention involves acquiring voice information to be queried received through a target input interface, converting the voice information into text information, and then using a pre-deployed large language model on the server to perform semantic parsing on the text information to obtain target response information. This allows the server to execute a target operation based on the target response information. The target operation involves determining the target response instruction corresponding to the target response information from a pre-deployed instruction database on the server, and determining the identifier information associated with the target response instruction. The solution also involves receiving the identifier information associated with the target response instruction sent by the server, and determining a target feedback instruction from a preset instruction library based on the identifier information associated with the target response instruction. The identifier information associated with the target response instruction is the same as that associated with the target feedback instruction. Finally, the solution executes the target feedback instruction and returns the execution result to the target input interface. This technical solution enables semantic parsing of voice input information based on a large language model and provides accurate and effective responses based on the semantic parsing results, making information interaction more intelligent and better meeting information interaction needs.
[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart of an information interaction method provided according to Embodiment 1 of the present invention;
[0025] Figure 2 This is a flowchart of an information interaction method provided according to Embodiment 2 of the present invention;
[0026] Figure 3 This is a schematic diagram of an information interaction process provided according to an embodiment of the present invention;
[0027] Figure 4 This is a schematic diagram of the structure of an information interaction device according to Embodiment 3 of the present invention;
[0028] Figure 5 This is a schematic diagram of the structure of an electronic device that implements an information interaction method according to an embodiment of the present invention. Detailed Implementation
[0029] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0030] It should be noted that the terms "first," "second," "target," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0031] Example 1
[0032] Figure 1 This is a flowchart of an information interaction method provided in Embodiment 1 of the present invention. This embodiment is applicable to intelligent information interaction based on a large language model. The method can be executed by an information interaction device, which can be implemented in hardware and / or software and can be configured in an electronic device with data processing capabilities. Figure 1 As shown, the method includes:
[0033] S110: Obtain the voice information to be queried received through the target input interface, and convert the voice information to be queried into text information to be queried.
[0034] The technical solution of this embodiment can be executed by an interactive all-in-one machine, such as a conference all-in-one machine. A conference all-in-one machine integrates multiple functions such as a projector, electronic whiteboard, speakers, television, and video conferencing terminal into one unit; it is an office device specifically designed for meetings. This technical solution can perform semantic parsing of input speech information based on a large language model and provide accurate responses based on the parsing results. It is applicable to both simple and complex speech input responses, effectively avoiding the problem of inaccurate responses due to the interactive all-in-one machine's inability to correctly understand speech input. A large language model (such as GPT) is an artificial intelligence model designed to understand and generate human language. They are trained on large amounts of text data and can perform a wide range of tasks, including text summarization, translation, sentiment analysis, and more.
[0035] The target input interface can be used to receive voice information input by the user. Specifically, the target input interface can be a voice input interface set on the interactive whiteboard or on other devices, as long as the interactive whiteboard can obtain the voice input. The specific interface can be set according to actual needs, and this embodiment does not impose any limitations on it. The voice information to be queried can refer to the voice information that the user wants to query through the target input interface. The text information to be queried can refer to the text information corresponding to the voice information to be queried.
[0036] In this embodiment, after a user inputs the voice information to be queried through the target input interface, the interactive all-in-one machine can acquire the voice information received through the target input interface and convert it into text information by calling a third-party service. The third-party service can be used to convert the voice information into corresponding text information.
[0037] S120 uses a pre-deployed large language model on the server to perform semantic parsing of the text information to be queried to obtain the target response information, so that the server can execute the target operation based on the target response information.
[0038] The target response information can refer to the information obtained after semantic parsing of the query text information using a large language model. For example, assuming the query text information is "I want to record the screen," semantic parsing using a large language model yields the target response information "Open screen recording." The target operation involves determining the target response instruction corresponding to the target response information from a pre-deployed instruction database on the server, as well as determining the identification information associated with the target response instruction. The instruction database is a pre-deployed database on the server that can store different functional instructions. The target response instruction can refer to the functional instruction in the instruction database that corresponds to the target response information. The identification information can be used to uniquely represent the instruction, such as ID information.
[0039] In this embodiment, after obtaining the text information to be queried, the interactive all-in-one machine can call a pre-deployed large language model on the server side. This large language model performs semantic parsing on the text information to obtain the target response information. After obtaining the target response information, the server can search for a matching function instruction from a pre-deployed instruction database as the target response instruction, determine the associated identification information of the target response instruction, and then send the associated identification information to the interactive all-in-one machine. In this embodiment, a corresponding identification information is pre-assigned to each instruction in the instruction database, which can be used to uniquely represent the instruction.
[0040] It should be noted that, due to the limited number of instructions in the instruction database—meaning it cannot include all the functional instructions a user wants to achieve—the server may be unable to find a matching instruction for the target response information in the database. This could result in the user's input of voice information not receiving a correct response. In such cases, the corresponding text information can be stored on the server. The server can then analyze the user's needs based on this text information to expand the instructions in the database later. If no matching instruction is found, the associated identifier can be designated as a preset identifier. This preset identifier can be used to indicate that no matching instruction can be found.
[0041] S130: Receive the identification information associated with the target response instruction sent by the server, and determine the target feedback instruction from the preset instruction library based on the identification information associated with the target response instruction.
[0042] The preset instruction library refers to a pre-deployed instruction library within the interactive whiteboard, used to store function execution instructions. It's important to note that the function execution instructions in the preset instruction library of the interactive whiteboard match the function instructions in the instruction database on the server side. The target feedback instruction refers to a function execution instruction in the preset instruction library whose identifier information matches the target response instruction. It's crucial to emphasize that the identifier information associated with the target response instruction and the target feedback instruction is the same. The target response instruction and the target feedback instruction have corresponding functions; the target response instruction focuses on function description, while the target feedback instruction focuses on function implementation.
[0043] In this embodiment, when the interactive all-in-one machine receives the identification information associated with the target response instruction sent by the server, it can search from a preset instruction library based on the identification information associated with the target response instruction, and determine the function execution instruction with the same identification information as the target response instruction as the target feedback instruction. Furthermore, if the identification information associated with the target response instruction is preset identification information, it indicates that no target response instruction corresponding to the target response information can be matched. In this case, the target feedback instruction can be determined as a preset feedback instruction, which can be used to indicate that there is no feedback instruction matching the queried voice information.
[0044] S140, execute the target feedback instruction and return the execution result of the target feedback instruction to the target input interface.
[0045] In this embodiment, after determining the target feedback instruction, the target feedback instruction can be executed through the interactive all-in-one machine, and the execution result of the target feedback instruction can be returned to the target input interface, thereby realizing intelligent information interaction. The execution result may be a voice message or an action. Furthermore, if the target feedback instruction is a preset feedback instruction, a preset execution result can be obtained when the target feedback instruction is executed. The preset execution result can be used to indicate that the queried voice information cannot be recognized.
[0046] The technical solution of this invention involves acquiring voice information to be queried received through a target input interface, converting the voice information into text information, and then using a pre-deployed large language model on the server to perform semantic parsing on the text information to obtain target response information. This allows the server to execute a target operation based on the target response information. The target operation involves determining the target response instruction corresponding to the target response information from a pre-deployed instruction database on the server, and determining the identifier information associated with the target response instruction. The solution also involves receiving the identifier information associated with the target response instruction sent by the server, and determining a target feedback instruction from a preset instruction library based on the identifier information associated with the target response instruction. The identifier information associated with the target response instruction is the same as that associated with the target feedback instruction. Finally, the solution executes the target feedback instruction and returns the execution result to the target input interface. This technical solution enables semantic parsing of voice input information based on a large language model and provides accurate and effective responses based on the semantic parsing results, making information interaction more intelligent and better meeting information interaction needs.
[0047] In this embodiment, optionally, the instruction database includes function names and response instructions associated with the function names; correspondingly, determining the target response instruction corresponding to the target response information from the instruction database pre-deployed on the server includes: determining the target function name from the instruction database on the server based on the target response information; and determining the response instruction associated with the target function name from the instruction database as the target response instruction.
[0048] The target function name can refer to the function name in the instruction database that matches the target response information. In this embodiment, the instruction database may specifically include function names and response instructions associated with those function names. Therefore, when determining the target response instruction corresponding to the target response information, the target function name matching the target response information can first be determined from the server's instruction database based on the target response information, and then the response instruction associated with the target function name can be determined as the target response instruction.
[0049] This solution, through its functional division of instructions in the instruction database, first identifies the functional module and then determines the specific instruction associated with that module when determining the target response instruction corresponding to the target response information. This avoids a global search of the instruction database, effectively shortening the instruction search time and improving the efficiency of determining the target response instruction.
[0050] In this embodiment, optionally, the method further includes: if the server detects an instruction addition event, obtaining the newly added instruction information from the instruction addition event through the server; adding the new instruction to the instruction database according to the newly added instruction information, and generating identification information for the new instruction; generating a list of newly added instructions according to the new instruction and the identification information of the new instruction.
[0051] The instruction addition event can refer to an operation instruction requesting the addition of a new instruction. The new instruction information can describe information related to the instruction to be added. For example, the new instruction information may include the instruction function name and instruction content description. The new instruction list can refer to a list generated based on the new instructions and their identification information, and may include one or more sets of new instructions and their corresponding identification information.
[0052] In this embodiment, a command registry can be pre-deployed on the server side. This command registry serves as a key module for command extensibility, used to register new commands and maintain existing commands. When a new command needs to be added, a command addition event can be generated through the external API interface provided by the command registry. When the command registry on the server side detects a command addition event, it automatically retrieves the information of the newly added command (such as the command function name and command content description) from the event. The server can then save the information of the newly added command obtained from the command registry in the command database, thereby realizing the addition of the new command and generating corresponding identification information for the new command. Furthermore, a list of newly added commands can be generated based on the new command and its identification information.
[0053] This solution, through this configuration, allows for the registration of new commands and the maintenance of old commands via a pre-deployed command registry on the server side. It also enables the expansion of commands as needed to better meet the requirements of actual applications.
[0054] In this embodiment, optionally, the method further includes: periodically obtaining a list of newly added instructions from the server; and updating a preset instruction library based on the list of newly added instructions.
[0055] In this embodiment, an instruction management center is set up in the interactive all-in-one machine. The instruction management center can periodically obtain a list of new instructions from the server, and then update the instructions in the preset instruction library according to the list of new instructions, so that the instructions in the instruction database of the server and the instructions in the preset database of the interactive all-in-one machine are functionally matched.
[0056] This solution allows for the periodic updating of the preset instruction library based on the list of new instructions generated by the server, ensuring that the instructions in the server's instruction database match the instructions in the preset database of the interactive all-in-one machine in terms of functionality, thereby better realizing the corresponding functions.
[0057] Example 2
[0058] Figure 2 This is a flowchart of an information interaction method provided in Embodiment 2 of the present invention. This embodiment is an optimization based on the above embodiment.
[0059] like Figure 2 As shown, the method in this embodiment specifically includes the following steps:
[0060] S210: Obtain the voice information to be queried received through the target input interface, and convert the voice information to be queried into text information to be queried.
[0061] S220: Within the scope of the functions provided by the query prompt information, the server performs semantic parsing of the query text information using a pre-deployed large language model to obtain the target response information, so that the server can perform the target operation based on the target response information.
[0062] The server-side pre-deploys query prompts, which characterize the scope of functions allowed for semantic parsing by the large language model. This scope can be determined based on the functional characteristics of the interactive all-in-one machine. For example, using a conference all-in-one machine, query prompts might include options such as start screen recording, pause screen recording, resume screen recording, stop screen recording, open whiteboard, close whiteboard, open annotations, and close annotations. The target operation involves determining the target response instruction corresponding to the target response information from the pre-deployed instruction database on the server, as well as determining the associated identification information of the target response instruction.
[0063] In this embodiment, query hints are pre-deployed on the server side. When the large language model performs semantic parsing on the query text, it needs to parse within the functional scope provided by the query hints. This ensures that the large language model parses within the functional scope of the interactive all-in-one machine, thereby realizing the specific functions provided by the interactive all-in-one machine. It should be noted that without query hints to limit the parsing content, the large language model will only perform parsing behavior similar to a word chain based on its own large dataset and parameter set. Once the query hints are added, the large language model can parse the query text according to the query hints, thereby generating target response information that conforms to the query hints. This avoids the problem that the large language model's unrestrained semantic parsing of the query text may lead to an inability to parse the user's true intent and thus fail to accurately parse the target response information.
[0064] For example, suppose the query text is "What is a strawberry?". Without query hints, the large language model might generate some explanations and descriptions about strawberries. However, if the query hints "animal, plant" are added, the large language model will accurately answer "plant". This demonstrates that query hints can constrain and control the large language model to semantically analyze the query text according to pre-defined hints, thereby generating a target response that more closely matches the user's true intent.
[0065] S230: Receive the identification information associated with the target response instruction sent by the server, and determine the target feedback instruction from the preset instruction library based on the identification information associated with the target response instruction.
[0066] The identification information associated with the target response command and the target feedback command is the same.
[0067] S240, execute the target feedback instruction and return the execution result of the target feedback instruction to the target input interface.
[0068] The specific implementation methods of S230-S240 can be found in the detailed description of S130-S140, and will not be repeated here.
[0069] The technical solution of this invention pre-deploys query prompt information on the server side. This query prompt information represents the functional scope allowed for semantic parsing by the large language model. When performing semantic parsing on the query text information using the pre-deployed large language model on the server side, the semantic parsing must be performed within the functional scope provided by the query prompt information to obtain the target response information. This technical solution can accurately perform semantic parsing on voice input information using the large language model under the constraints of the query prompt information, and provide accurate and effective responses based on the semantic parsing results. It avoids the problem of the large language model failing to parse the user's true intent due to unrestrained semantic parsing of the query text information. It can realize the specific functions provided by the interactive all-in-one machine, making information interaction more intelligent and better meeting information interaction needs.
[0070] Figure 3 This is a schematic diagram illustrating an information interaction process provided in an embodiment of the present invention. Figure 3 As shown, the interactive all-in-one machine first acquires the voice information to be queried received through the target input interface, then converts the voice information to text to obtain the text information to be queried, and sends the text information to the command parser. Upon receiving the text information, the command parser calls an API to perform semantic parsing of the text information under the constraints of the query prompt information using a large language model deployed on the server, obtaining the target response information. Then, the server determines the target response command corresponding to the target response information and the associated identifier information from a pre-deployed command database, and returns the identifier information associated with the target response command to the command parser of the interactive all-in-one machine. After receiving the identifier information associated with the target response command, the command parser determines the target feedback command from a preset command library based on the identifier information associated with the target response command, and sends the target feedback command to the command executor. The command executor then executes the target feedback command to obtain the corresponding execution result, and returns the execution result to the target input interface, thereby achieving intelligent information interaction. In addition, new instructions can be registered through the instruction registration center deployed on the server side, and the preset instruction library of the interactive all-in-one machine can be periodically updated based on the new instructions and their identification information in the instruction database, so that the instructions in the instruction database and the preset instruction library are functionally matched.
[0071] Example 3
[0072] Figure 4 This is a schematic diagram of an information interaction device provided in Embodiment 3 of the present invention. This device can execute the information interaction method provided in any embodiment of the present invention, and possesses the corresponding functional modules and beneficial effects of the method execution. For example... Figure 4 As shown, the device includes:
[0073] The text information determination module 310 is used to acquire the voice information to be queried received through the target input interface and convert the voice information to be queried into text information to be queried.
[0074] The target operation execution module 320 is used to perform semantic parsing on the text information to be queried using a large language model pre-deployed on the server to obtain target response information, so that the server can perform a target operation based on the target response information; wherein, the target operation is to determine the target response instruction corresponding to the target response information from the instruction database pre-deployed on the server, and to determine the identification information associated with the target response instruction;
[0075] The feedback instruction determination module 330 is used to receive the identification information associated with the target response instruction sent by the server, and determine the target feedback instruction from a preset instruction library according to the identification information associated with the target response instruction; wherein the identification information associated with the target response instruction is the same as that associated with the target feedback instruction.
[0076] The feedback instruction execution module 340 is used to execute the target feedback instruction and return the execution result of the target feedback instruction to the target input interface.
[0077] Optionally, the server pre-deploys query hint information, which is used to characterize the functional scope that allows the large language model to perform semantic parsing.
[0078] Optionally, the target operation execution module 320 is used for:
[0079] Within the scope of the functions provided by the query prompt information, the target response information is obtained by semantic parsing the text information to be queried through a large language model pre-deployed on the server.
[0080] Optionally, the instruction database includes function names and response instructions associated with the function names;
[0081] Accordingly, the target operation execution module 320 is further configured to:
[0082] The target function name is determined from the server's instruction database based on the target response information;
[0083] The response instruction associated with the target function name is determined from the instruction database as the target response instruction.
[0084] Optionally, the device further includes:
[0085] A new instruction information acquisition module is added, which is used to acquire the newly added instruction information in the instruction addition event through the server if the server detects an instruction addition event.
[0086] The new instruction addition module is used to add new instructions to the instruction database according to the new instruction information, and generate identification information for the new instructions;
[0087] A new instruction list generation module is added, which is used to generate a new instruction list based on the new instruction and the identification information of the new instruction.
[0088] Optionally, the device further includes:
[0089] A new instruction list acquisition module is added, which is used to periodically acquire the new instruction list from the server.
[0090] The preset instruction library update module is used to update the preset instruction library according to the newly added instruction list.
[0091] The information interaction device provided in this embodiment of the invention can execute an information interaction method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method execution.
[0092] Example 4
[0093] Figure 5 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0094] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0095] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0096] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as information interaction methods.
[0097] In some embodiments, the information interaction method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the information interaction method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the information interaction method by any other suitable means (e.g., by means of firmware).
[0098] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0099] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0100] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0101] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0102] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0103] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0104] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0105] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. An information exchange method, characterized in that, The method includes: Acquire the voice information to be queried received through the target input interface, and convert the voice information to be queried into text information to be queried; The query text information is semantically parsed using a pre-deployed large language model on the server to obtain target response information, so that the server can execute a target operation based on the target response information; wherein, the target operation is to determine the target response instruction corresponding to the target response information from the instruction database pre-deployed on the server, and to determine the identification information associated with the target response instruction; The system receives the identification information associated with the target response instruction sent by the server, and determines the target feedback instruction from a preset instruction library based on the identification information associated with the target response instruction; wherein the identification information associated with the target response instruction and the target feedback instruction is the same; wherein the target response instruction and the target feedback instruction have corresponding functions; the target response instruction focuses on function description, and the target feedback instruction focuses on function implementation; Execute the target feedback instruction and return the execution result of the target feedback instruction to the target input interface; The server-side component pre-deploys query hint information, which is used to characterize the functional scope that allows the large language model to perform semantic parsing. The target response information is obtained by semantically parsing the text information to be queried using a large language model pre-deployed on the server, including: Within the scope of the functions provided by the query prompt information, the target response information is obtained by semantic parsing the text information to be queried through a large language model pre-deployed on the server.
2. The method according to claim 1, characterized in that, The instruction database includes function names and response instructions associated with those function names; Accordingly, determining the target response instruction corresponding to the target response information from the instruction database pre-deployed on the server includes: The target function name is determined from the server's instruction database based on the target response information; The response instruction associated with the target function name is determined from the instruction database as the target response instruction.
3. The method according to claim 1, characterized in that, The method further includes: If the server detects a command addition event, it obtains the newly added command information from the command addition event through the server. Based on the newly added instruction information, a new instruction is added to the instruction database, and identification information is generated for the new instruction; A list of new instructions is generated based on the new instructions and their identifier information.
4. The method according to claim 3, characterized in that, The method further includes: The list of newly added instructions is periodically retrieved from the server. The preset instruction library is updated based on the newly added instruction list.
5. An information interaction device, characterized in that, The device includes: The text information determination module is used to acquire the voice information to be queried received through the target input interface and convert the voice information to be queried into text information to be queried. The target operation execution module is used to perform semantic parsing on the text information to be queried using a large language model pre-deployed on the server to obtain target response information, so that the server can execute a target operation based on the target response information; wherein, the target operation is to determine the target response instruction corresponding to the target response information from the instruction database pre-deployed on the server, and to determine the identification information associated with the target response instruction; The feedback instruction determination module is used to receive the identification information associated with the target response instruction sent by the server, and determine the target feedback instruction from a preset instruction library based on the identification information associated with the target response instruction; wherein the identification information associated with the target response instruction and the target feedback instruction is the same; wherein the target response instruction and the target feedback instruction have corresponding functions; the target response instruction focuses on function description, and the target feedback instruction focuses on function implementation; The feedback instruction execution module is used to execute the target feedback instruction and return the execution result of the target feedback instruction to the target input interface; The server-side component pre-deploys query hint information, which characterizes the scope of semantic parsing allowed by the large language model. The target response information is obtained by semantically parsing the queried text information using the pre-deployed large language model on the server, including: Within the scope of the functions provided by the query prompt information, the target response information is obtained by semantic parsing the text information to be queried through a large language model pre-deployed on the server.
6. An electronic device, characterized in that, The device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the information interaction method according to any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute the information interaction method according to any one of claims 1-4.
Citation Information
Patent Citations
Multi-round man-machine conversation method and device and apparatus
CN110196927A
Human-computer interaction method, device and system
CN116483980A