Method and device for generating reply information, electronic equipment and medium
By using a lightweight language model to predict dialogue information and cache responses when the user pauses input, and combining this with a large language model to optimize state machine state transitions, the problem of slow response generation by the large language model is solved, resulting in faster responses and lower latency, thus improving the user experience.
Patent Information
- Application Number
- CN202410459147.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-16
- Publication Date
- 2025-10-28
AI Technical Summary
Current large language models suffer from high computational complexity when generating response information, resulting in slow output speed and impacting user experience, especially in dialogue interactions where the delay time is too long.
By taking advantage of the pauses and thinking periods during user input, a lightweight language model is used to predict complete dialogue information in advance and generate and cache response information. The final response is then combined with a large language model, optimizing the client-side state machine state transitions to reduce frequent requests to the server.
It accelerates the generation of reply messages, reduces latency, improves user experience and interaction smoothness, and increases the hit rate of reply messages.
Smart Images

Figure CN120849536A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence, particularly to the fields of deep learning, intelligent search, and natural language processing, and specifically to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating response information. Background Technology
[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.
[0003] Human-computer interaction (HCI) is a way for humans to interact with machines using natural language. With the continuous development of artificial intelligence technology, machines have become capable of understanding human-generated information, comprehending its inherent meaning, and providing corresponding feedback. In these operations, the accuracy of semantic understanding, the speed of feedback, and the provision of appropriate opinions or suggestions all influence the smoothness of HCI interaction.
[0004] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention
[0005] This disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating response information.
[0006] According to one aspect of this disclosure, a method for generating response information executed by a server is provided, comprising: acquiring first dialogue information, wherein the first dialogue information is currently input dialogue information acquired when a preset time period is detected as a pause in input during the dialogue information input process; performing complete dialogue information prediction based on the first dialogue information to obtain predicted second dialogue information; generating first response information for replying to the second dialogue information using a first language model based on the second dialogue information; acquiring third dialogue information, wherein the third dialogue information is currently input dialogue information acquired when the dialogue information input process is detected as complete; and acquiring and displaying the first response information in response to determining that the third dialogue information semantically matches the second dialogue information.
[0007] According to another aspect of this disclosure, a server-based apparatus for generating response information is provided, comprising: a first acquisition unit configured to acquire first dialogue information, wherein the first dialogue information is currently input dialogue information acquired when input is stopped for a preset time period during the dialogue information input process; a first prediction unit configured to perform complete dialogue information prediction based on the first dialogue information to obtain predicted second dialogue information; a first response unit configured to generate first response information for responding to the second dialogue information based on the second dialogue information using a second language model; a second acquisition unit configured to acquire third dialogue information, wherein the third dialogue information is currently input dialogue information acquired when the dialogue information input process is detected to be completed; and a third acquisition unit configured to acquire and display the first response information in response to determining that the third dialogue information semantically matches the second dialogue information.
[0008] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory storing instructions executable by the at least one processor to enable the at least one processor to perform the methods described in this disclosure.
[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing a computer to perform the methods described in this disclosure.
[0010] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods described in this disclosure.
[0011] According to one or more embodiments of this disclosure, the pauses and thinking gaps during the user's typing or speaking process are fully utilized to predict the complete dialogue information in advance and generate and cache the response information, thereby achieving the purpose of accelerating the reasoning speed of response information and reducing latency.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0013] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0014] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein may be implemented according to embodiments of the present disclosure is shown;
[0015] Figure 2 A flowchart of a method for generating response information according to an embodiment of the present disclosure is shown;
[0016] Figure 3 A schematic diagram of a human-computer interaction interface according to an embodiment of the present disclosure is shown;
[0017] Figure 4 A schematic diagram of client state machine state transitions according to an embodiment of the present disclosure is shown;
[0018] Figure 5 A schematic diagram illustrating the effect of generating response information according to an embodiment of the present disclosure is shown;
[0019] Figure 6 A structural block diagram of an apparatus for generating response information according to embodiments of the present disclosure is shown; and
[0020] Figure 7 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0021] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0022] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.
[0023] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.
[0024] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0025] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1 The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.
[0026] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of methods for generating response information.
[0027] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105, and / or 106 under a Software as a Service (SaaS) model.
[0028] exist Figure 1In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the various methods described herein, and is not intended to be limiting.
[0029] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to obtain user input. The client devices can provide an interface that allows users to interact with the client devices. The client devices can also output information to the user through this interface. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.
[0030] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.
[0031] Network 110 can be any type of network well known to those skilled in the art, and can use any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0032] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.
[0033] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0034] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.
[0035] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.
[0036] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as predicted dialogue information and its corresponding responses. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located remotely to server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.
[0037] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.
[0038] Figure 1 The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.
[0039] With the explosive growth of AI native applications centered around large language models, more and more users are leveraging these models to meet their needs and solve problems in life and work. Large language models can meet users' needs in various fields and scenarios, such as text creation, academic research, education, emotional support, and language games, through dialogue.
[0040] However, current large language models have a significant problem: as the number of model parameters increases, so does the computational complexity, resulting in a slower output speed. Generally, a large language model with hundreds of billions of parameters takes about 10 seconds to complete a single question response. This severely impacts the user experience. Although AI-native applications provide a smooth waiting experience through client-side effects like a typewriter, the initial packet latency (the delay from sending a question to receiving the first result packet) is typically around 1.8-1.9 seconds. This waiting time still provides a poor user experience because the large language model doesn't provide any feedback during this period, leaving the user in a waiting state. Therefore, there is significant room for improvement in the current dialogue interaction experience based on large language models, especially regarding dialogue latency.
[0041] Current performance optimization methods for large language model inference latency typically include the following: 1) Starting from the model inference architecture, further improve the computational throughput of the prediction service through batch caching, optimized memory allocation, quantization, speculative sampling, etc., thereby reducing output latency; 2) Starting from the large language model structure, reduce the computational load of large language model inference, thereby effectively reducing output latency; 3) From a hardware perspective, further reduce model latency through faster and better-performing hardware.
[0042] However, whether starting from the architecture of model inference or from the structure of a large language model, the approach of reducing output latency is only considered after the user input question is completed and during backend processing, which has limited effect. On the other hand, from a hardware perspective, not only does it increase costs, but it is also likely that the goal of reducing output latency will be difficult to achieve due to the difficulty in finding better-performing hardware products.
[0043] Therefore, embodiments of this disclosure provide a method for generating response information. Figure 2 A flowchart of a method for generating response information according to an embodiment of the present disclosure is shown, such as... Figure 2 As shown, method 200 includes: acquiring first dialogue information, wherein the first dialogue information is the currently input dialogue information acquired when a preset time period is detected during the dialogue information input process (step 210); performing complete dialogue information prediction based on the first dialogue information to obtain predicted second dialogue information (step 220); generating first response information for replying to the second dialogue information through a first language model based on the second dialogue information (step 230); acquiring third dialogue information, wherein the third dialogue information is the currently input dialogue information acquired when the dialogue information input process is detected to be completed (step 240); and acquiring and displaying the first response information in response to determining that the third dialogue information semantically matches the second dialogue information (step 250).
[0044] According to embodiments of this disclosure, the pauses and thinking gaps during the user's typing or speaking process are fully utilized to predict complete dialogue information in advance and generate and cache response information, thereby achieving the purpose of accelerating the reasoning speed of response information and reducing latency.
[0045] In this disclosure, a dialogue information input process can start from the user entering dialogue information through a preset operation (e.g., typing dialogue text in a dialog box in the user interface) and end with the user sending out a complete dialogue information (e.g., clicking the send button in the user interface).
[0046] Figure 3A schematic diagram of a human-computer interaction interface according to an exemplary embodiment of the present disclosure is shown.
[0047] like Figure 3 As shown, when a user enters dialogue information 301 through the dialogue information input box on the electronic device, the electronic device can generate reply information 302 to reply to the dialogue information and display it on the interface of the electronic device.
[0048] In some embodiments, dialogue information can be input by the user in the form of text, voice, or other means. When inputting via voice, speech recognition technology can be used to convert the speech into text to obtain the corresponding dialogue information.
[0049] According to some embodiments, predicting complete dialogue information based on the first dialogue information may include: inputting the first dialogue information into a second language model to predict complete dialogue information, so as to obtain the predicted second dialogue information.
[0050] In this disclosure, the first language model and the second language model can be the same language model or different language models, and there is no restriction.
[0051] For example, in some embodiments, at least one of the first language model and the second language model can be trained on at least a predetermined scale of knowledge resources and dialogue data. At least one of the first language model and the second language model can be a knowledge-enhanced large language model for dialogue (e.g., ERNIE bot), trained on massive amounts of knowledge resources and dialogue data (e.g., including trillions of web pages, billions of search results, hundreds of millions of images, billions of voice requests per day, over 50 billion text requests, and over 550 billion pieces of factual knowledge).
[0052] Therefore, by applying this type of large language model, in addition to directly processing casual conversational dialogue information, it can also directly generate response information for dialogue information involving logical reasoning, common sense, and image generation, thereby improving generation efficiency while generating higher quality response information.
[0053] In some examples, the second language model can be a lightweight language model relative to the first language model. Leveraging the fast computational capabilities of the lightweight language model improves the prediction speed of dialogue information, thereby accelerating the reasoning speed of the first language model and reducing its response latency.
[0054] According to some embodiments, predicting complete dialogue information based on the first dialogue information may include: searching for complete dialogue information from historical inputs based on the first dialogue information to obtain complete dialogue information that semantically matches the first dialogue information, and using it as the second dialogue information.
[0055] In some examples, based on the first dialogue information, a search engine can be used to search the entire network for the user's historical search queries to retrieve queries that semantically match the first dialogue information, which are then used as the second dialogue information. In this way, by using historical search information as the predicted complete dialogue information, retrieval efficiency is improved, the prediction speed of dialogue information is increased, and thus the goal of accelerating the reasoning speed of the first language model and reducing its response latency is achieved.
[0056] According to some embodiments, before obtaining the third dialogue information, the method further includes: determining the sixth dialogue information for guiding the dialogue information input process based on the corresponding dialogue information predicted from the complete dialogue information, so that the sixth dialogue information is displayed on the client.
[0057] In some examples, dialogue information predicted by the second language model or complete historical dialogue information retrieved from search engines can be returned to the client, for example, through a suggested menu or speech bubble. This allows users to complete or speed up dialogue input by clicking on menus or speech bubbles.
[0058] This embodiment can improve the user experience in information input and also help improve the hit rate of prediction results and corresponding response information.
[0059] According to some embodiments, the method according to this disclosure further includes: in response to determining that the third dialogue information does not semantically match the second dialogue information, generating and displaying second response information for replying to the third dialogue information based on the third dialogue information using the first language model.
[0060] For example, when the predicted complete dialogue information does not semantically match the complete dialogue information input by the user, the first response information generated by the first language model based on the predicted complete dialogue information will be unusable. In this case, the cached first response information can be cleared, and a second response information for replying to the third dialogue information can be generated by the first language model based on the complete dialogue information input by the user (i.e., the third dialogue information).
[0061] In some examples, when generating a first response message for replying to the second dialogue information using a first language model based on the second dialogue information, contextual data related to the current dialogue information and appropriate prompts can also be input into the first language model to generate the first response message for replying to the second dialogue information.
[0062] In some examples, when it is determined that the third dialogue information does not semantically match the second dialogue information, when generating the second response information for replying to the third dialogue information through the first language model based on the third dialogue information, contextual data related to the current dialogue information and appropriate prompts can also be input into the first language model to generate the first response information for replying to the third dialogue information.
[0063] In some embodiments, when dialogue information is input by a user in text form or when dialogue information input in voice form can be converted into text by speech recognition technology, a text semantic similarity calculation algorithm can be used to determine whether the dialogue information is semantically matched.
[0064] Understandably, any suitable algorithm can be used to determine whether dialogue information is semantically matched, including but not limited to cosine similarity, Jaccard similarity, edit distance, word vector-based methods, BERT similarity calculation, and so on.
[0065] In some examples, the response information may include a summary of the dialogue information generated by rewriting the corresponding dialogue information, as well as information for responding to the corresponding dialogue information (e.g., knowledge information for answering the corresponding dialogue information).
[0066] According to some embodiments, the currently entered dialogue information is sent to the server by the client based on the state of a preset state machine, wherein the state of the state machine includes an input state, a waiting state, and a sending state, and the client triggers the sending operation of the currently entered dialogue information when the state of the state machine changes to the waiting state and the sending state.
[0067] Specifically, for any dialogue information input process, the state of the state machine is determined by the client based on the following operations: in response to detecting the input of dialogue information, the state of the state machine is changed to the input state; in response to detecting that the input of dialogue information has stopped for a preset time period during the dialogue information input process, the state of the state machine is changed to the waiting state; in response to detecting the completion of the dialogue information input process, the state of the state machine is changed to the sending state.
[0068] For example, the operation to complete this dialogue input process could be: clicking the send button in a preset area on the client, clicking a preset key (such as the Enter key), etc. It is understood that other operations that can be used to complete this dialogue input process are also possible, and are not limited here.
[0069] Figure 4 A schematic diagram of client state machine state transitions according to an embodiment of the present disclosure is shown. For example... Figure 4 As shown, when a user initiates the current dialogue input process by entering text (or other forms, such as voice), the state machine enters the "input state," which remains active as text is continuously entered. If the user pauses input for a preset time period (e.g., 2 seconds), the state machine transitions to the "waiting state," and remains in this state if no input continues. When the user enters text again, the state transitions back to the "input state." When the user ends the current dialogue input process, for example, by clicking the send button, the state machine transitions to the "send state," and the current dialogue input process is complete.
[0070] In some examples, the operation of sending the currently entered dialogue information to the server is triggered when the state machine enters the "waiting state" and "send state".
[0071] In the above embodiment, when the client transitions between the waiting state and the sending state in the state machine, it triggers the sending operation of the currently entered dialogue information. That is, each time the client enters the "waiting state" and each time it enters the "sending state," it triggers only one operation to send the currently entered dialogue information to the server. This avoids sending the same dialogue information to the server frequently in a short period of time.
[0072] In the embodiments of this disclosure, the operation based on the client state machine can effectively avoid the client frequently sending dialogue information to the server, thereby increasing the computational burden on the server and causing problems such as lag and delay.
[0073] It is understandable that during the "waiting state," the user is likely thinking about what they will input next. Through embodiments of this disclosure, by sending currently entered dialogue information to the server during the "waiting state," the server can effectively utilize the gap between the user's pause for thought and subsequent input. Based on a language model (e.g., a lightweight language model), it can quickly predict the complete dialogue information the user might need to input, allowing the large language model used to generate the response information to perform the reasoning operation for generating the response information in advance. This, in turn, achieves the goal of accelerating reasoning speed and reducing latency in returning responses.
[0074] According to some embodiments, before obtaining the third dialogue information, the method of this disclosure may further include: in response to obtaining fourth dialogue information during the process of the first language model generating first response information for replying to the second dialogue information, determining whether the fourth dialogue information and the second dialogue information semantically match, wherein the fourth dialogue information is a new currently input dialogue information obtained when the input of the dialogue information is detected to have stopped for the preset time period; performing complete dialogue information prediction based on the fourth dialogue information to obtain predicted fifth dialogue information; in response to determining that the fifth dialogue information and the second dialogue information semantically match, continuing the process of the first language model generating the first response information for replying to the second dialogue information; and in response to determining that the fifth dialogue information and the second dialogue information semantically do not match, stopping the process of the first language model generating the first response information for replying to the second dialogue information, and generating third response information for replying to the fifth dialogue information through the first language model based on the fifth dialogue information.
[0075] It is understood that the operation of predicting complete dialogue information based on the fourth dialogue information is similar to the operation of predicting complete dialogue information based on the first dialogue information, and can be implemented by any suitable embodiment described above, which will not be repeated here.
[0076] In this embodiment, only the latest predicted complete dialogue information and its corresponding response information can be retained, thereby saving computation and storage space; at the same time, since the latest predicted dialogue information is generally determined based on the user's latest input, its prediction results and the resulting response information will be more accurate.
[0077] In some embodiments, the first language model can be any suitable Large Language Model (LLM). A LLM typically transforms the input text (such as a user's question or the context of a dialogue) into a sequence of tokens, each token representing a word, punctuation mark, or sub-word fragment in the text. The LLM infers the next most likely token based on the current context (i.e., the already generated token sequence). This process is iterative, inferring and adding a new token to the sequence each time, until a preset stopping condition is met (such as reaching a maximum length or encountering a specific end marker). The LLM decodes the generated token sequence back into text format and outputs it as a response. This process may include merging sub-word fragments back into complete words and adding appropriate punctuation. Therefore, in some embodiments, the LLM can determine whether to pause the current response generation process after each token is generated. If it is determined that the current inference process needs to be paused (e.g., obtaining fourth dialogue information), the current response generation process is paused; otherwise, the inference of the next token continues.
[0078] In some examples, speculative decoding can be combined with the process of generating response information through first-language model inference. Speculative decoding accelerates inference by increasing the parallelism of LLM computation in each decoding step, thereby reducing the total number of decoding steps (i.e., reducing repeated reads and writes of LLM parameters). As an emerging LLM inference acceleration technique, speculative decoding achieves acceleration of LLM inference without sacrificing the decoding quality of the LLM.
[0079] According to some embodiments, in response to determining that the third dialogue information and the second dialogue information semantically match, obtaining the first response information and displaying it includes: when it is determined that the fifth dialogue information and the second dialogue information do not semantically match, determining whether the third dialogue information and the fifth dialogue information semantically match; and in response to determining that the third dialogue information and the fifth dialogue information semantically match, obtaining the third response information and displaying it.
[0080] According to some embodiments, obtaining the third dialogue information may further include: in response to obtaining the third dialogue information during the process of the first language model generating first response information for replying to the second dialogue information, determining whether the third dialogue information and the second dialogue information semantically match; in response to determining that the third dialogue information and the second dialogue information semantically match, continuing the process of the first language model generating first response information for replying to the second dialogue information; and in response to determining that the third dialogue information and the second dialogue information semantically do not match, stopping the process of the first language model generating first response information for replying to the second dialogue information.
[0081] In this embodiment, if the first language model obtains complete dialogue information (i.e., third dialogue information) input by the user during the process of generating response information based on the predicted dialogue information, and if the predicted dialogue information semantically matches the obtained complete dialogue information input by the user, then the process of generating response information continues. The generated response information is then displayed.
[0082] Otherwise, the process of generating response information is stopped, and instead, response information is generated using the first language model based on the complete dialogue information obtained from the user input (i.e., the third dialogue information). Then, the generated response information is displayed.
[0083] According to embodiments of this disclosure, if, after obtaining complete dialogue information input by the user (i.e., third dialogue information), it is determined that there is semantically matching predicted dialogue information and its corresponding response information, then the response information is directly obtained, thereby significantly shortening the response delay time. If, after obtaining complete dialogue information input by the user (i.e., third dialogue information), it is determined that there is semantically matching predicted dialogue information and its corresponding response information is being generated, then the response information is waited for to be generated and obtained directly after generation, thereby also shortening the response delay time to a certain extent.
[0084] In some examples, retrieving and displaying the corresponding response information may include returning the response information generated by the first language model to the client in the form of streaming packets. In this streaming implementation, the response data generated by the first language model is divided into a series of small data packets (streaming packets) for transmission. Each data packet contains a portion of the response information, and these packets are returned sequentially according to the order of generation. In this case, the first language model continuously generates and returns data, rather than waiting for all data to be generated before returning it all at once. This streaming approach is particularly useful for applications requiring real-time responses, such as online chat and voice assistants. Based on the streaming packet format, user waiting time can be further reduced, improving the smoothness of interaction and user experience.
[0085] According to embodiments of this disclosure, it is determined whether the dialogue information input process has been completed based on a preset flag bit of a data packet containing currently entered dialogue information.
[0086] It is understandable that determining whether the dialogue information input process has been completed is equivalent to determining whether the obtained dialogue information is the first or the third dialogue information. Based on the information in the preset flag bits of the data packet, the server can easily determine whether the current dialogue information input process has been completed or not, thereby determining the corresponding subsequent operations (e.g., calling the second language model to predict the complete dialogue information or directly calling the first language model for response reasoning).
[0087] In an exemplary embodiment according to this disclosure, after obtaining the corresponding dialogue information, the server can first determine whether the obtained dialogue information was obtained before or after the dialogue information input process was completed.
[0088] If the dialogue information is obtained before the input process is complete, a second language model (such as a lightweight language model) is invoked to predict the complete dialogue information. Furthermore:
[0089] In some examples, it can be further determined whether the first language model is currently performing an inference task (i.e., generating response information). If it is not inference, the first language model is invoked to generate response information from the currently acquired dialogue information, and the predicted complete dialogue information and its corresponding response information are associated and cached. If there are other cached results at this time, they can be directly overwritten so that only the latest prediction result is stored in the cache.
[0090] Furthermore, in some examples, if an inference task is in progress and the dialogue information upon which the current inference task is based semantically matches the currently predicted complete dialogue information, the inference task can be allowed to complete before the predicted complete dialogue information and its corresponding response information are cached together. If inference is in progress and the dialogue information upon which the current inference task is based does not semantically match the currently predicted complete dialogue information, the current inference task is paused, and the first language model is invoked to generate response information based on the currently predicted complete dialogue information. The predicted complete dialogue information and its corresponding response information are then cached together.
[0091] If the information is obtained after the dialogue input process has been completed, in some examples, it can be further determined whether the first language model is currently performing an inference task (i.e., generating response information). Furthermore:
[0092] In some examples, if no reasoning is in progress, it can be further determined whether the cache contains the predicted complete dialogue information and its corresponding response information. If the cache does not contain the predicted complete dialogue information and its corresponding response information, or if the predicted complete dialogue information stored in the cache does not semantically match the currently acquired dialogue information, then the first language model is invoked to generate the response information based on the currently acquired dialogue information.
[0093] In some examples, if reasoning is in progress and the dialogue information upon which the current reasoning task is based semantically matches the currently acquired dialogue information, the process can wait for the current reasoning task to complete before obtaining the reasoned response information. If reasoning is in progress and the dialogue information upon which the current reasoning task is based does not semantically match the currently acquired dialogue information, the current reasoning task is paused, and the first language model is then invoked to generate the response information based on the currently acquired dialogue information.
[0094] The embodiments of this disclosure achieve an effect similar to communication between friends, relatives, or lovers, where both parties understand each other's thoughts. The intelligent question-and-answer system implemented based on the methods of the embodiments of this disclosure has very high question-and-answer efficiency.
[0095] Figure 5 A schematic diagram illustrating the effect of generating response information according to an embodiment of this disclosure is shown. Figure 5 As shown, after optimization based on the method described in the embodiments of this disclosure, the time taken to return the first token of the reply information and the total time taken to return the reply information are both reduced.
[0096] In some embodiments, when predicting complete dialogue information based on a second language model, the second language model can be further optimized based on the dialogue information input by the user, making the model more consistent with the current user's questioning habits. This will result in more accurate prediction of complete dialogue information based on the second language model, helping to improve the hit rate of prediction results and their corresponding response information, and reducing response delay.
[0097] According to embodiments of this disclosure, such as Figure 6As shown, a server-side device 600 for generating response information is also provided, comprising: a first acquisition unit 610 configured to acquire first dialogue information, wherein the first dialogue information is the currently input dialogue information acquired when a preset time period is detected as a pause in input during the dialogue information input process; a first prediction unit 620 configured to predict complete dialogue information based on the first dialogue information to obtain predicted second dialogue information; a first response unit 630 configured to generate first response information for responding to the second dialogue information based on the second dialogue information using a first language model; a second acquisition unit 640 configured to acquire third dialogue information, wherein the third dialogue information is the currently input dialogue information acquired when the dialogue information input process is detected as complete; and a third acquisition unit 650 configured to acquire and display the first response information in response to determining that the third dialogue information semantically matches the second dialogue information.
[0098] Here, the operation of each of the above units 610 to 650 of the server-side device 600 for generating reply information is similar to the operation of steps 210 to 250 described above, and will not be repeated here.
[0099] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0100] According to embodiments of this disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.
[0101] refer to Figure 7 The present invention describes a structural block diagram of an electronic device 700 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0102] like Figure 7As shown, the electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. The RAM 703 may also store various programs and data required for the operation of the electronic device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0103] Multiple components in electronic device 700 are connected to I / O interface 705, including: input unit 706, output unit 707, storage unit 708, and communication unit 709. Input unit 706 can be any type of device capable of inputting information to electronic device 700. Input unit 706 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 707 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 708 may include, but is not limited to, disk and optical disk. Communication unit 709 allows electronic device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0104] The computing unit 701 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as method 200. For example, in some embodiments, method 200 may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of method 200 described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform method 200 by any other suitable means (e.g., by means of firmware).
[0105] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0106] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0107] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0108] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0109] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.
[0110] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0111] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0112] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.
Claims
1. A method for generating response information executed by a server, comprising: Acquire first dialogue information, wherein the first dialogue information is the currently entered dialogue information acquired when a preset time period is detected as a pause in input during the dialogue information input process; Based on the first dialogue information, perform complete dialogue information prediction to obtain the predicted second dialogue information; Based on the second dialogue information, a first response message is generated using a first language model to respond to the second dialogue information. Acquire third dialogue information, wherein the third dialogue information is the currently entered dialogue information acquired when the dialogue information input process is detected to be completed; and In response to determining that the third dialogue information semantically matches the second dialogue information, the first reply information is obtained and displayed.
2. The method of claim 1, further comprising: In response to determining that the third dialogue information does not semantically match the second dialogue information, a second response information is generated and displayed based on the third dialogue information using the first language model.
3. The method as described in claim 1 or 2, wherein, The currently entered dialogue information is sent to the server by the client based on the state of a preset state machine. The state machine includes an input state, a waiting state, and a send state. The client triggers the send operation of the currently entered dialogue information when the state machine transitions to the waiting state or the send state. Furthermore, for any dialogue information input process, the state of the state machine is determined by the client based on the following operations: In response to the detection of input dialogue information, the state of the state machine is changed to the input state; In response to detecting that the input of dialogue information has stopped for a preset time period during the dialogue information input process, the state of the state machine is changed to the waiting state; In response to detecting the completion of the dialogue information input process, the state of the state machine is changed to the sending state.
4. The method according to any one of claims 1-3, wherein, Before obtaining information from the third dialogue, it also includes: In response to obtaining fourth dialogue information during the process of generating first response information for replying to the second dialogue information in the first language model, it is determined whether the fourth dialogue information and the second dialogue information semantically match, wherein the fourth dialogue information is a new currently input dialogue information obtained when the input of the dialogue information is detected to have stopped for the preset time period again during the input of the dialogue information; Based on the fourth dialogue information, perform complete dialogue information prediction to obtain the predicted fifth dialogue information; In response to determining that the fifth dialogue information semantically matches the second dialogue information, the process of generating first response information from the first language model to respond to the second dialogue information continues; and In response to determining that the fifth dialogue information does not semantically match the second dialogue information, the process of generating the first response information for replying to the second dialogue information by the first language model is stopped, and a third response information for replying to the fifth dialogue information is generated by the first language model based on the fifth dialogue information.
5. The method of claim 4, wherein, In response to determining that the third dialogue information semantically matches the second dialogue information, obtaining the first response information and displaying it includes: When it is determined that the fifth dialogue information does not semantically match the second dialogue information, it is determined whether the third dialogue information semantically matches the fifth dialogue information; and In response to determining that the third dialogue information and the fifth dialogue information semantically match, the third response information is obtained and displayed.
6. The method according to any one of claims 1-5, wherein, Obtaining third-party dialogue information includes: In response to obtaining the third dialogue information during the process of generating first response information for replying to the second dialogue information in the first language model, it is determined whether the third dialogue information and the second dialogue information semantically match; In response to determining that the third dialogue information semantically matches the second dialogue information, the process of generating first response information from the first language model to respond to the second dialogue information continues; and In response to determining that the third dialogue information does not semantically match the second dialogue information, the process of generating the first response information for replying to the second dialogue information by the first language model is stopped.
7. The method of claim 1, wherein, The system determines whether the dialogue input process has been completed based on a preset flag bit in a data packet containing currently entered dialogue information.
8. The method according to any one of claims 1-7, wherein, Predicting complete dialogue information based on the first dialogue information includes: The first dialogue information is input into the second language model to predict the complete dialogue information, so as to obtain the predicted second dialogue information.
9. The method according to any one of claims 1-7, wherein, Predicting complete dialogue information based on the first dialogue information includes: Based on the first dialogue information, a search is performed on the complete dialogue information of the historical input to obtain the complete dialogue information that semantically matches the first dialogue information, which is then used as the second dialogue information.
10. The method of any one of claims 1-9, further comprising, before obtaining the third dialogue information: Based on the corresponding dialogue information obtained from the prediction of complete dialogue information, a sixth dialogue information is determined to guide the dialogue information input process, so that the sixth dialogue information is displayed on the client.
11. A server-side apparatus for generating response information, comprising: The first acquisition unit is configured to acquire first dialogue information, wherein the first dialogue information is the currently input dialogue information acquired when a preset time period is detected as a stop in input during the input process of dialogue information; The first prediction unit is configured to predict complete dialogue information based on the first dialogue information in order to obtain the predicted second dialogue information. The first response unit is configured to generate first response information for replying to the second dialogue information based on the second dialogue information using a first language model. The second acquisition unit is configured to acquire third dialogue information, wherein the third dialogue information is the currently input dialogue information acquired when the dialogue information input process is detected to be completed; and The third acquisition unit is configured to acquire and display the first response information in response to determining that the third dialogue information semantically matches the second dialogue information.
12. The apparatus of claim 11, further comprising a second recovery unit configured as follows: In response to determining that the third dialogue information does not semantically match the second dialogue information, a second response information is generated and displayed based on the third dialogue information using the first language model.
13. The apparatus of claim 11 or 12, wherein, The currently entered dialogue information is sent to the server by the client based on the state of a preset state machine. The state machine includes an input state, a waiting state, and a send state. The client triggers the send operation of the currently entered dialogue information when the state machine transitions to the waiting state or the send state. Furthermore, for any dialogue information input process, the state of the state machine is determined by the client based on the following operations: In response to the detection of input dialogue information, the state of the state machine is changed to the input state; In response to detecting that the input of dialogue information has stopped for a preset time period during the dialogue information input process, the state of the state machine is changed to the waiting state; In response to detecting the completion of the dialogue information input process, the state of the state machine is changed to the sending state.
14. The apparatus according to any one of claims 11-13, wherein, Before obtaining information from the third dialogue, it also includes: The first determining unit is configured to, in response to obtaining fourth dialogue information during the process of the first language model generating first response information for replying to the second dialogue information, determine whether the fourth dialogue information and the second dialogue information semantically match, wherein the fourth dialogue information is a new currently input dialogue information obtained when the input of the dialogue information is detected to have stopped for the preset time period again during the input of the dialogue information; The second prediction unit is configured to predict complete dialogue information based on the fourth dialogue information to obtain the predicted fifth dialogue information. The second determining unit is configured to, in response to determining that the fifth dialogue information semantically matches the second dialogue information, continue the process of generating first response information from the first language model to reply to the second dialogue information; and The third determining unit is configured to, in response to determining that the fifth dialogue information does not semantically match the second dialogue information, stop the process of generating a first response information for replying to the second dialogue information using the first language model, and generate a third response information for replying to the fifth dialogue information based on the fifth dialogue information using the first language model.
15. The apparatus of claim 14, wherein, The third acquisition unit includes: A unit for determining whether the third dialogue information and the fifth dialogue information semantically match when it is determined that the fifth dialogue information does not semantically match the second dialogue information; and A unit for responding to determining that the third dialogue information and the fifth dialogue information semantically match, obtaining the third response information, and displaying it.
16. The apparatus according to any one of claims 11-15, wherein, The second acquisition unit includes: A unit for responding to obtaining the third dialogue information and determining whether the third dialogue information and the second dialogue information are semantically matched during the process of generating first response information for replying to the second dialogue information in the first language model; A unit for responding to determining that the third dialogue information semantically matches the second dialogue information and continuing the process of generating a first response information for replying to the second dialogue information using the first language model; and A unit for stopping the generation of a first response message for replying to the second dialogue information in response to determining that the third dialogue information does not semantically match the second dialogue information.
17. The apparatus of claim 11, wherein, The system determines whether the dialogue input process has been completed based on a preset flag bit in a data packet containing currently entered dialogue information.
18. The apparatus according to any one of claims 11-17, wherein, The first prediction unit includes a unit for inputting the first dialogue information into a second language model to predict complete dialogue information and obtain the predicted second dialogue information.
19. The apparatus according to any one of claims 11-17, wherein, The first prediction unit includes: a unit for searching historical input complete dialogue information based on the first dialogue information to obtain searched complete dialogue information that semantically matches the first dialogue information, and serving as the second dialogue information.
20. The apparatus of any one of claims 11-19, further comprising: The prompting unit is configured to determine a sixth dialogue information for guiding the dialogue information input process based on the corresponding dialogue information predicted from the complete dialogue information, so that the sixth dialogue information is displayed on the client.
21. An electronic device, comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method of any one of claims 1-10.
22. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-10.
23. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the method of any one of claims 1-10.