A data processing method and device and related equipment

By lowering the verification standard for draft tokens and using a pre-trained model to perform probabilistic matching verification of draft tokens in the draft library, the problems of low efficiency and long response time of generative language models are solved, and more efficient answer generation is achieved.

CN122633800APending Publication Date: 2026-08-25HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510214168.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing generative language models are inefficient in answering questions, have long response times, and have strict draft validation standards that result in low draft token validation pass rates.

Method used

By lowering the verification standard for draft tokens, a pre-trained model is used to verify the matching probability of draft tokens in the draft library. Draft tokens with a matching probability higher than the target threshold are selected to generate an answer.

Benefits of technology

This increases the probability of successful draft token verification, improves the efficiency of generating answers, and reduces the response time of the language generation model system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122633800A_ABST
    Figure CN122633800A_ABST
Patent Text Reader

Abstract

The application provides a data processing method for improving the efficiency of answering questions. The method comprises: obtaining a first input sent by a user; determining a plurality of target drafts matching the first input from a draft library according to the first input, the target drafts comprising at least one draft token; calling a pre-trained model to obtain a matching probability of each draft token in each target draft, the matching probability being a probability of the pre-trained model outputting the draft token when the first input is input; determining a verified draft token, the verified draft token being a draft token with a matching probability greater than a target threshold; and determining an available target draft in the plurality of target drafts, the verified draft token in the target available draft being used to generate an answer to the first input. The application also provides corresponding devices, computing device clusters, computer-readable storage media, and computer program products.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a data processing method, apparatus and related equipment. Background Technology

[0002] Currently, most common Large Language Models (LLMs) are generative models with an autoregressive architecture. An autoregressive generative model is a type of generative model that predicts future outputs based on historical information, and can derive the output result using inputs and historical outputs.

[0003] The ability to generate answers based on a large language model allows for the construction of a language generation model system. This system can receive questions and contextual information from the dialogue user. By invoking the large language model based on the user's input and contextual information, an answer to the user's question can be obtained.

[0004] However, existing generation methods suffer from low generation efficiency and long response times. Summary of the Invention

[0005] In view of this, this application provides a data processing method for improving the efficiency of answering questions. This application also provides corresponding apparatus, computing device clusters, computer-readable storage media, and computer program products.

[0006] Firstly, this application provides a data processing method. This method can improve the efficiency of generating answers by lowering the verification criteria for drafts. Specifically, after obtaining the first input, a search can be performed on the draft library based on the first input. Multiple target drafts related to the first input are determined from multiple first candidate drafts stored in the draft library. Then, a pre-trained model can be invoked to verify the multiple target drafts, determining the matching probability of each draft token in each target draft. The pre-trained model can be a generative model with an autoregressive structure, such as a large language model. The matching probability is the probability that the output of the pre-trained model is a draft token given the first input. After obtaining the probability of each target draft output by the pre-trained model, the multiple draft tokens are verified according to the matching probability to determine the verified draft tokens. The matching probability of the verified draft tokens is greater than a target threshold. Next, based on the number of verified draft tokens in the target drafts, the target draft with the highest number of verified draft tokens among multiple target drafts can be identified as the usable target draft. Then, an answer for the first input is generated based on the verified draft tokens in the usable target drafts. In other words, the decision to retain a draft token is based on the relationship between the matching probability of the draft token and a threshold, without requiring strict consistency between the draft token and the output of the large language model. This effectively lowers the verification standard for draft tokens, thereby increasing the probability of successful verification and allowing more draft tokens to pass verification. Compared to speculative decoding techniques, this approach can identify more verified draft tokens, increasing the number of tokens available in the usable target drafts for generating the answer. Thus, a single call to the pre-trained model can identify more tokens, improving the efficiency of answer generation and reducing the response time of the language generation model system.

[0007] In some possible implementations, the target threshold can be user-configurable. Specifically, the server may include a threshold configuration interface. If a user triggers an operation to configure the target threshold on the client, the client can send a threshold configuration request to the server through the threshold configuration interface. The threshold configuration request is used to configure the target threshold. Correspondingly, the server can obtain the threshold configuration request sent by the user through the threshold configuration interface and determine the target threshold based on the threshold configuration request. In this way, the user can adjust the size of the target threshold according to actual needs, thereby validating draft tokens with a stricter or more lenient standard. Thus, by adjusting the size of the target threshold, a balance can be achieved between efficiency and accuracy, adapting to the needs of different tasks.

[0008] In some possible implementations, users may lack awareness of the target threshold size and be unable to select an appropriate one. To address this, threshold recommendation information can be provided to inform users of the relationship between the target threshold size and the task type. Specifically, at least one set of threshold recommendations can be determined first. Each set includes information on a reference threshold and the corresponding task type, indicating that the reference threshold is suitable for that task type to determine whether the draft token verification passes. The server can provide a threshold recommendation interface to the client. The client can display this interface, showing at least one set of threshold recommendations. By viewing the threshold recommendation information, users can understand the relationship between the task type and the target threshold size, and thus set an appropriate target threshold according to their needs.

[0009] In some possible implementations, the drafts in the draft library can be determined based on the context information of the input. Specifically, after obtaining the first input, the context information of the first input can be further obtained. Then, multiple first candidate drafts can be determined based on the context information of the first input, and these first candidate drafts are added to the draft library. In this way, when determining the target draft later, one (or more) first candidate drafts can be selected as the target candidate draft. Since the first candidate drafts come from the context information of the first input, there is a strong correlation between the first candidate drafts and the first input. Therefore, the probability of the draft token passing verification is relatively high among the first candidate drafts selected as the target draft. Furthermore, using the context information of the first input as the drafts in the draft library can better preserve the context information of the first input and avoid the loss of context information during compression.

[0010] In some possible implementations, if new contextual information is obtained, it can also be added to the draft library. Specifically, based on the first input, if the user inputs a second input in the same session, the contextual information of both the second and second inputs can be obtained. Based on the contextual information of the second input, multiple second candidate drafts can be obtained, and these second candidate drafts are also added to the draft library. Since the draft library includes drafts formed from the contextual information of both the first and second inputs, when selecting a target draft, it is possible to select either a fragment from the contextual information of the first input or a fragment from the contextual information of the second input. This preserves the contextual information of both the first and second inputs, avoids the loss of contextual information, and improves the accuracy of the generated answer.

[0011] In some possible implementations, the length of the context information of the first input may exceed the maximum input length of the pre-trained model. Therefore, the context information of the first input can be compressed, and compression-related information can be displayed to the user. Specifically, if the length of the context information of the first input exceeds the maximum input length of the pre-trained model, the context information of the first input can be compressed to ensure that the compressed context information can be input into the pre-trained model. To help users understand the situation before and after compression, a compression ratio display page can also be provided. The compression ratio display page is used to display the compression ratio of the first input. The compression ratio of the first input is determined based on the length of the context information of the first input before and after compression. In this way, not only can the context information be compressed, but the user can also understand the compression status of the context information.

[0012] Secondly, this application provides a data processing apparatus, the apparatus comprising:

[0013] The acquisition unit is used to acquire the first input sent by the user;

[0014] The draft determination unit is used to determine multiple target drafts that match the first input from the draft library based on the first input. The target drafts include at least one draft token.

[0015] The probability determination unit is used to call the pre-trained model to obtain the matching probability of each draft token in each target draft. The matching probability is the probability that the pre-trained model outputs the draft token when the first input is used as input.

[0016] The verification unit is used to determine the draft tokens that pass the verification. The draft tokens that pass the verification are those with a matching probability greater than the target threshold.

[0017] The available draft determination unit is used to determine the available target drafts among multiple target drafts. The number of verified draft tokens in the available target drafts is greater than that of other drafts in the multiple target drafts besides the available target drafts. The verified draft tokens in the target available drafts are used to generate an answer for the first input.

[0018] In some possible implementations, the acquisition unit is also used to acquire a threshold configuration request sent by the user through the threshold configuration interface, which is used to configure the target threshold.

[0019] In some possible implementations, the apparatus further includes a recommendation information determination unit; the recommendation information determination unit is used to determine at least one set of threshold recommendation information, each set of threshold recommendation information including task type information and reference threshold information, the threshold recommendation information being used to provide a reference in the process of determining the target threshold according to the task type corresponding to the first input; and to provide a threshold recommendation interface for displaying at least one set of threshold recommendation information.

[0020] In some possible implementations, the apparatus further includes a draft library configuration unit; an acquisition unit, further configured to acquire context information of the first input; a draft library configuration unit, configured to determine multiple first candidate drafts based on the context information of the first input; and to add the first candidate drafts to the draft library.

[0021] In some possible implementations, the acquisition unit is also used to acquire the context information of the second input and the first input, the second input and the first input belonging to the same session; the draft library configuration unit is also used to determine multiple second candidate drafts based on the context information of the second question; and to add the multiple second candidate drafts to the draft library.

[0022] In some possible implementations, the apparatus further includes a compression unit; specifically, the compression unit is configured to compress the context information of the first input in response to the fact that the length of the context information of the first input is greater than the maximum input length of the pre-trained model; input the compressed context information of the first input, the first input, and multiple target drafts into the pre-trained model; and provide a compression ratio display page, which displays the compression ratio of the first input, the compression ratio of the first input being determined based on the length of the context information of the first input before and after compression.

[0023] Thirdly, this application provides a computing device, the computing device including at least one processor and at least one memory; the at least one memory is used to store instructions, and the at least one processor executes the instructions stored in the at least one memory to cause the computing device to perform the method in the first aspect or any possible implementation thereof. It should be noted that the memory may be integrated into the processor or may be independent of the processor. The at least one computing device may also include a bus. The processor is connected to the memory via the bus. The memory may include readable storage and random access memory.

[0024] Fourthly, this application provides a computing device cluster, the computing device including at least one computing device, the at least one computing device including at least one processor and at least one memory; the at least one memory is used to store instructions, and the at least one processor executes the instructions stored in the at least one memory to cause the computing device cluster to perform the method in the first aspect or any possible implementation of the first aspect. It should be noted that the memory can be integrated into the processor or can be independent of the processor. The at least one computing device may also include a bus. The processor is connected to the memory via the bus. The memory may include readable storage and random access memory.

[0025] Fifthly, this application provides a computer-readable storage medium storing instructions that, when executed on at least one computing device, cause the at least one computing device to perform the method described in the first aspect or any implementation thereof.

[0026] In a sixth aspect, this application provides a computer program product containing instructions that, when run on at least one computing device, cause the at least one computing device to perform the method described in the first aspect or any implementation thereof.

[0027] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description

[0028] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0029] Figure 1a A schematic diagram illustrating an application scenario provided in this application embodiment;

[0030] Figure 1b This is a schematic diagram illustrating another application scenario provided in the embodiments of this application;

[0031] Figure 2 A schematic flowchart of a data processing method provided in an embodiment of this application;

[0032] Figure 3 Another flowchart illustrating the data processing method provided in this application embodiment;

[0033] Figure 4 A schematic diagram of a page displayed by the client 10 provided in an embodiment of this application;

[0034] Figure 5 This is a schematic diagram of the structure of a video generation apparatus provided in an embodiment of this application;

[0035] Figure 6 A schematic diagram of the structure of a computing device provided in an embodiment of this application;

[0036] Figure 7 This is a schematic diagram of a computing device cluster provided in an embodiment of this application;

[0037] Figure 8 This is a schematic diagram illustrating one implementation of a computing device cluster provided in an embodiment of this application. Detailed Implementation

[0038] The solutions in the embodiments provided in this application will now be described with reference to the accompanying drawings.

[0039] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application.

[0040] First, let me introduce some of the terms used in this application.

[0041] Language Generation Model System: The language generation model system provides dialogue capabilities. Dialogue tenants can input questions and contextual information through the language generation model system's client. The language generation model system's server can invoke the large language model, obtain the answer to the question based on the question and context information, and display the answer to the dialogue tenant through the client. Each invocation of the large language model generates a token. Users of the language generation model system are also referred to as dialogue tenants.

[0042] Conversation: A conversation refers to the continuous communication process between a dialogue tenant and a large language model using natural language. In a single conversation, a dialogue tenant can engage in one round of dialogue with the large language model through a language generation model system, or multiple rounds of dialogue.

[0043] Context information refers to additional information relevant to the current session. Context information can be uploaded by the conversation tenant or retrieved based on the conversation tenant's question.

[0044] Draft: A draft refers to a preliminary prediction based on the question. A draft may include one or more tokens. In a language generation model system with a draft-then-verify structure, multiple drafts can be determined based on the question. These drafts are then validated by a large language model. The tokens from the draft with the most validated tokens are selected as the output of the current step of the large model, and these tokens are used as part of the answer generation.

[0045] Draft token: A draft token is a single token in a draft. Each draft can include one or more draft tokens. During validation, it can be determined whether each draft token in the draft is the output of the large language, and the probability of passing for each draft token can be determined.

[0046] The language model generation system can invoke a large language model to generate answers to questions. Specifically, the language model generation system can send the question and its context information to the large language model through its external interface. The large language model can then synthesize the question and context information, select the most suitable token, and output it. Each time the large language model is invoked, the language model generation system receives one token. Multiple invocations can yield multiple tokens, which are then combined to form the answer and displayed to the dialogue tenant.

[0047] In the above implementations, the large language model generates one token at a time, resulting in low efficiency and long response times. To address these issues, some possible implementations employ speculative decoding techniques to improve the efficiency of the large language model, thereby reducing the response time of the language model generation system.

[0048] Specifically, in speculative decoding techniques, multiple drafts can first be generated using a draft model. Next, these drafts, along with the question and context information, can be input into a large language model. The large language model simulates the output when the input consists of the question and context information, and the output of the large language model is compared with the draft tokens in each draft to determine the number of validated draft tokens in each draft. The draft with the most validated draft tokens is identified as the usable draft, and the validated draft tokens from the usable drafts are used as tokens in the answer, thereby generating an answer to the question.

[0049] For example, in retrieval-based speculative decoding (REST) ​​technology, a general corpus can be pre-built as a draft library. When a draft needs to be generated, it can be searched in the draft library based on the input question to determine multiple draft tokens. Multiple drafts can be obtained through different search methods. These multiple drafts, the question, and the question's context information are then input into a large language model. The large language model can determine whether the input is a question and its context information, and output whether it is a specific draft token. In this way, it can be determined whether each draft token passes verification, the number of verified draft tokens in each draft can be determined, and the draft with the most verified draft tokens is identified as the usable draft. The verified draft tokens from the usable drafts are then identified as the tokens in the answer.

[0050] In speculative decoding techniques, large language models can perform parallel verification of multiple drafts at once. This allows for the verification of multiple drafts with a single call to the large language model, improving the efficiency of draft verification and consequently increasing the efficiency of the language model generation system in generating responses.

[0051] However, the performance of traditional speculative decoding techniques is still insufficient, and language generation model systems still suffer from slow response times.

[0052] On the one hand, speculative decoding techniques have strict verification of draft tokens. Only when the draft token is completely consistent with the output of the large language model is it considered to be the output of the large language model. In this way, only a small number of draft tokens pass verification each time the large language model is called, resulting in low efficiency and long response time.

[0053] On the other hand, speculative decoding techniques are highly dependent on draft models. If the draft model is an artificial intelligence model, it will increase the load and response time of the language generation model system. If the draft model is an information retrieval model, a dedicated draft library needs to be built so that the information retrieval model can retrieve draft tokens from the draft library.

[0054] Large language models often have a maximum input length. Content exceeding this maximum length cannot be processed by the large language model. Therefore, for language generation model systems, when the question and its context are too long, the input to the large language model needs to be compressed. Generally, the context information can be compressed to ensure that the total length of the question and the compressed context information does not exceed the maximum input length of the large language model. However, the answer is based on the content within the context information, and compressing the context information will result in the loss of this content. Therefore, compressing the context information can lead to inaccurate generated answers and reduce the quality of the answers.

[0055] Based on this, this application provides a data processing method. This method can improve the efficiency of generating answers by reducing the validation criteria for drafts. Specifically, after obtaining the first input, a search can be performed on the draft library based on the first input. Multiple target drafts related to the first input are determined from multiple first candidate drafts stored in the draft library. Then, a pre-trained model can be invoked to validate the multiple target drafts and determine the matching probability of each draft token in each target draft. The pre-trained model can be a generative model with an autoregressive structure, such as a large language model. The matching probability is the probability that the output of the pre-trained model is a draft token given the first input. After obtaining the probability of each target draft output by the pre-trained model, the multiple draft tokens are validated according to the matching probability to determine the validated draft tokens. The matching probability of the validated draft tokens is greater than a target threshold. Next, based on the number of verified draft tokens in the target drafts, the target draft with the highest number of verified draft tokens among multiple target drafts can be identified as the usable target draft. Then, an answer for the first input is generated based on the verified draft tokens in the usable target drafts. In other words, the decision to retain a draft token is based on the relationship between the matching probability of the draft token and a threshold, without requiring strict consistency between the draft token and the output of the large language model. This effectively lowers the verification standard for draft tokens, thereby increasing the probability of successful verification and allowing more draft tokens to be verified. Compared to speculative decoding techniques, this approach can identify more verified draft tokens, increasing the number of tokens available in the usable target drafts for generating the answer. Thus, a single call to the pre-trained model can identify more tokens, improving the efficiency of answer generation and reducing the response time of the language generation model system.

[0056] Next, various non-limiting specific implementation methods of the data processing procedure will be described in detail.

[0057] First, an exemplary application scenario is introduced. The data processing method provided in this application can be applied to the server side, which can be the server side of a language generation model system.

[0058] Refer to Figure 1a , Figure 1a which is a schematic diagram of an application scenario of the data processing method provided in the embodiment of this application. In the Figure 1a application scenario shown, it includes a client 10, a server 20, and a pre-trained model 30. Among them, the client 10 can be a client of a language generation model system, and the server 20 can be a server of a generation model system. The server 20 includes a data processing device 21, a draft library 22, and an answer generation device 23. The data processing device 21 includes an acquisition unit, a draft determination unit, a probability determination unit, a verification unit, and an available draft determination unit.

[0059] Specifically, the client 10 can run on the terminal device of the dialogue tenant. The dialogue tenant using the client 10 can input a first input on the client 10 and can also input context information related to the first input. The client 10 can send the first input and the context information to the server 20 through the network. The acquisition unit can acquire the first input. The draft determination unit can retrieve from the draft library 22 according to the first input and determine multiple target drafts related to the first input from the draft library 22. The probability determination unit can call the pre-trained model 30 through the interface of the pre-trained model 30 to verify the multiple target drafts and determine the matching probability of each draft token. The verification unit can determine the draft tokens that pass the verification. The available draft determination unit can determine the available target drafts. The data processing device 21 can send the available target drafts or the draft tokens that pass the verification in the available target drafts to the answer generation device 23. The answer generation device 23 can generate answer information according to the draft tokens that pass the verification in the available target drafts. After generating the answer, the answer generation device 23 can generate an answer according to the available information and return it to the client 10. The client 10 can display the answer to the first input to the dialogue tenant to complete the dialogue with the dialogue tenant.

[0060] In an actual application scenario, the client 10 can be a cloud service client, and the dialogue tenant can call a cloud service based on a pre-trained model, so as to have one or more rounds of conversations with the pre-trained model through the cloud service. The server 20 is a model supporting the pre-trained model 30, used for data preprocessing and decoding, and realizes the conversation with the dialogue tenant by calling the pre-trained model 30. Optionally, the server 20 and the pre-trained model 30 can be implemented based on different computing devices or computing device clusters. For example, the server 20 can run on device cluster A, and the pre-trained model 30 can run on device cluster B. The probability determination unit calls the pre-trained model 30 through the network interface to verify the target drafts.

[0061] It should be noted that the above application scenarios are merely examples, and the data processing method provided in this application embodiment can be applied to any application scenario that calls a pre-trained model for dialogue. For example, in some other possible application scenarios, the above data processing device can also be integrated into the client.

[0062] The following section provides a detailed introduction to the specific implementation methods in the data processing process.

[0063] See Figure 2 , Figure 2 A flowchart illustrating the data processing method provided in this application. This method can be applied to... Figure 1a The application scenarios shown can also be applied to other applicable application scenarios.

[0064] Specifically, Figure 2 The data processing methods shown may specifically include:

[0065] S201: Get the first input sent by the user.

[0066] In this embodiment, the data processing device can first obtain the first input from the user. The user can be the aforementioned dialogue tenant. Optionally, if Figure 2 The data processing method shown is implemented by the server, which can receive the first input sent by the client over the network.

[0067] Optionally, the first input may include a question raised by the user, and may also include contextual information about the question. Contextual information refers to information associated with the first input, such as historical dialogue information about the question and search information obtained based on the question.

[0068] Historical dialogue information refers to the historical questions and responses between the dialogue tenant and the language generation model system in the current conversation before the first input. In other words, if the dialogue tenant and the language generation model system have conducted multiple rounds of dialogue in a certain conversation, when the dialogue tenant raises a new question, the existing questions and answers in that conversation can be used as contextual information for the newly raised question.

[0069] The retrieved information is obtained based on the question input by the user. Specifically, after determining the question raised by the dialogue user, the language generation model system can search based on the question, identifying relevant information as context information in the first input. Specifically, the language generation model system can call a search engine via an interface, using the question and / or keywords from the question as search keywords to retrieve information, and use the retrieved results as context information. Optionally, the language generation model system can perform information retrieval on the web based on search keywords, identifying one or more web pages related to the question, and extracting text information from these web pages as context information in the first input.

[0070] S202: Based on the first input, determine multiple target drafts from the draft library that match the first input.

[0071] After obtaining the first input, a search can be performed in the draft library to identify multiple target drafts that match the first input. The draft library includes multiple first candidate drafts, and the target drafts are selected from these first candidate drafts. In other words, after obtaining the first input, multiple first candidate drafts that match the first input can be selected as the target drafts.

[0072] In this embodiment of the application, a draft refers to a set of tokens. That is, a first candidate draft may include one token or multiple tokens. Tokens in the draft may be referred to as draft tokens.

[0073] When searching from the draft library, the first candidate draft related to the first input can be selected as the target draft.

[0074] Optionally, retrieval can be performed using similarity matching. For example, the similarity between each first candidate draft and the first input can be calculated, and the first candidate draft with the highest similarity can be selected as the target draft. Alternatively, the first input can be split into multiple tokens, the target token can be determined from these tokens, and a matching process can be performed on the drafts based on the target token, selecting the first candidate draft that includes the target token as the target draft. For instance, the last k tokens of the first input can be selected as the target token. k can be a pre-configured positive integer.

[0075] Alternatively, a search can be conducted from a draft library based on semantic matching. For example, the draft determination unit can invoke a semantic understanding model to perform semantic analysis on the first input, and search the draft library based on the semantics of the first input, identifying the first candidate draft whose semantics match the first input as the target draft.

[0076] The above describes how to retrieve a target draft from a draft library. Below, we will introduce some ways to obtain a draft library.

[0077] In some possible implementations, the draft library can be pre-defined. Specifically, a corpus can be pre-acquired, and the corpus can be split into multiple phrases, with each phrase stored as a first-choice draft in the draft library.

[0078] In the above implementation, the draft library is derived from the corpus, and different sessions can use the same draft library without additional configuration. However, because a general corpus is used as the draft library, the matching degree between the first candidate draft and the first input is limited. Thus, in some application scenarios, the target draft determined from the draft library may have a low matching degree with the first input. Consequently, the matching degree between the available information determined based on the target draft and the first input may also be relatively poor, and the answer obtained based on the available information may not be able to answer the first input.

[0079] Therefore, in some other possible implementations, the draft library can be determined based on the context information of the first input. Specifically, after obtaining the context information of the first input, this context information can be split into multiple first candidate drafts, thus obtaining the draft library. In this way, the first candidate drafts in the draft library are determined based on the context information of the first input and have a high degree of correlation with the first input. Therefore, determining the target draft from the first candidate drafts and thus determining the available information can ensure the matching degree between the available information and the first input, thereby improving the ability to answer the first input.

[0080] Specifically, the context information of the first input can be split into multiple first candidate drafts based on a preset number of tokens. For example, each first candidate draft can be pre-configured to have a length of n tokens. Then, the context information of the first input can be split, with each n tokens forming a first candidate draft. Here, n is a positive integer greater than 1, representing the number of draft tokens in the candidate drafts. Optionally, the context information of the first input may include separators, such as punctuation marks or other semantically distinguishing elements. When splitting the context information of the first input, if a separator is detected, the tokens before the separator can be grouped into a first candidate draft, and the token after the separator can be used as the starting point for the next first candidate draft.

[0081] Alternatively, when creating the draft library, the context information of the first input can be used as the draft library without splitting it. Furthermore, the length of the target draft can be pre-configured, for example, it can be configured to have m tokens. When determining the target draft, a window of length m (where m is a positive integer) can be created, and a sliding window method can be used to find multiple sentence segments of length m that match the first input within the context information of the first input, and these segments can be used as the target draft.

[0082] In real-world applications, dialogue tenants may engage in multi-turn dialogues with the language generation model. Within the same session, the contextual information for different dialogues may differ. If new contextual information is introduced in a dialogue, it can be added to the draft library. That is, suppose that after answering the first input, the dialogue tenant raises a second question in the session corresponding to the first input, and the language generation model retrieves new contextual information based on the second question. Then, multiple second candidate drafts can be determined based on the contextual information of the second question, and these second candidate drafts can also be added to the draft library. Thus, when determining the target draft for the second question, a candidate draft associated with the second question can be found from the multiple first and second candidate drafts and used as the target draft for the second question. Storing the contextual information of each question in the session in the draft library not only allows for the generation of new candidate drafts based on new contextual information but also preserves the contextual information of historical questions in the session, preventing the loss of contextual information.

[0083] S203: Call the pre-trained model to obtain the matching probability of each draft token in each target draft.

[0084] After identifying multiple target drafts, a pre-trained model can be invoked to validate these drafts and determine the matching probability of each draft token. The pre-trained model can be a pre-trained large language model, and the matching probability of a draft token can be the probability that the actual output of the pre-trained model is a target draft. As mentioned earlier, each target draft includes one or more draft tokens; during validation, the matching probability can be calculated for each draft token in each target draft.

[0085] Optionally, if the pre-trained model is a generative model with an autoregressive structure, multiple target drafts can be validated in a single call to the pre-trained model. When validating target drafts, multiple target drafts, a first input, and context information of the first input can be input to the pre-trained model. The pre-trained model can simulate the actual output when the input is the first input and its context information, and obtain the probability of each draft token as the output. By receiving the information returned by the pre-trained model, the matching probability of each draft token can be obtained. Thus, multiple target drafts can be validated in a single call to the pre-trained model.

[0086] In some application scenarios, the length of the context information of the first input may exceed the maximum input length of the pre-trained model, making it impossible to completely input the context information of the first input into the pre-trained model. Therefore, the context information of the first input can be compressed. Optionally, the context information of the first input can be compressed according to the maximum input length of the pre-trained model to ensure that the compressed context information is less than the maximum input length of the pre-trained model, thereby ensuring that both the first input and the compressed context information can be input into the pre-trained model. Optionally, the context information of the first input can be compressed using an encoder-structured model.

[0087] In some application scenarios, the dialogue tenant may be particularly concerned about the compression of context information. Therefore, some implementations can calculate the compression rate of the context information and display it to the dialogue tenant.

[0088] Specifically, after compressing the context information of the first input, the compression ratio of the first input can be calculated based on the length of the context information of the first input before compression and the length of the context information of the first input after compression. Furthermore, the compression ratio of the first input can be displayed to the dialogue tenant corresponding to the first input through the client, so that the dialogue tenant of the first input understands that the context information of the first input has been compressed.

[0089] In the above implementation, the compression ratio of the first input is determined after the context information is compressed. It is the compression ratio of the context information of the first input, and can also be called the actual compression ratio of the first input.

[0090] In some application scenarios, the dialogue tenant may want to know the context information of the first input before compression. For example, the dialogue tenant might manually delete some content from the context information when the compression ratio is already high. Therefore, the dialogue tenant needs to know the compression ratio of the context information in advance. Alternatively, after determining the context information of the first input, the possible compression ratio of the first input can be calculated and displayed.

[0091] Specifically, after determining the context information of the first input, the compression ratio of the first input's context information under the current input can be calculated based on the length of the first input's context information and the maximum input length of the pre-trained model. Since the first input's context information has not yet been compressed at this point, this compression ratio can also be called the simulated compression ratio. By displaying the simulated compression ratio on the client side, the dialogue tenant can understand the extent to which the first input's context information will be compressed, thereby determining whether adjustments to the first input's context information are needed. If not, subsequent processes can continue, and the first input's context information will be compressed when validating the target draft.

[0092] It is understandable that compressing the context information of the first input inevitably results in some content loss. Therefore, the higher the compression rate of the first input, the more content will be lost from its context information. This content loss will lead to distortion of the verification results. To compensate for the distortion of verification results, some of the aforementioned implementation methods can establish a draft library based on the context information before compression. In this way, the target draft comes from lossless context information, and even if the context information is compressed during the verification process, no content loss will occur, thus avoiding distortion of the verification results.

[0093] Optionally, to display the compression ratio of the first input to the user, the server can provide a compression ratio display page. The compression ratio display page includes displaying the compression ratio of the first input. Specifically, the server can generate the front-end code for the compression ratio display page based on a preset front-end page architecture and the first input compression ratio, and provide the corresponding front-end code to the client. The client can receive the front-end code from the server, render it, obtain a visual representation of the compression ratio display page, and display it to the user.

[0094] S204: Confirm the verified draft token.

[0095] After obtaining the matching probability of each draft token, each target draft can be verified based on the matching probability to determine the verified draft tokens. Validated draft tokens are those with a matching probability greater than a target threshold. Optionally, the target threshold is a positive number less than 1.

[0096] Specifically, we can check one by one whether the matching probability of each draft token is greater than the target threshold. If the matching probability of a draft token is greater than the target threshold, the draft token can be considered to have passed the verification, that is, the verification result of the draft token is successful.

[0097] In the above process, the validation of a draft token is determined based on whether its matching probability is greater than a target threshold, rather than on the actual output of the pre-trained model. This threshold-based validation filters out draft tokens with low matching probabilities, ensuring the accuracy of the answer. Furthermore, using the target threshold allows multiple draft tokens to pass validation, increasing the number of valid draft tokens and improving the efficiency of determining usable information. This approach ensures that more target drafts pass validation, increasing the length of usable information obtained with each call to the pre-trained model, thereby improving the efficiency of answer generation and reducing the response time of the language generation model system.

[0098] Optionally, the target threshold mentioned above can be a pre-configured threshold or a threshold selected by the dialogue tenant based on the actual situation. Specifically, the dialogue tenant can configure an appropriate target threshold for the first input based on the content of the dialogue. For example, if the dialogue tenant has high requirements for the accuracy of the first input, the target tenant can choose a higher target threshold to filter the draft tokens more strictly. If the target tenant has lower requirements for the accuracy of the first input but higher requirements for the response speed of the language generation model system, then a lower target threshold can be selected to increase the length of usable information determined each time. Optionally, the dialogue tenant can configure the target threshold through the client. Further details on this can be found below and will not be repeated here.

[0099] S205: Identify the available target drafts among multiple target drafts.

[0100] After obtaining the verification result of the draft token, the usable target draft can be determined from multiple target drafts based on the verification result. Specifically, the number of verified draft tokens in the usable target draft is greater than the number of verified draft tokens in the other target drafts. In other words, the usable target draft is the one with the largest number of verified draft tokens among the multiple target drafts.

[0101] The number of draft tokens in the verification results of each target draft can be counted, and the target draft with the most verified draft tokens is identified as the usable target draft.

[0102] After identifying the available target drafts, an answer to the first input can be generated based on the draft tokens in the available target drafts. Specifically, an answer to the first input can be generated based on the verified draft tokens in the available target drafts. Specifically, the verified draft tokens in the available target drafts can be identified as tokens in the answer. Thus, by repeatedly executing steps S202-S204, multiple tokens in the answer can be obtained, generating an answer to the first input.

[0103] This application provides a data processing method. This method can improve the efficiency of generating answers by lowering the validation criteria for drafts. Specifically, after obtaining the first input, a search can be performed on the draft library based on the first input. Multiple target drafts related to the first input are determined from multiple first candidate drafts stored in the draft library. Next, a pre-trained model can be invoked to validate the multiple target drafts, determining the matching probability of each draft token in each target draft. The pre-trained model can be a generative model with an autoregressive structure, such as a large language model. The matching probability is the probability that the output of the pre-trained model is a draft token given the first input. After obtaining the probability of each target draft output by the pre-trained model, the multiple draft tokens are validated according to the matching probability to determine the validated draft tokens. The matching probability of the validated draft tokens is greater than a target threshold. Then, based on the number of validated draft tokens in the target drafts, the target draft with the most validated draft tokens among the multiple target drafts can be determined as the available target draft, and an answer for the first input can be generated based on the validated draft tokens among the available target drafts. In other words, the decision to retain a draft token is based on the relationship between its matching probability and a threshold, without requiring strict consistency between the draft token and the output of the large language model. This effectively lowers the validation threshold for draft tokens, thereby increasing their probability of passing validation and allowing more draft tokens to be validated. Compared to speculative decoding techniques, this approach can identify more draft tokens as valid, increasing the number of tokens available for generating answers from the target drafts. Thus, a single call to the pre-trained model can identify more tokens, improving the efficiency of answer generation and reducing the response time of the language generation model system.

[0104] The following section provides a detailed description of the data processing method provided in this application embodiment, using the application scenario of multiple dialogues between the dialogue tenant and the language generation model system.

[0105] First, combine Figure 1b The application scenarios will be introduced.

[0106] exist Figure 1a On this basis, Figure 1b Data processing device 21 ( Figure 1b (Not shown) It also includes an information retrieval unit, a draft library configuration unit, and a compression unit. The information retrieval unit is used to retrieve the context information of the first input, the draft library configuration unit is used to determine the draft library based on the context information of the first input, and the compression unit is used to compress the context information of the first input. Figure 1b For introductions to the remaining units, please refer to [link / reference]. Figure 1a I won't go into details here.

[0107] See Figure 3 , Figure 3 A flowchart illustrating the data processing method provided in this application. This method can be applied to... Figure 1b The application scenarios shown. Specifically, Figure 3 The data processing methods shown may specifically include:

[0108] S301: Client 10 obtains the first input and the target threshold.

[0109] If a dialogue tenant needs to interact with the language generation model system, the dialogue tenant can input questions and contextual information into client 10, and can also configure a target threshold through client 10. Client 10 can obtain the first input and the target threshold from the user, and send the first input and the target threshold to server 20. In this embodiment, the information configured by the dialogue tenant through client 10 for inputting into the large language model can be referred to as the first input.

[0110] Specifically, client 10 can display input controls, such as input boxes. The user can input a question through the input box. The user-inputted question can be provided to server 20 as the first input. Optionally, client 10 can also be used to input context information for the question. For example, client 10 can display a file upload control, which the target user can use to upload a file. The content of the file can serve as context information for the question. Optionally, the file can include one or more of the following: text files, image files, code files, and online document files.

[0111] The client 10 can also configure a target threshold. Specifically, the client 10 can display a threshold configuration control for configuring the target threshold. This allows users to adjust the target threshold according to actual needs, thereby validating draft tokens with stricter or more lenient standards. Thus, by adjusting the target threshold, a balance can be struck between efficiency and accuracy to suit the needs of different tasks.

[0112] Optionally, the threshold configuration control may include an input field where the dialog tenant can enter a target threshold. And / or, the threshold configuration control may also include a slider control. The slider control includes a virtual slider and a virtual slider rail. The relative position of the virtual slider on the virtual slider rail represents the size of the target threshold. The dialog tenant can adjust the target threshold size by dragging the virtual slider to change its position on the virtual slider rail.

[0113] Considering that some users may not understand the impact of matching probability on accuracy and therefore cannot set an appropriate target threshold, client 10 can also display at least one set of threshold recommendation information. Specifically, client 10 can display a threshold recommendation interface to the user. The threshold recommendation interface displays at least one set of threshold recommendation information. Each set of threshold recommendation information includes task type information and reference threshold information. Within the same set of threshold recommendation information, the task type corresponds to the reference threshold. The reference threshold applies to tasks of that task type. (The last sentence appears to be a separate, unrelated statement about task types.)

[0114] For example, common speech generation model systems might correspond to task types such as question answering, code completion, and summary generation. Among these, the "question answering" task has the highest accuracy requirement, the "chat" task has the lowest accuracy requirement, and the "summary generation" task has a moderate accuracy requirement. Therefore, the reference threshold for the "question answering" task is higher than that for the "summary generation" task, and the reference threshold for the "summary generation" task is higher than that for the "chat" task.

[0115] Optionally, the threshold recommendation information can be provided by the server 20 to the client 10. Specifically, the server 20 can first determine at least one set of threshold recommendation information. For example, the server 20 can read at least one set of pre-stored threshold recommendation information from a database. Alternatively, the server 20 can determine the possible dialogue purpose of the dialogue tenant based on the dialogue tenant's historical dialogue behavior, and thus determine at least one set of threshold recommendation information based on the dialogue purpose. Then, the server 20 can provide a threshold recommendation interface to the client 10, so that the client 10 can display at least one set of threshold recommendation information to the dialogue tenant through the threshold recommendation interface.

[0116] Optionally, the threshold recommendation interface provided by the server 20 may include the server 20 generating front-end code for the threshold recommendation interface and sending it to the client 10. The client 10 can render a visual representation of the threshold recommendation interface based on the front-end code and display it to the dialogue tenant, thereby achieving the effect of displaying threshold recommendation information to the dialogue tenant.

[0117] In some possible implementations, the dialogue tenant can select a pre-trained model. That is, the pre-trained model 30 can be configured by the dialogue tenant through the client 10. Specifically, the client 10 can display a model configuration control to the dialogue tenant. The model configuration control may include, for example, a drop-down menu or an input box. The drop-down menu includes multiple selectable models. The dialogue tenant can select the pre-trained model 30 through the drop-down menu. Alternatively, the dialogue tenant can also enter the address of the pre-trained model 30 in the input box.

[0118] As mentioned earlier, the contextual information of the problem may be compressed. Figure 4 In the implementation shown, compression-related information can also be displayed on the client 10. Specifically, the client 10 can display the maximum input length of the pre-trained model and the length of the current context information. Furthermore, the client 10 can also display the length of the compressed context information and the compression ratio of the context information.

[0119] In some possible implementations, the dialogue tenant can configure a standard format for the response. Specifically, the client can include an input field. The dialogue tenant can enter a standard response in the input field. When generating a response, the standard response entered by the dialogue tenant can be referenced to generate a response with the same or similar format.

[0120] Optionally, client 10 can display as follows: Figure 4 Page 40 is shown. Figure 4 In the implementation shown, the page 40 viewed by the dialogue tenant includes a question input control 41, a model selection control 42, a compressed information display area 43, a threshold configuration control 44, and a standard answer configuration control 45.

[0121] The question input control 41 includes an input box. The dialog tenant can input the first input through the input box of the question input control 41, and can also input the context information corresponding to the first input.

[0122] The model selection control 42 includes a drop-down menu, through which the dialogue tenant can select a pre-trained model 30 as the model for generating the response to the first input.

[0123] The compression information display area 43 is used to display the maximum input length of the pre-trained model 30, the current input length of the pre-trained model, the compressed input length, and the compression ratio. Figure 4 In the implementation shown, the maximum input length of the pre-trained model 30 is 4096 tokens, the current input length is 8192 tokens, and the compressed input length is 2867 tokens, with a compression rate of 0.35. Optionally, the current input length can be the length of the content entered by the dialogue tenant in the question input control 41, or it can be the sum of the length of the content entered by the dialogue tenant in the question input control 41 and the total length of the context information retrieved in step S302.

[0124] The threshold configuration control 44 includes a sliding configuration control 441 and a recommendation information display area 442. The sliding configuration control 441 includes a virtual slider 4411 and a virtual slider 4412. When the virtual slider 4412 is at the leftmost position of the virtual slider 4411, the target probability is set to 0. When the virtual slider 4412 is at the rightmost position of the virtual slider 4411, the target probability is set to 1. The recommendation information display area 442 displays three sets of threshold recommendation information: a reference threshold of 0.1 for the task type "Question and Answer," a reference threshold of 0.2 for the task type "Summary Generation," and a reference threshold of 0.05 for the task type "Chat." The user can select an appropriate threshold based on the actual situation of the first input and set the target threshold by moving the position of the virtual slider 4411.

[0125] The standard answer configuration control 45 includes an input box for entering a standard answer. The dialog tenant can enter a standard answer in the input box to standardize the format of the answer generated by the language generation model system.

[0126] It should be noted that, Figure 4 As an example only, in real-world applications, the client may display more or less content. Furthermore, the aforementioned question input control 41, model selection control 42, compressed information display area 43, threshold configuration control 44, and standard answer configuration control 45 can be displayed on one page or multiple pages.

[0127] After the client 10 configures the first input and target threshold, the client 10 can send the user-configured information to the server 20 via the network. The acquisition unit in the server 21 can obtain the first input and target threshold. After obtaining the first input and target threshold, the acquisition unit can send the first input and target threshold to each unit in the data processing device 21.

[0128] It is understandable that each control provided by client 10 corresponds to the configuration interface between client 10 and server 20. That is, if a client tenant performs a certain configuration through client 10, the client tenant's configuration request or the specific information configured by the client tenant can be sent to server 20 through the interface between client 10 and server 20, so that server 20 can obtain the information configured by the client tenant.

[0129] For example, if a tenant configures a target threshold through client 10, client 10 can send the target threshold to server 20 through the threshold configuration interface, or send a threshold configuration request to server 20 through the threshold configuration interface. Server 20 can determine the target threshold based on the threshold configuration request.

[0130] S302: The information retrieval unit performs a retrieval based on the first input and determines the context information of the first input.

[0131] After receiving the first input, the acquisition unit can provide the first input to the information retrieval unit. The information retrieval unit can then perform a retrieval based on the first input to obtain the context information of the first input.

[0132] For example, in some implementations, the information retrieval unit can call the search interface of a search engine, using the first input as a search keyword to retrieve multiple web pages related to the first input. The information retrieval unit can then select web pages with high relevance to the first input from these multiple pages and use the content of these web pages as contextual information for the first input.

[0133] For example, in some other implementations, the information retrieval unit can also retrieve information from the knowledge base based on the first input, and use the retrieved content as context information for the first input. The knowledge base can be a local knowledge base of the server 20 or a remote knowledge base.

[0134] Optionally, after obtaining the context information of the first input, the input length and compression ratio displayed by the client 10 can be updated.

[0135] S303: The draft library configuration unit determines multiple first candidate drafts based on the context information of the first input, and adds the multiple first candidate drafts to the draft library.

[0136] After retrieving the context information of the first input, the draft library configuration unit can determine multiple first candidate drafts based on the context information of the first input and add the first candidate drafts to the library. When determining the first candidate drafts, the context information of the first input can be divided according to a preset length to obtain multiple first candidate drafts.

[0137] If the first input is the first question in the current session, the draft library configuration unit can create a new draft library for the current session, determine multiple first candidate drafts based on the context information of the first input, and add the multiple first candidate drafts to the draft library. If the first input is not the first question in the current session, the draft library configuration unit can determine the corresponding draft library based on the session to which the first input belongs, and add the multiple first candidate drafts to the draft library corresponding to the session to which the first input belongs.

[0138] S304: The draft determination unit determines multiple target drafts from the draft library based on the first input.

[0139] After configuring the draft library, the draft determination unit can determine the target draft based on the first input. For details on this part, please refer to the above text; it will not be repeated here.

[0140] S305: The compression unit compresses the context information of the first input.

[0141] If the context information of the first input is too long, exceeding the maximum input length of the pre-trained model 30, the context information of the first input can be compressed. Specifically, a compression unit can be invoked to compress the context information of the first input to ensure that the total length of the content input to the pre-trained model 30 does not exceed the maximum input length of the pre-trained model 30.

[0142] S306: The probability determination unit sends the first input, compressed context information and multiple target drafts to the pre-trained model 30, and receives the matching probability of each draft token returned by the pre-trained model.

[0143] The probability determination unit can invoke the pre-trained model 30 to validate multiple target drafts. Optionally, the probability determination unit can construct a prompt based on the first input, compressed context information, and multiple target drafts, and send the prompt to the pre-trained model 30. The pre-trained model 30 can receive the prompt and, based on the prompt, calculate the probability of each target draft when the input is the first input and compressed context information, outputting the probability of matching each draft token and returning it to the probability determination unit.

[0144] For a detailed explanation of how to determine the matching probability, please refer to the above text, which will not be repeated here.

[0145] S307: The verification unit determines the draft token that has passed verification.

[0146] S308: The available draft determination unit determines the available target draft among multiple target drafts.

[0147] After obtaining the matching probability of each draft token through the probability determination unit, the verification unit can determine the verified draft tokens, and the available draft determination unit can determine the available target drafts from multiple target drafts based on the number of verified draft tokens. For a detailed explanation of this part, please refer to the above text; it will not be repeated here.

[0148] After determining the target draft, the available draft determination unit can send the verified draft tokens in the available target draft to the response generation device 23, or it can send the available target draft and the matching probability of each draft token in the available target draft to the response generation device 23.

[0149] S309: The answer generation device 23 generates an answer to the first input based on the draft token that has been verified in the available target draft.

[0150] After obtaining the available information, the response generation device 23 can generate a response to the first input based on the draft token that has been verified in the available target draft.

[0151] Specifically, the response generation device 23 can determine whether the verified draft token in the available drafts can generate a complete response to the first input. If not, the data processing device can be invoked to re-execute steps S304-S308 to determine a new available target draft. If yes, a response to the first input can be generated based on the verified draft token in the available target drafts. Optionally, if the dialogue tenant has configured a standard response on the client 10, the response generation device 23 can generate a response to the first input by referring to the format of the standard response.

[0152] S310: Client 10 displays the response to the first input.

[0153] After the response generation device 23 receives the response to the first input, it can return the response to the first input to the client 10. The client 10 can then display the response to the first input to the dialogue tenant. Thus, by filtering the target draft using target probability, each call to the pre-trained model 30 can determine more available information, improving the efficiency of response generation.

[0154] Optionally, the efficiency of generating an answer when using a target threshold to filter the target draft and the efficiency of generating an answer when not using a target threshold to filter the target draft can be calculated separately, and the efficiency can be displayed to the dialogue tenant so that the dialogue tenant understands the efficiency improvement.

[0155] If the dialogue tenant needs to conduct further dialogue, the dialogue tenant can input further questions in client 10. Optionally, the dialogue tenant can also adjust the target threshold. Accordingly, client 10 can obtain the second input and the adjusted target threshold, and send the second input and the adjusted target threshold to server 20 via the network. The acquisition unit can obtain the second input and the adjusted target threshold. The information retrieval unit can retrieve the context information of the second question based on the second question. The draft library configuration unit can determine multiple second candidate drafts based on the context information of the second question, and add the multiple second candidate drafts to the draft library. The draft determination unit can determine multiple target drafts corresponding to the second question from the draft library based on the second question. The probability determination unit can call the pre-trained model 30 to verify the multiple target drafts corresponding to the second question based on the context information of the second question and the compressed second question, and obtain the matching probability of each draft token corresponding to the second question. The verification unit can verify each draft token based on the adjusted target threshold, and judge whether each draft token is a verified draft token according to the new standard. The available draft determination can determine new available target drafts so that the answer to the second input can be obtained based on the new available target drafts.

[0156] In other words, within a multi-turn dialogue within the same session, each turn can determine new candidate drafts based on the context information of the question and add them to the draft library. Furthermore, candidate drafts in the draft library are not deleted until the dialogue ends. Thus, the context information of the question can be preserved through the draft library, preventing the loss of contextual information.

[0157] This application also provides a data processing apparatus. The data processing apparatus can be applied to... Figure 1a or Figure 1b The server 20 in the implementation shown. Specifically, as... Figure 5 As shown, the data processing apparatus 500 includes:

[0158] Acquisition unit 510 is used to acquire the first input sent by the user;

[0159] The draft determination unit 520 is configured to determine, based on the first input, a plurality of target drafts matching the first input from the draft library, wherein the target drafts include at least one draft token;

[0160] The probability determination unit 530 is used to call a pre-trained model to obtain the matching probability of each draft token in each target draft, wherein the matching probability is the probability that the pre-trained model outputs the draft token when the first input is used as input;

[0161] Verification unit 540 is used to determine the verified draft token, wherein the verified draft token is the draft token whose matching probability is greater than the target threshold;

[0162] Available draft determination unit 550 is used to determine available target drafts among the plurality of target drafts, wherein the number of verified draft tokens in the available target drafts is greater than that in other drafts among the plurality of target drafts besides the available target drafts, and the verified draft tokens in the available target drafts are used to generate an answer to the first input.

[0163] The acquisition unit 510, draft determination unit 520, probability determination unit 530, verification unit 540, and available draft determination unit 550 can all be implemented in software or in hardware. For example, the implementation of the draft determination unit 520 will be described below. Similarly, the implementation of the acquisition unit 510, probability determination unit 530, verification unit 540, and available draft determination unit 550 can refer to the implementation of the draft determination unit 520.

[0164] As an example of a software functional unit, the draft determination unit 520 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the aforementioned computing instance may be one or more. For example, the draft determination unit 520 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.

[0165] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0166] As an example of a hardware functional unit, the draft determination unit 520 may include at least one computing device, such as a server. Alternatively, it may be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0167] The draft determination unit 520 includes multiple computing devices that can be distributed in the same region or in different regions. Similarly, the draft determination unit 520 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the draft determination unit 520 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0168] It should be noted that, in other embodiments, the acquisition unit 510 is used to execute any step in the data processing method, the draft determination unit 520 is used to execute any step in the data processing method, the probability determination unit 530 is used to execute any step in the data processing method, the verification unit 540 is used to execute any step in the data processing method, and the usable draft determination unit 550 can be used to execute any step in the data processing method. The steps implemented by the acquisition unit 510, the draft determination unit 520, the probability determination unit 530, the verification unit 540, and the usable draft determination unit 550 can be specified as needed. Thus, the acquisition unit 510, the draft determination unit 520, the probability determination unit 530, the verification unit 540, and the usable draft determination unit 550 respectively implement all the functions of the data processing device based on different steps in the data processing method.

[0169] This application also provides a computing device. For example... Figure 6As shown, the computing device 100 includes a bus 102, a processor 104, a memory 106, and a communication interface 108. The processor 104, the memory 106, and the communication interface 108 communicate with each other via the bus 102. The computing device 100 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 100.

[0170] Bus 102 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 6 The bus 102 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 102 may include a path for transmitting information between various components of the computing device 100 (e.g., memory 106, processor 104, communication interface 108).

[0171] The processor 104 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0172] Memory 106 may include volatile memory, such as random access memory (RAM). Processor 104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0173] The memory 106 stores executable program code, which the processor 104 executes to implement the functions of the aforementioned acquisition unit 510, draft determination unit 520, probability determination unit 530, and available draft determination unit 550, thereby realizing the data processing method. That is, the memory 106 stores instructions for executing this storage method.

[0174] Alternatively, the memory 106 stores executable code, which the processor 104 executes to implement the functions of the aforementioned data processing device, thereby implementing the data processing method. That is, the memory 106 stores instructions for executing the data processing method.

[0175] The communication interface 108 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 100 and other devices or communication networks.

[0176] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0177] like Figure 7 As shown, the computing device cluster includes at least one computing device 100. The memory 106 of one or more computing devices 100 in the computing device cluster may store the same instructions for executing data processing methods.

[0178] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing data processing methods. In other words, a combination of one or more computing devices 100 can jointly execute instructions for performing basic data processing methods.

[0179] It should be noted that the memory 106 in different computing devices 100 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the data processing device 500. That is, the instructions stored in the memory 106 of different computing devices 100 can implement the functions of one or more units among the acquisition unit 510, draft determination unit 520, probability determination unit 530, verification unit 540, and available draft determination unit 550.

[0180] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 8 One possible implementation is shown. For example... Figure 8As shown, the two computing devices 100A and 100B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this possible implementation, the memory 106 in computing device 100A stores instructions for executing the probability determination unit 530. Meanwhile, the memory 106 in computing device 100B stores instructions for executing the functions of the acquisition unit 510, the draft determination unit 520, the verification unit 540, and the available draft determination unit 550.

[0181] Figure 8 The connection method between the computing device clusters shown can be based on the fact that the data processing method provided in this application can be divided into two parts: internal processing and interaction with the pre-trained model. Therefore, it is considered that the function of calling the cloud training model is executed by computing device 100A, and other functions are executed by computing device 100B.

[0182] It should be understood that Figure 8 The functions of the computing device 100A shown can also be performed by multiple computing devices 100. Similarly, the functions of the computing device 100B can also be performed by multiple computing devices 100.

[0183] This application also provides another computing device cluster. The connection relationships between the computing devices in this computing device cluster can be similarly referred to... Figure 7 and Figure 8 The connection method of the computing device cluster. The difference is that the memory 106 of one or more computing devices 100A in the computing device cluster can store the same instructions for executing data processing methods.

[0184] In some possible implementations, the memories of one or more computing devices 100B in the computing device cluster may also each store a portion of the instructions for executing the data processing method. In other words, a combination of one or more computing devices can jointly execute the instructions for executing the data processing method.

[0185] It should be noted that the memory 106 in different computing devices 100A within the computing device cluster can store different instructions for executing some functions of the data processing device. That is, the instructions stored in the memory 106 of different computing devices 100A can implement the functions of one or more units in the data processing device.

[0186] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform a data processing method.

[0187] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium capable of being stored by a computing device, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform a data processing method, or instruct the computing device to perform a data processing method.

[0188] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A data processing method, characterized in that, The method includes: Get the first input sent by the user; Based on the first input, a plurality of target drafts matching the first input are determined from the draft library, wherein the target drafts include at least one draft token; The matching probability of each draft token in each target draft is obtained by calling the pre-trained model. The matching probability is the probability that the pre-trained model outputs the draft token when the first input is used as input. The draft tokens that pass verification are identified as those whose matching probability is greater than the target threshold. The available target drafts among the plurality of target drafts are determined, wherein the number of verified draft tokens in the available target drafts is greater than that in other drafts among the plurality of target drafts besides the available target drafts, and the verified draft tokens in the available target drafts are used to generate an answer to the first input.

2. The method according to claim 1, characterized in that, Before determining the available information, the method further includes: The threshold configuration request sent by the user is obtained through the threshold configuration interface, and the threshold configuration request is used to configure the target threshold.

3. The method according to claim 2, characterized in that, The method further includes: At least one set of threshold recommendation information is determined, each set of threshold recommendation information including task type information and reference threshold information, the threshold recommendation information being used to provide a reference in the process of determining the target threshold based on the task type corresponding to the first input; A threshold recommendation interface is provided, which is used to display the at least one set of threshold recommendation information.

4. The method according to any one of claims 1 to 3, characterized in that, Before retrieving from the draft library based on the first input, the method further includes: Obtain the context information of the first input; Multiple first candidate drafts are determined based on the context information of the first input; Add the first candidate draft to the draft library.

5. The method according to claim 4, characterized in that, The method further includes: Obtain the second input and its context information, wherein the second input belongs to the same session as the first input; Based on the context information of the second question, multiple second candidate drafts were identified; Add the plurality of second candidate drafts to the draft library.

6. The method according to any one of claims 1 to 5, characterized in that, The step of calling the pre-trained model to validate the multiple target drafts includes: In response to the fact that the length of the context information of the first input is greater than the maximum input length of the pre-trained model, the context information of the first input is compressed; The compressed context information of the first input, the first input, and the multiple target drafts are input into the pre-trained model; The method further includes: A compression ratio display page is provided, which is used to display the compression ratio of the first input. The compression ratio of the first input is determined based on the length of the context information of the first input before and after compression.

7. A data processing apparatus, characterized in that, The device includes: The acquisition unit is used to acquire the first input sent by the user; A draft determination unit is configured to determine, based on the first input, multiple target drafts matching the first input from a draft library, wherein the target drafts include at least one draft token; The probability determination unit is used to call a pre-trained model to obtain the matching probability of each draft token in each target draft, wherein the matching probability is the probability that the pre-trained model outputs the draft token when the first input is used as input; A verification unit is used to determine the draft tokens that have passed verification, wherein the draft tokens that have passed verification are the draft tokens whose matching probability is greater than the target threshold; A usable draft determination unit is used to determine usable target drafts among the plurality of target drafts, wherein the number of verified draft tokens in the usable target drafts is greater than that in other drafts among the plurality of target drafts besides the usable target drafts, and the verified draft tokens in the target usable drafts are used to generate an answer to the first input.

8. The apparatus according to claim 7, characterized in that, The acquisition unit is further configured to acquire the threshold configuration request sent by the user through the threshold configuration interface, and the threshold configuration request is used to configure the target threshold.

9. The apparatus according to claim 8, characterized in that, The device also includes a recommendation information determination unit; The recommendation information determination unit is used to determine at least one set of threshold recommendation information, each set of threshold recommendation information including task type information and reference threshold information, the threshold recommendation information being used to provide a reference in the process of determining the target threshold according to the task type corresponding to the first input; A threshold recommendation interface is provided, which is used to display the at least one set of threshold recommendation information.

10. The apparatus according to any one of claims 7 to 9, characterized in that, The device also includes a draft library configuration unit; The acquisition unit is further configured to acquire the context information of the first input; The draft library configuration unit is used to determine multiple first candidate drafts based on the context information of the first input; and to add the first candidate drafts to the draft library.

11. The apparatus according to claim 10, characterized in that, The acquisition unit is further configured to acquire the second input and the context information of the second input, wherein the second input and the first input belong to the same session; The draft library configuration unit is further configured to determine multiple second candidate drafts based on the context information of the second problem, and add the multiple second candidate drafts to the draft library.

12. The apparatus according to any one of claims 7 to 11, characterized in that, The device also includes a compression unit; The compression unit is specifically configured to compress the context information of the first input in response to the fact that the length of the context information of the first input is greater than the maximum input length of the pre-trained model; and input the compressed context information of the first input, the first input, and the plurality of target drafts into the pre-trained model. A compression ratio display page is provided, which is used to display the compression ratio of the first input. The compression ratio of the first input is determined based on the length of the context information of the first input before and after compression.

13. A computing device, characterized in that, The computing device includes a processor and memory; The processor is configured to execute instructions stored in the memory to cause the computing device to perform the operational steps of the method as described in any one of claims 1 to 6.

14. A computing device cluster, characterized in that, The computing device cluster includes at least one computing device, each computing device including a processor and memory: The memory is used to store instructions; The processor is configured to, according to the instructions, cause the computing device cluster to perform the operational steps of the method according to any one of claims 1 to 6.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a computing device, cause the computing device to perform the operational steps of the method as described in any one of claims 1 to 6.

16. A computer program product comprising instructions that, when run on a computing device, cause the computing device to perform the operational steps of the method as described in any one of claims 1 to 6.