Text generation method, electronic equipment and storage medium
By updating the candidate set in the lookahead framework, using the matching priority of target prediction information and text sequences, the problem of n-gram pool capacity limitation is solved, improving the efficiency and quality of text generation and improving user experience.
Patent Information
- Application Number
- CN202510400565.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
AI Technical Summary
When maintaining n-gram pools, the existing lookahead framework is limited by capacity, resulting in important n-grams being abandoned, which reduces the model's inference speed and matching rate and affects the user experience.
By updating the candidate set based on the target text sequence and the initial candidate set based on the input information, the target prediction information is used for matching, ensuring that the latest input information or the text sequence that generates the prediction information is the highest matching priority, reducing the waste of storage resources and computing needs, and improving matching efficiency.
Improve the quality and user experience of text generation, and maintain the consistency and logic of text generation by updating candidate collections in real time, improving the inference speed and resource utilization.
Smart Images

Figure CN120256580A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical fields of machine learning and artificial intelligence, and particularly to a text generation method, an electronic device, and a storage medium. Background Art
[0002] In the field of natural language processing, the lookahead framework can process multiple possible word sequences (branches) in parallel during the decoding process by maintaining an n-gram pool, enabling the model to generate multiple tokens at each decoding step, which improves the decoding speed to a certain extent. However, due to capacity limitations, the current method of maintaining the n-gram pool using a queue may cause some important n-grams to be discarded during updates, reducing the inference speed of the model. Summary of the Invention
[0003] One aspect of the present disclosure provides a text generation method, including: obtaining target prediction information based on a target text sequence of input information and an initial candidate set, the initial candidate set including text sequences with matching priorities; updating the initial candidate set using the target text sequence and the target prediction information to obtain a target candidate set, the target prediction information being the information with the highest matching priority in the target candidate set; matching the target prediction information with the information in the target candidate set based on the matching priority to obtain a matching result; and outputting a target text based on the matching result.
[0004] According to an embodiment of the present disclosure, updating the initial candidate set using the target text sequence and the target prediction information to obtain a target candidate set includes: updating the initial candidate set using the target text sequence to obtain an intermediate candidate set; and updating the intermediate candidate set using the target prediction information to obtain a target candidate set.
[0005] According to an embodiment of the present disclosure, updating the intermediate candidate set using the target prediction information to obtain a target candidate set includes: when the target prediction information exists in the intermediate candidate set, determining the priority of the target prediction information as the highest priority to obtain a target candidate set; when the target prediction information does not exist in the intermediate candidate set, adding the target prediction information to the intermediate candidate set and determining it as the highest priority to obtain a target candidate set.
[0006] According to an embodiment of the present disclosure, the method further includes: when the resource consumption information corresponding to the intermediate candidate set is greater than or equal to a preset consumption threshold, deleting the text sequence with the lowest priority from the intermediate candidate set based on the usage frequency.
[0007] According to an embodiment of the present disclosure, obtaining target prediction information based on a target text sequence and an initial candidate set of input information includes: determining a plurality of initial prediction information corresponding to the input information and probability values respectively corresponding to the plurality of initial prediction information; and determining target prediction information from the plurality of initial prediction information based on the probability values.
[0008] According to an embodiment of the present disclosure, the target candidate set includes a plurality of text sequences; matching the target prediction information with the information in the target candidate set based on a matching priority to obtain a matching result, including: generating a plurality of prediction sequences based on the plurality of target prediction information; and matching the plurality of prediction sequences with the text sequences in the target candidate set based on the matching priority to obtain a matching result.
[0009] According to an embodiment of the present disclosure, outputting a target text based on the matching result includes: generating a plurality of intermediate sequences corresponding to different time points when the matching result indicates that the prediction sequence matches the text sequence; and outputting the target text when the intermediate sequence includes a stop symbol and / or the text length formed by the plurality of intermediate sequences is greater than or equal to a preset length threshold.
[0010] According to an embodiment of the present disclosure, the method further includes: when the matching result indicates that the prediction sequence does not match the text sequence, adding the prediction sequence to the target candidate set and determining it as the highest priority.
[0011] Another aspect of the present disclosure provides a text generation device, including: an obtaining module, configured to obtain target prediction information based on a target text sequence and an initial candidate set of input information, where the initial candidate set includes text sequences with matching priorities; an updating module, configured to update the initial candidate set with the target text sequence and the target prediction information to obtain a target candidate set, where the target prediction information is the information with the highest matching priority in the target candidate set; a matching module, configured to match the target prediction information with the information in the target candidate set based on the matching priority to obtain a matching result; and a text output module, configured to output a target text based on the matching result.
[0012] Another aspect of the present disclosure provides an electronic device, including at least one processor and at least one processing model capable of running on the processor, where the processing model can be called by a target application to perform at least one of the following: obtaining target prediction information based on a target text sequence and an initial candidate set of input information, where the initial candidate set includes text sequences with matching priorities; updating the initial candidate set with the target text sequence and the target prediction information to obtain a target candidate set, where the target prediction information is the information with the highest matching priority in the target candidate set; matching the target prediction information with the information in the target candidate set based on the matching priority to obtain a matching result; and outputting a target text based on the matching result.
[0013] Another aspect of the present disclosure provides a non - volatile storage medium storing computer - executable instructions that, when executed, are used to implement the method of any one of the above.
[0014] Another aspect of the present disclosure provides a computer program that includes computer - executable instructions that, when executed, are used to implement the method of any one of the above. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To more fully understand the present disclosure and its advantages, reference will now be made to the following description taken in conjunction with the accompanying drawings, in which:
[0016] Figure 1 shows an application scenario of a text generation method, apparatus, and electronic device according to an embodiment of the present disclosure;
[0017] Figure 2 schematically shows a flowchart of a text generation method according to an embodiment of the present disclosure;
[0018] Figure 3A schematically shows an example diagram of a text generation process according to an embodiment of the present disclosure;
[0019] Figure 3B schematically shows a flowchart of another text generation method according to an embodiment of the present disclosure;
[0020] Figure 4 shows a block diagram of a text generation apparatus according to an embodiment of the present disclosure;
[0021] Figure 5 schematically shows a diagram of an electronic device according to an embodiment of the present disclosure;
[0022] Figure 6 shows a schematic block diagram of an example electronic device that can be used to implement the text generation method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. It should be understood, however, that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In addition, in the following description, descriptions of well - known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.
[0024] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0025] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.
[0026] Some block diagrams and / or flowcharts are shown in the drawings. It should be understood that some blocks in the block diagrams and / or flowcharts, or combinations thereof, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, these instructions can create a device for implementing the functions / operations illustrated in these block diagrams and / or flowcharts.
[0027] Therefore, the technology of the present disclosure can be implemented in the form of hardware and / or software (including firmware, microcode, etc.). Additionally, the technology of the present disclosure can take the form of a computer program product on a computer-readable medium storing instructions, which can be used by or in conjunction with an instruction execution system. In the context of the present disclosure, a computer-readable medium can be any medium that can contain, store, transmit, propagate, or transport instructions. For example, a computer-readable medium can include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, components, or propagation media. Specific examples of computer-readable media include: magnetic storage devices, such as magnetic tapes or hard disk drives (HDDs); optical storage devices, such as compact discs (CD-ROMs); memories, such as random access memories (RAMs) or flash memories; and / or wired / wireless communication links.
[0028] Figure 1 An application scenario of a text generation method, apparatus, and electronic device according to an embodiment of the present disclosure is shown.
[0029] It should be noted that Figure 1 The shown is only an example of a scenario where the embodiments of the present disclosure can be applied, to help those of ordinary skill in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments, or scenarios.
[0030] As Figure 1 shown, the application scenario according to this embodiment can include a terminal device 101, a network 102, and a server 103. The network 102 is used to provide a medium for a communication link between the terminal device 101 and the server 103. The network 102 can include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0031] The terminal device 101 may be installed with a text generation application capable of receiving text information input by a user, enabling the user to input text words or text sentences through this application. The terminal device 101 may also display or play the target text generated by the server 103 in response to the input information acquisition request of the terminal device 101.
[0032] The terminal device 101 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, desktop computers, and the like.
[0033] The server 103 may be a processor providing various services. For example, the server 103 may generate a text sequence and a candidate set according to the information input by the user. For another example, the server 103 may further process the generated text sequence and the initial candidate set to obtain prediction information. For yet another example, the service 105 may perform a matching process between the prediction information and the information in the candidate set to obtain a matching result.
[0034] It should be noted that the text generation method provided by the embodiments of the present disclosure may generally be executed by the server 103. Correspondingly, the text generation device provided by the embodiments of the present disclosure may generally be disposed in the server 103. The text generation method provided by the embodiments of the present disclosure may also be executed by a server or a server cluster different from the server 103 and capable of communicating with the terminal device 101 and / or the server 103. Correspondingly, the text generation device provided by the embodiments of the present disclosure may also be disposed in a server or a server cluster different from the server 103 and capable of communicating with the terminal device 101 and / or the server 103.
[0035] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in
[0036] are merely illustrative. According to the implementation requirements, there may be any number of terminal devices, networks, and servers.
[0037] However, when the n-gram pool approaches the capacity threshold, the addition of new n-grams will cause the old n-grams to be removed. If a simple first-in-first-out strategy is adopted, it may lead to the removal of some high-priority n-grams that are very important for the subsequent decoding process, resulting in the subsequent verification branches being unable to utilize these important n-grams when matching n-grams, reducing the matching rate and affecting the inference speed of the model.
[0038] In some examples, in the application scenario of multi-turn conversations, the lookahead method optimizes dialogue generation through policy planning. For example, a lookahead heuristic method similar to the A search algorithm is adopted to predict future dialogue policy sequences and their possible user feedbacks to select the best dialogue policy.
[0039] However, this method requires predicting future dialogue policy sequences and user feedbacks, involving a large number of computational operations, which may slow down the system's response speed and affect the user experience.
[0040] Based on the above problems, the present disclosure provides a text generation method, including: obtaining target prediction information based on a target text sequence of input information and an initial candidate set, where the initial candidate set includes text sequences with matching priorities; updating the initial candidate set with the target text sequence and the target prediction information to obtain a target candidate set, where the target prediction information is the information with the highest matching priority in the target candidate set; matching the target prediction information with the information in the target candidate set based on the matching priority to obtain a matching result; and outputting a target text based on the matching result.
[0041] According to an embodiment of the present disclosure, by matching the generated target prediction information with the information in the target candidate set based on the matching priority, a matching result is obtained to output the target text. Since the target candidate set takes the text sequence of the latest input information or the generated prediction information as the highest matching priority, the hit rate of the target prediction information and the information in the target candidate set can be improved. The text sequence of the latest input information or the generated prediction information can make the information retained in the target candidate set be the most relevant to subsequent predictions, avoiding waste of storage resources and the need for additional computational resources, improving the matching efficiency and the generation quality of the target text, and further enhancing the user's usage experience.
[0042] Figure 2 Schematically shows a flowchart of the text generation method according to an embodiment of the present disclosure.
[0043] As Figure 2 shown, the method includes operations S201 to S204.
[0044] In operation S201, target prediction information is obtained based on a target text sequence of input information and an initial candidate set, where the initial candidate set includes text sequences with matching priorities.
[0045] In an embodiment of the present disclosure, the input information may be any information input by a user through a terminal device according to actual needs. The target text sequence may be a text sequence obtained by processing the input information. The initial candidate set may be a set composed of multiple text sequences to be matched with different matching priorities, so as to accelerate generation through parallel prediction and parallel verification. The target prediction information may be information obtained through parallel prediction based on the input information and the initial candidate set.
[0046] For example, a serialization tool is used to perform word segmentation processing on the input information, and the result after word segmentation is converted into a corresponding text sequence, thereby obtaining the target text sequence. For example, the initial candidate set may include multiple basic text sequences to be matched. After the target text sequence is generated, the target text sequence may be added to the basic text sequences to obtain the initial candidate set. Alternatively, when initially generating text, an initial candidate set is constructed using the target text sequence of the input information.
[0047] In operation S202, the initial candidate set is updated using the target text sequence and the target prediction information to obtain a target candidate set, and the target prediction information is the information with the highest matching priority in the target candidate set.
[0048] In an embodiment of the present disclosure, the target candidate set may be a sequence set including target text sequences with different matching priorities and prediction sequences based on the target prediction information. The matching priorities corresponding to the target text sequences and the prediction sequences in the target candidate set may be determined by the least recently or most recently used policy.
[0049] For example, when generating a new target text sequence at the current moment, the initial candidate set may be updated using the target text sequence, so as to determine that the target text sequence at the current moment is the sequence with the highest matching priority based on the "least recently used" principle in the cache eviction algorithm. For example, at the current moment, the target text sequence corresponding to the input information is obtained, and the initial candidate combination is updated using the target text sequence at the current moment to obtain the target candidate set.
[0050] In operation S203, the target prediction information is matched with the information in the target candidate set based on the matching priority to obtain a matching result.
[0051] In an embodiment of the present disclosure, the information in the target candidate set may include the target text sequence of the input information at the current moment and the prediction sequence of the target prediction information at the previous moment, which may be referred to as the sequence to be matched for subsequent matching. The matching process may be regarded as a process of verifying the rationality of the target prediction information at the current moment. The matching result may include mutual matching and non-matching based on the level of the matching degree. For example, when the prefix information of the prediction sequence at the current moment matches the sequence to be matched, the rationality degree of the prediction sequence is the highest.
[0052] In operation S204, based on the matching result, the target text is output.
[0053] In an embodiment of the present disclosure, the target text may be a predicted text composed of multiple predicted sequences obtained when the multiple matching results are successful in matching, and when the stop generation condition is satisfied, the target text is output. The stop generation condition may include any one of a stop symbol condition and a text length condition based on different types.
[0054] For example, when the sequence to be matched at the current moment matches the predicted sequence, the sequence result is continuously generated. When multiple sequence results satisfy the stop symbol condition, or the entire text length formed by multiple sequence results satisfies the text length condition, the generation is stopped and the target text is output.
[0055] According to an embodiment of the present disclosure, by matching the generated target prediction information with the information in the target candidate set based on the matching priority, a matching result is obtained to output the target text. Since the target candidate set takes the text sequence of the latest input information or the generated prediction information as the highest matching priority, the hit rate of the target prediction information and the information in the target candidate set can be improved. The text sequence of the latest input information or the generated prediction information can make the information retained in the target candidate set the most relevant to the subsequent prediction, avoiding waste of storage resources and the need for additional computing resources, improving the matching efficiency and the generation quality of the target text, and further enhancing the user experience.
[0056] It can be understood that how to output the target text has been described above. Below, an exemplary description of how to obtain the target candidate set will be given.
[0057] According to an embodiment of the present disclosure, the initial candidate set is updated using the target text sequence and the target prediction information to obtain the target candidate set, including: updating the initial candidate set using the target text sequence to obtain an intermediate candidate set; and updating the intermediate candidate set using the target prediction information to obtain the target candidate set.
[0058] In an embodiment of the present disclosure, the intermediate candidate set may be a candidate set obtained by adding the target text sequence of the input information at the current moment to the initial candidate set. The target candidate set may be a candidate set obtained by adding the predicted sequence of the target prediction information at the current moment to the intermediate candidate set.
[0059] For example, the initial candidate set at the k-2 moment is updated based on the target text sequence of the input information at the k-1 moment to obtain the (k-1)th candidate set at the k-1 moment, and then the (k-1)th candidate set is updated using the predicted sequence at the k-1 moment to obtain the kth candidate set at the k moment.
[0060] For example, after obtaining the intermediate candidate set, it is determined whether there is intersection information between the intermediate candidate set and the target prediction information to determine the target candidate set and the matching priority of the text sequences in the target candidate set.
[0061] It can be understood that in the context of multi-round conversations and long-understanding texts, updating candidate sets at different stages with the help of input information can improve the hit rate of text sequences in the lookahead inference process and further improve the inference speed of the processing model. By saving the candidate sets of the user's historical habitual questions and answers, the speed of answering the user's questions can be accelerated, and the user experience can be improved.
[0062] For example, each time the user interacts with the processing model, the corresponding text sequences in the user input and the target prediction information generated by the model can be extracted and stored in a global candidate set, which can use a Trie tree structure for efficient storage and retrieval. During each new conversation or text generation process, the sequences are dynamically added to the pool to ensure that the content in the pool is always relevant to the user's latest interaction. Then, when generating a new response, the text sequences relevant to the current input are first retrieved from the target candidate set. By matching these text sequences, reasonable candidate sequences can be quickly generated, reducing invalid calculations and backtracking.
[0063] Furthermore, by updating the candidate sequences in real time, the generated text can be made consistent with the previous context, improving the coherence of multi-round conversations and the logic of long texts.
[0064] According to an embodiment of the present disclosure, the text generation method can be implemented by a text generation model based on the lookahead framework. By adding the target text sequence and the prediction sequence to the candidate set and updating it in real time, the number of sequences related to the current context in the candidate set can be increased, enabling the text generation model to more quickly find appropriate text sequences in subsequent prediction and verification processes, thereby improving the hit rate of text sequences and the inference speed.
[0065] According to an embodiment of the present disclosure, updating the intermediate candidate set with the target prediction information to obtain the target candidate set includes: when the target prediction information exists in the intermediate candidate set, determining the priority of the target prediction information as the highest priority to obtain the target candidate set; when the target prediction information does not exist in the intermediate candidate set, adding the target prediction information to the intermediate candidate set and determining it as the highest priority to obtain the target candidate set.
[0066] In an embodiment of the present disclosure, when there is intersection information between the intermediate candidate set and the target prediction information, it can be indicated that the target prediction information already exists in the intermediate candidate set. When there is no intersection information between the intermediate candidate set and the target prediction information, it can be indicated that the target prediction information does not exist in the intermediate candidate set.
[0067] For example, a prefix tree structure is used to store and match text sequences. By looking up and matching prefixes, it is determined whether the prediction sequence already exists in the intermediate candidate set; if the prediction sequence already exists in the intermediate candidate set, its priority can be updated to ensure that it is given priority in subsequent generation and verification processes.
[0068] For example, it is implemented by maintaining a priority field in the prefix tree structure. Whenever a text sequence is detected to exist, its priority is set to the highest value in the current candidate set.
[0069] For example, if the prediction sequence is not in the current candidate set, it is inserted into the prefix tree structure and its priority is set to the highest, so that the newly inserted sequence can be quickly identified and used in subsequent prediction and verification processes.
[0070] According to an embodiment of the present disclosure, before adding the prediction sequence to the candidate set, setting a way to detect whether there is intersection information between the detection candidate set and the prediction sequence can reduce invalid calculations and backtracking, and further accelerate the entire inference process.
[0071] According to an embodiment of the present disclosure, the method further includes: when the resource consumption information corresponding to the intermediate candidate set is greater than or equal to a preset consumption threshold, deleting the text sequence with the lowest priority from the intermediate candidate set based on the usage frequency.
[0072] In an embodiment of the present disclosure, the resource consumption information can be used to characterize the storage space information of the candidate set at different stages. Considering the case of maintaining the candidate set in the form of a queue and the limited storage space of the candidate set, the present disclosure controls the resource usage of the candidate set by setting a consumption threshold (for example, 90%). It can be understood that the consumption threshold can be determined according to the actual situation and is not specifically limited herein.
[0073] For example, real-time monitoring of the resource consumption of the intermediate candidate set, including memory occupancy, computing resources, etc., can obtain monitoring information through system monitoring tools or custom resource monitoring modules; thus, according to the limitations of system resources and task requirements, a resource consumption threshold is preset in advance. When the resource consumption reaches or exceeds this threshold, the least recently used cache policy is triggered to delete the sequence with the lowest priority.
[0074] For example, maintain an access record for each text sequence in the intermediate candidate set, recording the timestamp of its most recent access or use. A hash table can be used to store these timestamps for quick lookup and update. When a text sequence needs to be deleted, the hash table can be traversed to find the text sequence with the earliest timestamp (i.e., the least recently accessed), which is the text sequence with the lowest priority. Then, the determined text sequence with the lowest priority is deleted from the intermediate candidate set, and the hash table is updated to remove its corresponding access record.
[0075] According to an embodiment of the present disclosure, based on the least recently used caching policy, by timely deleting text sequences with low priority, system resources are released, resource utilization is improved, and it is ensured that the system can operate efficiently. The size of the intermediate candidate set is reduced, the computational complexity is lowered, the inference process is accelerated, and the response speed of the model is increased.
[0076] According to an embodiment of the present disclosure, based on the target text sequence of the input information and the initial candidate set, target prediction information is obtained, including: determining a plurality of initial prediction information corresponding to the input information and probability values corresponding to each of the plurality of initial prediction information; and determining the target prediction information from the plurality of initial prediction information based on the probability values.
[0077] In an embodiment of the present disclosure, the probability value can be the probability value corresponding to the initial prediction branch of the initial prediction information. Based on the Lookahead method, a plurality of initial prediction branches can be generated and the probability value corresponding to each initial prediction branch can be calculated. Thus, based on the corresponding probability values, the one with the highest probability value is selected or selected based on a certain probability to obtain the selected sequence, and then these selected sequences are used as the target prediction information.
[0078] For example, when generating prediction branch tokens, sampling can be performed according to the probability value of each token. By calculating the generation probability of each possible token, random sampling is performed according to these probability values to select a token as the next predicted token. This probability-based sampling method can ensure that the generated token sequence not only conforms to the prediction of the language model but also has a certain degree of diversity.
[0079] It can be understood that the above has given an exemplary description of how to determine the target prediction information. Next, how to obtain a matching result using the target prediction information and the target candidate set will be described.
[0080] According to an embodiment of the present disclosure, the target candidate set includes a plurality of text sequences; the target prediction information is matched with the information in the target candidate set based on the matching priority to obtain a matching result, including: matching the plurality of prediction sequences in the target prediction information with the text sequences of the target candidate set based on the matching priority to obtain a matching result.
[0081] In an embodiment of the present disclosure, the prediction sequence at the current moment can be respectively matched with the text sequences in the target candidate set. By identifying the longest subsequence in the text sequence, when the subsequence matches the prediction sequence, this part of the subsequence is determined as the sequence result that meets the matching degree requirement.
[0082] For example, the text sequence is arranged in relative order to obtain multiple subsequences, and the common subsequence in different two subsequences is determined from the multiple subsequences. Furthermore, the ratio of the length of the common subsequence to the length of the longest one among the two subsequences is used as the similarity value. When the similarity is greater than the preset similarity threshold, the matching prediction sequence and text sequence are obtained.
[0083] According to an embodiment of the present disclosure, based on the matching result, the target text is output, including: when the matching result indicates that the prediction sequence matches the text sequence, generating multiple intermediate sequences corresponding to different moments; and when the intermediate sequence contains a stop symbol and / or the text length composed of multiple intermediate sequences is greater than or equal to the preset length threshold, outputting the target text.
[0084] In an embodiment of the present disclosure, the intermediate sequence can represent the text sequence continuously generated when the prediction sequence matches the text sequence, and detect the total length of the text sequence at the current moment or whether there is a stop symbol in the sequence each time a new text sequence is generated, so as to determine whether to end the prediction process and generate the text.
[0085] For example, in the Lookahead method, by setting a preset length threshold or defining a stop symbol to control the length of the generated text, the generated text can be neither too long nor end naturally semantically.
[0086] For example, by setting a preset length threshold (for example, 50 tokens), it can be ensured that the generated text does not exceed this length. Or, by defining one or more stop symbols (for example, the full stop “。”), it can end naturally in the generated text.
[0087] Figure 3A Schematically shows an example diagram of the text generation process according to an embodiment of the present disclosure.
[0088] As Figure 3A shown, taking the input information as “The weather is nice today, go...” as an example, assuming the sequence length n of the text sequence is 3, the input information is serialized to obtain the target text sequence 301 (“Today”, “weather”, “nice”, “,”, “go”); thus, the initial candidate set 303 is constructed by using the target text sequence 301 and the base sequence 302.
[0089] After obtaining the initial candidate set 303, prediction can be performed based on the target text sequence 301 and the initial candidate set 303 to obtain multiple initial prediction information 304 ("park" and "ride a bike"). Through the rationality or matching degree corresponding to each of the multiple initial prediction information, the prediction information with the highest rationality (for example, the matching degree of "park" is 90% and the matching degree of "ride a bike" is 70%) ("park") is used as the target prediction information 305.
[0090] After obtaining the target prediction information 305 ("park"), the intermediate prediction sequence ending with "park" ["very good", "go", "park"] can be added to the initial candidate set 303 to obtain the intermediate candidate set 306. It can be understood that before adding this intermediate prediction sequence, it can be first detected whether this sequence already exists in the initial candidate set 303. If it exists, there is no need to add it, and the matching priority of this sequence is determined to be the highest; if it does not exist, then add it and determine it to be the highest priority.
[0091] Further, the target prediction information 305 can be matched with the information in the intermediate candidate set 306 based on the matching priority to verify whether the target prediction information 305 hits. If it hits, prediction of the next sequence can be performed. If it does not hit (assuming the target prediction information at this time is "ride a bike"), then the sequence of length n = 3 ["very good", "go", "ride a bike"] ending with the unhit target prediction information (that is, "ride a bike") is used to update the intermediate candidate set 306 in real time to obtain the target candidate set 307.
[0092] In the process of matching the target prediction information 305 with the information in the target candidate set 307, it can be determined whether to stop prediction by detecting the stop symbol or the text length 308 to output the target text 309.
[0093] According to an embodiment of the present disclosure, the method further includes: in the case where the matching result indicates that the prediction sequence does not match the text sequence, adding the prediction sequence to the target candidate set and determining it to be the highest priority.
[0094] In an embodiment of the present disclosure, the longest common subsequence algorithm can be used to determine the matching degree between the prediction sequence and the text sequence. If the similarity is lower than the matching degree threshold, it can be determined that the prediction sequence does not match the text sequence.
[0095] For example, in the case where it is determined that the predicted sequence does not match the text sequence, the predicted sequence is added to the candidate set. In the candidate set, a priority field can be maintained for each sequence. When adding a new predicted sequence, its priority can be set to the highest value in the current candidate set. Further, when adding or removing a sequence each time, the priorities in the candidate set are dynamically updated to ensure that the latest predicted sequence has the highest priority.
[0096] Figure 3B The flowchart of another text generation method according to an embodiment of the present disclosure is schematically shown.
[0097] As Figure 3B shown, the method includes operations S310 to S350.
[0098] In operation S310, the input information is initialized. The input information of the user is obtained as the starting point of generation. A maximum text generation length is set, and the generation stops when the generated text reaches this preset length threshold. Alternatively, by defining a stop symbol, the generation stops when the stop symbol appears in the generated text.
[0099] In operation S320, multiple target prediction information is generated. The input information is converted into a target text sequence and added to the initial candidate set. Multiple target prediction information is generated by processing the model, the target text sequence, and the initial candidate set. The target prediction information includes multiple predicted sequences, and each predicted sequence can represent a possible subsequent text. The initial candidate set is updated in real time, and based on the resource consumption data of the initial candidate set, the text sequences with low matching priorities are removed to optimize the performance.
[0100] In operation S330, the initial candidate set is updated using the target prediction information and the target text sequence to obtain a target candidate set.
[0101] In operation S340, the target prediction information is matched with the information in the target candidate set. Based on the forward propagation verification strategy, the forward propagation of the language model is performed on each target prediction information to calculate its generation probability; by identifying the longest correct subsequence in the predicted sequence, its consistency with the prediction of the language model is ensured to obtain the matching result.
[0102] In operation S350, the target text is generated and output. Multiple target prediction information corresponding to different moments is continuously generated until the stop condition is met. The stop condition judgment can include: checking whether the length of the generated text reaches the preset threshold; checking whether the generated text contains a stop symbol; when any of the above stop conditions is met, the generated text is output as the target text.
[0103] Based on the above text generation method, the present disclosure also provides a text generation device. The following will be combined withFigure 4 Describe the device in detail.
[0104] Figure 4 The block diagram of the text generation device according to an embodiment of the present disclosure is shown.
[0105] As Figure 4 shown, the text generation device includes an obtaining module 401, an updating module 402, a matching module 403, and a text output module 404.
[0106] According to some embodiments of the present disclosure, the text generation device can be used to implement the text generation method of the embodiments of the present disclosure.
[0107] The obtaining module 401 can execute, for example, operation S201 to obtain target prediction information based on the target text sequence of the input information and the initial candidate set, where the initial candidate set includes text sequences with matching priorities.
[0108] The updating module 402 can execute, for example, operation S202 to update the initial candidate set using the target text sequence and the target prediction information to obtain a target candidate set, where the target prediction information is the information with the highest matching priority in the target candidate set.
[0109] The matching module 403 can execute, for example, operation S203 to match the target prediction information with the information in the target candidate set based on the matching priority to obtain a matching result.
[0110] The text output module 404 can execute, for example, operation S204 to output the target text based on the matching result.
[0111] According to the embodiments of the present disclosure, based on the obtaining module 401, the updating module 402, the matching module 403, and the text output module 404 in the text generation device, by matching the generated target prediction information with the information in the target candidate set based on the matching priority, a matching result is obtained to output the target text. Since the target candidate set takes the text sequence of the latest input information or the generated prediction information as the highest matching priority, the hit rate of the target prediction information and the information in the target candidate set can be improved. The text sequence of the latest input information or the generated prediction information can make the most relevant information be retained in the target candidate set, avoiding the waste of storage resources and the need for additional computing resources, improving the matching efficiency and the generation quality of the target text, and further enhancing the user experience.
[0112] According to an embodiment of the present disclosure, the update module 402 includes: a first update sub-module and a second update sub-module. The first update sub-module is configured to update the initial candidate set by using the target text sequence to obtain an intermediate candidate set; and the second update sub-module is configured to update the intermediate candidate set by using the target prediction information to obtain a target candidate set.
[0113] According to an embodiment of the present disclosure, the second update sub-module includes: a priority determination unit and an addition unit. The priority determination unit is configured to determine the priority of the target prediction information as the highest priority to obtain a target candidate set when the target prediction information exists in the intermediate candidate set; the addition unit is configured to add the target prediction information to the intermediate candidate set and determine it as the highest priority to obtain a target candidate set when the target prediction information does not exist in the intermediate candidate set.
[0114] According to an embodiment of the present disclosure, the apparatus further includes: a deletion module, configured to delete the text sequence with the lowest priority from the intermediate candidate set based on the usage frequency when the resource consumption information corresponding to the intermediate candidate set is greater than or equal to a preset consumption threshold.
[0115] According to an embodiment of the present disclosure, the acquisition module 401 includes: a probability value determination sub-module and a prediction information determination sub-module. The probability value determination sub-module is configured to determine a plurality of initial prediction information corresponding to the input information and probability values corresponding to each of the plurality of initial prediction information; and the prediction information determination sub-module is configured to determine the target prediction information from the plurality of initial prediction information based on the probability values.
[0116] According to an embodiment of the present disclosure, the target candidate set includes a plurality of text sequences; the matching module 403 includes: a sequence generation sub-module and a sequence matching sub-module. The sequence generation sub-module is configured to generate a plurality of prediction sequences based on the plurality of target prediction information; and the sequence matching sub-module is configured to match the plurality of prediction sequences with the text sequences of the target candidate set based on the matching priority to obtain a matching result.
[0117] According to an embodiment of the present disclosure, the text output module 404 includes: a prediction information generation sub-module and a target text output sub-module. The prediction information generation sub-module is configured to generate a plurality of intermediate sequences corresponding to different time points when the matching result indicates that the prediction sequence matches the text sequence; and the target text output sub-module is configured to output the target text when the intermediate sequence includes a stop symbol and / or the text length composed of the plurality of intermediate sequences is greater than or equal to a preset length threshold.
[0118] According to an embodiment of the present disclosure, the apparatus further includes: a sequence addition module, configured to add the prediction sequence to the target candidate set and determine it as the highest priority when the matching result indicates that the prediction sequence does not match the text sequence.
[0119] Figure 5 Schematically shows a schematic diagram of an electronic device according to an embodiment of the present disclosure.
[0120] As Figure 5 shown, the electronic device 50 includes at least one processor 51 and at least one processing model capable of running on the processor 51. The processing model can be called by a target application to perform at least one of the following: obtaining target prediction information based on a target text sequence of input information and an initial candidate set, where the initial candidate set includes text sequences with matching priorities; updating the initial candidate set using the target text sequence and the target prediction information to obtain a target candidate set, where the target prediction information is the information with the highest matching priority in the target candidate set; matching the target prediction information with the information in the target candidate set based on the matching priority to obtain a matching result; and outputting a target text based on the matching result.
[0121] Figure 6 Shows a schematic block diagram of an example electronic device that can be used to implement the text generation method according to an embodiment of the present disclosure.
[0122] The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0123] As Figure 6 shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0124] Multiple components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disc, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0125] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the virtual avatar driving method. For example, in some embodiments, the virtual avatar driving method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the virtual avatar driving method described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the virtual avatar driving method in any other suitable way (e.g., by means of firmware).
[0126] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0127] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0128] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0129] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0130] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0131] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client - server relationship is created by computer programs running on respective computers and having a client - server relationship with each other. Among them, the server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with a blockchain.
[0132] Those skilled in the art will understand that the features recited in the various embodiments and / or claims of the present disclosure can be combined or / and combined in many ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in the various embodiments and / or claims of the present disclosure can be combined and / or combined in many ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0133] Although the present disclosure has been shown and described with reference to particular exemplary embodiments thereof, those skilled in the art should understand that various changes in form and detail can be made therein without departing from the spirit and scope of the present disclosure as defined by the appended claims and their equivalents. Therefore, the scope of the present disclosure should not be limited to the above - described embodiments, but should be determined not only by the appended claims but also by the equivalents of the appended claims.
Claims
1. A text generation method, comprising: Obtaining target prediction information based on a target text sequence of input information and an initial candidate set, where the initial candidate set includes text sequences with matching priorities; Updating the initial candidate set using the target text sequence and the target prediction information to obtain a target candidate set, where the target prediction information is the information with the highest matching priority in the target candidate set; Matching the target prediction information with the information in the target candidate set based on the matching priority to obtain a matching result; And Outputting a target text based on the matching result.
2. The method according to claim 1, wherein updating the initial candidate set using the target text sequence and the target prediction information to obtain a target candidate set includes: Updating the initial candidate set using the target text sequence to obtain an intermediate candidate set; And Updating the intermediate candidate set using the target prediction information to obtain the target candidate set.
3. The method according to claim 2, wherein updating the intermediate candidate set using the target prediction information to obtain the target candidate set includes: When the target prediction information exists in the intermediate candidate set, determining the priority of the target prediction information as the highest priority to obtain the target candidate set; When the target prediction information does not exist in the intermediate candidate set, adding the target prediction information to the intermediate candidate set and determining it as the highest priority to obtain the target candidate set.
4. The method according to claim 3, further comprising: When the resource consumption information corresponding to the intermediate candidate set is greater than or equal to a preset consumption threshold, deleting the text sequence with the lowest priority from the intermediate candidate set based on the usage frequency.
5. The method according to claim 1, wherein obtaining target prediction information based on a target text sequence of input information and an initial candidate set includes: Determining a plurality of initial prediction information corresponding to the input information and probability values corresponding to each of the plurality of initial prediction information; And Determining the target prediction information from the plurality of initial prediction information based on the probability values.
6. The method according to claim 1, wherein the target candidate set includes a plurality of text sequences; Matching the target prediction information with the information in the target candidate set based on the matching priority to obtain a matching result includes: Matching a plurality of prediction sequences in the target prediction information with the text sequences of the target candidate set based on the matching priority to obtain the matching result.
7. The method according to claim 6, wherein outputting a target text based on the matching result includes: When the matching result indicates that the prediction sequence matches the text sequence and / or the prediction information, generating a plurality of target prediction information corresponding to different moments; And When the target prediction information includes a stop symbol and / or the text length of the target prediction information is greater than or equal to the preset length threshold, outputting the target text.
8. The method according to claim 7, wherein the method further comprises: In a case where the matching result indicates that the predicted sequence does not match the text sequence in the target candidate set, adding the predicted sequence to the target candidate set and determining it as the highest priority.
9. An electronic device, comprising at least one processor and at least one processing model capable of running on the processor, the processing model being callable by a target application to perform at least one of the following: Obtaining target prediction information based on a target text sequence of input information and an initial candidate set, the initial candidate set including text sequences with matching priorities; Updating the initial candidate set with the target text sequence and the target prediction information to obtain a target candidate set, the target prediction information being the information with the highest matching priority in the target candidate set; Matching the target prediction information with the information in the target candidate set based on the matching priority to obtain a matching result; And Outputting a target text based on the matching result.
10. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, The computer program or instruction is executed by the processor to perform at least one of the following: Obtaining target prediction information based on a target text sequence of input information and an initial candidate set, the initial candidate set including text sequences with matching priorities; Updating the initial candidate set with the target text sequence and the target prediction information to obtain a target candidate set, the target prediction information being the information with the highest matching priority in the target candidate set; Matching the target prediction information with the information in the target candidate set based on the matching priority to obtain a matching result; And Outputting a target text based on the matching result.