Parallel Generation Method of Reply Text, Terminal Device, and Storage Medium

Through the parallel generation method and multi-candidate sequence processing of the A-G module, combined with position embedding and multi-head self-attention, the problems of traditional text generation speed and semantic faults are solved, and faster and more natural reply generation is achieved.

CN119862892BActive Publication Date: 2025-07-08SHENZHEN XUNFANG TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510354487.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-08
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

Traditional text generation methods rely on word-by-word or sentence-by-word generation methods, resulting in slow generation speed, difficult to meet real-time interaction requirements, and easy to cause semantic faults due to local optimal solutions.

Method used

The parallel generation method is adopted to generate multiple candidate sequences through the A-G module, combining the timing relationship and deep semantics of the dialogue context between position embedding and multi-head self-attention capture, and combining the overall evaluation mechanism of candidate sequences, the generation process is optimized.

Benefits of technology

It significantly shortens the generation delay, improves the BLEU value of the reply, reduces the confusion, avoids semantic faults, and makes the generated reply more natural and logically coherent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119862892B_ABST
    Figure CN119862892B_ABST
Patent Text Reader

Abstract

The present invention is applicable to the field of data processing, and discloses a method for parallel generation of response texts, a terminal device, and a storage medium. The method for parallel generation of response texts includes: when input information is detected, generating a global vector corresponding to the dialogue context and generating an initial response sequence corresponding to the global vector; inputting the global vector and the initial response sequence into the A-G module for parallel processing to obtain multiple candidate sequences; and generating a response content according to the candidate sequences. Through the parallel processing mechanism of the A-G module, multiple candidate sequences are generated in a single iteration, shortening the time delay required by the traditional word-by-word generation method. The problem of delay caused by multiple iterations in the traditional sequence generation method is solved. In the stage of generating the global vector, through the cooperation of positional embedding and multi-head self-attention, the temporal relationship and deep semantics of the dialogue context can be captured. Combining with the overall evaluation mechanism of the candidate sequences of the A-G module, the phenomenon of semantic discontinuity caused by local optimal solutions in the traditional method is effectively avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data processing, and in particular relates to a method for parallel generation of reply texts, a terminal device, and a storage medium. Background Art

[0002] In the field of text generation, intelligent dialogue systems are widely used in many fields such as customer service, personal assistants, educational tutoring, medical consultations, and smart home control, improving the convenience of life and work efficiency. However, traditional text generation methods still have some limitations.

[0003] Traditional text generation methods often rely on a sequential generation method word by word or sentence by sentence. This method has a relatively slow generation speed. For example, it may take up to 2 seconds to generate a sentence word by word. A new technical means is needed to solve the above technical problems. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a method for parallel generation of reply texts, a terminal device, and a storage medium, which can solve the problem of relatively slow text generation speed in related technologies.

[0005] The first aspect of the present invention provides a method for parallel generation of reply texts, including:

[0006] When input information is detected, generate a global vector corresponding to the dialogue context, and generate an initial reply sequence corresponding to the global vector;

[0007] Input the global vector and the initial reply sequence into the A-G module for parallel processing to obtain multiple candidate sequences;

[0008] Generate a reply content according to the candidate sequences.

[0009] Optionally, in the first implementation manner of the first aspect of the present invention, when input information is detected, generating a global vector corresponding to the dialogue context further includes:

[0010] When input information is detected, obtain the dialogue context;

[0011] Encode each sentence in the dialogue context through a bidirectional gated recurrent unit to obtain a discourse feature vector, and assign a continuous position number to each sentence to obtain a position number;

[0012] According to each position number, calculate the position embedding PE of each position through alternating sine waves and cosine waves;

[0013] Concatenate each position embedding PE with each discourse feature vector to obtain a context representation including the dialogue order;

[0014] Generate the global vector corresponding to the dialogue context according to each of the context representations.

[0015] Optionally, in the second implementation manner of the first aspect of the present invention, the step of generating the global vector corresponding to the dialogue context according to each of the context representations includes:

[0016] Input each of the concatenated context representations into a multi-head self-attention module to generate the global vector corresponding to the dialogue context containing global semantic information.

[0017] Optionally, in the third implementation manner of the first aspect of the present invention, the step of inputting the global vector and the initial reply sequence into the A-G module for parallel processing to obtain a plurality of candidate sequences includes:

[0018] Calculate the degree of association between the initial reply sequence and the global vector to obtain an intermediate result integrating context information;

[0019] Input the intermediate result into a recurrent neural network unit with a memory function to obtain the hidden representation of the word to be generated;

[0020] Calculate the generation probability of each candidate word according to the hidden representation to form a list of candidate words sorted by probability;

[0021] Generate a plurality of the candidate sequences according to the list of candidate words.

[0022] Optionally, in the fourth implementation manner of the first aspect of the present invention, the step of generating a reply content according to the candidate sequences includes:

[0023] Assign gradually increasing weight values to each unit in the list of candidate words;

[0024] Multiply the word probability corresponding to each unit in the list of candidate words by the weight value corresponding to each unit and then accumulate to obtain a plurality of comprehensive probabilities;

[0025] Select a plurality of the candidate sequences according to the plurality of comprehensive probabilities.

[0026] Optionally, in the fifth implementation manner of the first aspect of the present invention, the step of assigning gradually increasing weight values to each unit in the list of candidate words includes:

[0027] Calculate the initial weight value according to the position index of each unit in the list of candidate words and a preset positive real number parameter through an exponential function, and the greater the position index of each, the greater the growth amplitude of the corresponding initial weight value;

[0028] Divide each initial weight value by the sum of all initial weight values to obtain the weight value corresponding to each of the units.

[0029] Optionally, in a sixth implementation manner of the first aspect of the present invention, after the step of generating a reply content according to the candidate sequence, the method further includes:

[0030] If a reply operation is being executed and new content input by the user is detected, stop the reply operation, and update the conversation context according to the complete semantic segments in the currently replied content, where the currently replied content is the content that is being replied but not yet fully replied.

[0031] Return to execute the step of generating the global vector corresponding to the conversation context.

[0032] Optionally, in a seventh implementation manner of the first aspect of the present invention, the step of generating a global vector corresponding to the conversation context when the input information is detected includes:

[0033] When the input information is detected, obtain the historical conversation context;

[0034] According to the input information, retain N rounds of conversations in the historical conversation context as the conversation context;

[0035] Generate the global vector corresponding to the conversation context.

[0036] In a second aspect, an embodiment of the present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned parallel generation method for reply text are implemented.

[0037] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned parallel generation method for reply text are implemented.

[0038] In a fourth aspect, an embodiment of the present invention provides a computer program product, which, when running on a terminal device, causes the terminal device to execute the above-mentioned parallel generation method for reply text.

[0039] The beneficial effects of the embodiments of the present invention compared with the prior art are as follows: Through the parallel processing mechanism of the A-G module, multiple candidate sequences are generated in a single iteration, shortening the time delay required by the traditional word-by-word generation method. The problem of delay caused by multiple iterations in the traditional sequence generation method is solved. In the stage of generating the global vector, through the synergistic effect of positional embedding and multi-head self-attention, the temporal relationship and deep semantics of the dialogue context can be captured. Combining with the overall evaluation mechanism of the candidate sequences of the A-G module, the BLEU value of the generated reply can be improved and the perplexity (PPL) can be reduced, effectively avoiding the semantic discontinuity phenomenon caused by local optimal solutions in the traditional method. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0041] Figure 1 It is a schematic diagram of an embodiment of the parallel generation method of the reply text in the embodiments of the present invention;

[0042] Figure 2 It is a schematic diagram of a specific embodiment of step S101 of the parallel generation method of the reply text in the embodiments of the present invention;

[0043] Figure 3 It is a schematic diagram of a specific embodiment of step S102 of the parallel generation method of the reply text in the embodiments of the present invention;

[0044] Figure 4 It is a schematic diagram of a specific embodiment of step S103 of the parallel generation method of the reply text in the embodiments of the present invention;

[0045] Figure 5 It is a schematic diagram of an embodiment of the terminal device in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] In order to make the purpose, technical solutions and advantages of the present invention clearer, the following further details the present invention in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0047] It should be noted that the terms "comprising", "including" and "having" and any variations thereof in the description, claims and drawings of the present invention are intended to cover non-exclusive inclusion. For example, a process, method, terminal, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products or devices. In the claims, description and drawings of the present invention, relational terms such as "first" and "second" are only used to distinguish one entity / operation / object from another entity / operation / object, and do not necessarily require or imply any such actual relationship or order between these entities / operations / objects.

[0048] The mention of "embodiment" in this context means that a particular feature, structure or characteristic described in connection with an embodiment may be included in at least one embodiment of the present invention. The phrase appears at various places in the description and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.

[0049] In the field of text generation, intelligent dialogue systems are widely used in many fields such as customer service, personal assistants, educational tutoring, medical consultations, and smart home control, improving the convenience of life and work efficiency. However, traditional text generation methods still have some limitations.

[0050] Traditional text generation methods often rely on a sequential generation method word by word or sentence by sentence. This method has a relatively slow generation speed. For example, it may take up to 2 seconds to generate a sentence word by word. A new technical means is needed to solve the above technical problems.

[0051] In view of this, the embodiments of the present invention provide a parallel generation method, terminal device and storage medium for reply text. Through the parallel processing mechanism of the A-G module, multiple candidate sequences are generated in a single iteration, shortening the time delay required by the traditional word-by-word generation method. It solves the delay problem caused by multiple iterations in the traditional sequence generation method. In the stage of generating the global vector, through the cooperation of positional embedding and multi-head self-attention, the temporal relationship and deep semantics of the dialogue context can be captured. Combining with the overall evaluation mechanism of the candidate sequences of the A-G module, the BLEU value of the generated reply can be improved and the perplexity (PPL) can be reduced, effectively avoiding the semantic discontinuity phenomenon caused by local optimal solutions in traditional methods.

[0052] In order to illustrate the technical solution of the present invention, specific embodiments will be described below.

[0053] Figure 1 The figure shows a schematic flowchart of an implementation of a method for parallel generation of response text provided by an embodiment of the present invention. This method can be applied to a terminal device. The terminal device can be a mobile phone, a tablet computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, etc.

[0054] Specifically, the above method for parallel generation of response text may include the following steps S101 to S103.

[0055] Step S101: When input information is detected, generate a global vector corresponding to the dialogue context and generate an initial response sequence corresponding to the global vector.

[0056] Specifically, when the system detects user input (such as a text message, a voice command), it starts a dialogue processing flow. Extract the context information of the current dialogue from the storage module, including the content of the user's latest input and the preset reserved historical dialogue turns (for example, the last 3 dialogue turns). Automatically filter out irrelevant historical information and focus on the dialogue content directly related to the current input.

[0057] Use a bidirectional recurrent neural network to deeply encode each sentence in the context and convert each sentence into a high-dimensional vector containing semantic features. For example, the user's question "How to relieve a headache?" is encoded into a 256-dimensional feature vector.

[0058] Assign consecutive position numbers to each sentence (such as the first sentence numbered 0, the second sentence numbered 1), generate corresponding position encoding vectors through a mathematical function to ensure that the model can recognize the order of the dialogue. Concatenate the semantic feature vectors and the position encoding vectors to form an enhanced representation that contains both semantic and temporal information.

[0059] Based on the global vector, generate an initial response segment.

[0060] Convert the global vector and the initial response segment into a vector form that can be processed by the model. Start multiple parallel generation units (A-G units), and each unit independently predicts the subsequent possible words. Each unit associates the global vector with the generated content through an attention mechanism to ensure that the prediction results conform to the context logic. Each unit outputs a list of candidate words sorted by probability, and combines them to form multiple candidate response sequences.

[0061] In a specific example, convert the input text data into a numerical vector representation. For each utterance in the dialogue context, first perform word segmentation and then obtain its vector representation through Embedding. Then use Bi-GRU to encode each utterance feature vector and use its last hidden state as the representation of this utterance. Finally, obtain the vector representations (global vectors) of N utterances.

[0062] Multi-Head Self-Attention (MHSA) can capture important information in the sequence and generate a context representation for each utterance by aggregating the features of other utterances from the conversation history. At the same time, Position Embedding (PE) is used to distinguish different positions between utterances. The specific operation is to concatenate them with the utterance representation, that is . Then the vector representation of the dialogue context is obtained . Among them, h N is the utterance feature vector of the Nth sentence in the dialogue context (N is a natural number sequence number, representing the Nth round of sentences in the conversation history). The Nth sentence is encoded by a bidirectional gated recurrent unit to obtain Bi-GRU, and forward and backward semantic analyses are performed simultaneously. After concatenating with the position embedding PE N , an enhanced representation [h N ; PE N is formed. The generation of h N will dynamically associate all relevant sentences in the conversation history (for example, h2 may strengthen the entity reference relationship with h5).

[0063] Step S102, input the global vector and the initial reply sequence into the A-G module for parallel processing to obtain multiple candidate sequences.

[0064] In a specific example, in order to generate the t-th to t + n-th words , first use the previously generated t - 1 words as input to obtain a representation with multi-head self-attention (future masking):

[0065] ;

[0066] where is the embedding vector of the generated words for the target response , that is .

[0067] There are n parallel A-G units in the A-G module, and each A-G unit contains: attention, GRU, and Softmax layers. Input the vector representation of the generated sequence and the dialogue context vector representation output by the encoder into the A-G module. For the k-th A-G unit in the A-G module, its working process each time has the following three steps:

[0068] The first step, use the response history representation as the query, and use the dialogue context vector representation Obtain the output representation through attention as the key and value :

[0069] ;

[0070] In the second step, use the GRU model as the decoder to generate the response. The decoding process is represented by the following formula:

[0071] ;

[0072] In the third step, use the Softmax layer to obtain the word probabilities through the output vector representation of the GRU as follows:

[0073] ;

[0074] where are trainable parameters

[0075] Since all A-G units in the A-G module receive the same input vector but it is easier to accurately predict the token for the A-G units arranged in the front. For the A-G units arranged in the back, since the token span between them and the front A-G units is relatively large, the generated tokens are less likely to be accurately predicted than those in the front. To address this challenge, the present invention introduces an exponential growth weight mechanism when designing the loss function. This method assigns different weights according to the positions of the A-G units, defined by an exponential function, to ensure that the later A-G units obtain larger weights. This strategy not only emphasizes the importance of the later part of the sequence but also can highlight some more critical sequence parts in specific application scenarios, thereby optimizing the overall generation effect. This method helps to balance the training signals of tokens at different positions and improve the consistency and naturalness of the entire response sequence

[0076] Now, if the model has n A-G units and each unit is responsible for generating a token, the exponential growth weight of the A-G unit at the i-th position can be designed as follows:

[0077] ;

[0078] where i is the position index of the A-G unit (from 0 to n - 1), is a positive real number parameter used to control the growth rate of the weights. For the choice of the value of , if you hope that the gap between the weights is not particularly large, you can choose a smaller value, which will make the growth of the weights relatively gentle. If you significantly emphasize the importance of the later GRU, you can choose a large The value, which will result in very small weights for the early GRUs and relatively large weights for the later GRUs.

[0079] Since directly using the above formula may lead to overly large weight differences, thus affecting the stability of gradient updates during training, the weights are normalized to ensure that their sum is 1 or within a reasonable range. The normalization method adopted in the present invention is to divide all weights by their sum, that is:

[0080] ;

[0081] This can ensure that the sum of all weights is 1 while maintaining the trend of exponential growth.

[0082] For each A-G unit (responsible for generating a token), the loss function can be expressed as:

[0083] ;

[0084] where is the conditional probability of the i-th token given the dialogue context and the previously generated token sequence . Finally, the total weighted loss function can be written as:

[0085] ;

[0086] Step S103, generate a response content according to the candidate sequence.

[0087] Optionally, higher weights are assigned to the later generation units to emphasize long-distance semantic coherence. The candidate sequences are weighted and scored, and the sequences with higher accuracy of the tail words are preferentially selected. When a complete sentence is generated or a new user input is detected, the generation is immediately terminated and the optimal response is output. Immediately stop the current response and turn to handle the new question.

[0088] The beneficial effects of the embodiments of the present invention compared with the prior art are: through the parallel processing mechanism of the A-G module, multiple candidate sequences are generated in a single iteration, shortening the time delay required by the traditional word-by-word generation method. Solved the delay problem caused by multiple iterations in the traditional sequence generation method. In the stage of generating the global vector, through the synergistic effect of positional embedding and multi-head self-attention, the temporal relationship and deep semantics of the dialogue context can be captured. Combined with the overall evaluation mechanism of the candidate sequences of the A-G module, the semantic discontinuity phenomenon caused by local optimal solutions in the traditional method is effectively avoided.

[0089] Refer to Figure 2 , Figure 2It is a schematic diagram of a specific embodiment of step S101 of the method for parallel generation of reply text in the embodiment of the present invention. Step S101 further includes the following specific implementations.

[0090] Step S1011, when the input information is detected, obtain the conversation context.

[0091] In the embodiment of the present invention, when user input (such as a text message, a voice command) is detected, a conversation context processing flow is started. The complete conversation history related to the current input is extracted from the storage module, including the content newly submitted by the user and the preset reserved historical conversation turns (for example, the last 5 conversation turns are reserved). Through an intelligent filtering mechanism, irrelevant or outdated conversation content is automatically removed to ensure that the context focuses on the core theme of the current interaction.

[0092] Step S1012, encode each sentence in the conversation context through a bidirectional gated recurrent unit to obtain a discourse feature vector, and assign a consecutive position number to each sentence to obtain a position number.

[0093] In the embodiment of the present invention, the conversation context is segmented into independent units according to natural sentences. For example, when the user inputs "I want to book a flight to Beijing tomorrow, economy class", it is split into four semantic units: "book a flight", "time: tomorrow", "destination: Beijing", and "class: economy class".

[0094] Use a bidirectional gated recurrent network (Bi-GRU) to deeply encode each sentence. This network analyzes the sentence from both the forward and backward directions simultaneously to capture context dependencies. For example, when encoding "economy class", not only its own semantics but also the intention of the previous sentence "book a flight" is associated. Output the semantic feature vector of each sentence (such as a 256-dimensional vector) to accurately represent the core meaning of the sentence.

[0095] Step S1013, according to each of the position numbers, calculate the position embedding PE of each position through alternating sine waves and cosine waves.

[0096] In the embodiment of the present invention, consecutive increasing position numbers are assigned to each sentence (for example, the first sentence is numbered 0, the second sentence is numbered 1) to clarify the temporal relationship of the conversation. For example, in a multi-turn conversation, "What time do you need to depart?" (numbered 2) is before "I want to book a morning flight" (numbered 3).

[0097] Based on the position numbers, calculate the position embedding vector through alternating sine wave and cosine wave functions. This method makes the encodings of adjacent positions continuous, while the encodings of distant positions are significantly different, helping the model distinguish the order of the conversation. For example, the position embedding of number 0 highlights the starting feature of the conversation, and the embedding of number 5 strengthens the tail information.

[0098] Step S1014: Concatenate each of the position embeddings PE with each of the utterance feature vectors to obtain a context representation that includes the conversation order.

[0099] In an embodiment of the present invention, the semantic feature vector (256 - dimensional) of each sentence is concatenated with the corresponding position embedding vector (128 - dimensional) to form a 384 - dimensional enhanced semantic representation. This operation enables the model to simultaneously understand the content of the statement and its position meaning in the conversation. For example, after concatenating the semantic vector of "economy class" with the embedding of position number 3, its occurrence order in the current conversation is clarified.

[0100] Step S1015: Generate the global vector corresponding to the conversation context according to each of the context representations.

[0101] In an embodiment of the present invention, the concatenated enhanced semantic representation is input into a multi - head self - attention module. This module analyzes the association strength between different statements through a parallel multi - path attention mechanism. For example, it identifies the strong correlation between "Beijing" and "flight" (attention weight 0.9) and weakens the non - direct association between "tomorrow" and "economy class" (attention weight 0.2). The finally output global vector (384 - dimensional) integrates the core intention, entity relationship, and temporal logic of the conversation, providing a complete context representation for subsequent generation.

[0102] In the embodiments of the present invention, through the synergistic effect of semantic encoding and position modeling, a context representation that combines in - depth semantic understanding and precise temporal reasoning is provided for the dialogue system, which is the core basis of the parallel generation mechanism.

[0103] In traditional technologies, encoders based on Bi - GRU or unidirectional RNNs are difficult to effectively model cross - sentence dependencies in long conversations. For example, in scenarios where users ask multiple follow - up questions or the topic changes, traditional models may generate off - topic responses due to ignoring early key information. In addition, retrieval - based models rely on local matching and cannot dynamically capture the global dialogue logic. Based on this, an alternative embodiment of the present invention is proposed.

[0104] Step S1015 further includes the following specific embodiments.

[0105] Step S10151: Input each of the concatenated context representations into a multi - head self - attention module to generate the global vector corresponding to the conversation context that includes global semantic information.

[0106] In an embodiment of the present invention, each spliced context representation (including the utterance feature vector and the position embedding) is input into the multi-head self-attention module. By calculating the correlation weights between each utterance feature and other utterance features in parallel, the semantic information at different positions in the dialogue history is aggregated; the features are weighted and summed based on the weights to generate a vector representation that fuses the global semantics; finally, the global vector of the dialogue context is output, which completely represents the overall semantics and the context logical relationship of the dialogue.

[0107] In an embodiment of the present invention, the long-distance dependence relationship across utterances in the dialogue history is captured through the multi-head self-attention mechanism, effectively avoiding the problem of semantic attenuation caused by the sequence length limitation of the traditional recurrent neural network. The multi-head mechanism allows the model to independently learn dialogue features from different subspaces (such as grammar, semantics, emotion, etc. dimensions), enhancing the ability to understand complex contexts and improving the accuracy and relevance of the responses. Combining the position embedding and the multi-head attention, while integrating the global semantics, the utterance order information is retained, avoiding the logical confusion problem caused by insufficient position encoding in traditional methods.

[0108] Traditional generative models (such as Seq2Seq) need to generate responses word by word iteratively, resulting in a high response delay and being difficult to meet the requirements of real-time interaction. At the same time, word-by-word generation is easily affected by local probability biases, resulting in inconsistent logic or semantic redundancy before and after the response. For example, in a long dialogue scenario, the traditional model may generate responses that deviate from the topic due to insufficient modeling of the global dependence. Based on this, an alternative embodiment of the present invention is proposed.

[0109] Refer to Figure 3 , Figure 3 which is a schematic diagram of a specific embodiment of step S102 of the parallel generation method of the response text in an embodiment of the present invention. Step S102 further includes the following specific embodiments.

[0110] Step S1021, calculate the degree of correlation between the initial response sequence and the global vector to obtain an intermediate result that fuses the context information.

[0111] Step S1022, input the intermediate result into a recurrent neural network unit with a memory function to obtain the hidden representation of the word to be generated.

[0112] Step S1023, according to the hidden representation, calculate the generation probability of each candidate word to form a candidate word list sorted by probability.

[0113] Step S1024, according to the candidate word list, generate a plurality of the candidate sequences.

[0114] In an embodiment of the present invention, during the parallel processing of the A-G module, the semantic correlation degree between the initial reply sequence and the global vector is calculated through an attention mechanism to generate an intermediate feature representation that fuses context information; subsequently, the intermediate features are input into a recurrent neural network unit with a memory function to dynamically capture the state dependencies during the sequence generation process, and a context-aware hidden representation of the word to be generated is output; based on the hidden representation, a prediction probability of each candidate word is generated through a probability distribution calculation module to form a list of candidate words sorted by probability; finally, according to the diversity strategy of the candidate word list, multiple semantically coherent candidate sequences are generated in parallel.

[0115] In an embodiment of the present invention, through the parallel processing mechanism of the A-G module, multiple candidate words can be generated in a single iteration, significantly reducing the number of iterations required for traditional word-by-word generation. During the generation process, the overall semantic coherence of multiple candidate sequences is evaluated synchronously, effectively avoiding the semantic discontinuity or repetition problems caused by local optimality in traditional methods, and improving the naturalness and logical rationality of the reply. Through the synergistic effect of the attention mechanism and the recurrent neural network, the generation of each candidate word fully combines historical dialogue information, enhancing the context relevance of the reply.

[0116] During the word-by-word generation process of traditional generative models (such as Seq2Seq), the loss function is usually calculated using uniform weights or fixed decay weights, resulting in insufficient attention of the model to the latter part of the sequence. For example, when generating long replies, traditional methods may suffer from semantic repetition or logical discontinuity due to the lack of global constraints in the generation of later tokens. Based on this, an alternative embodiment of the present invention is proposed.

[0117] Refer to Figure 4 , Figure 4 FIG. is a schematic diagram of a specific embodiment of step S103 of the parallel generation method of the reply text in an embodiment of the present invention, and step S103 further includes the following specific embodiments.

[0118] Step S1031: Assign gradually increasing weight values to each unit in the candidate word list.

[0119] Step S1032: Multiply the word probability corresponding to each unit in the candidate word list by the weight value corresponding to each unit and then accumulate to obtain multiple comprehensive probabilities.

[0120] Step S1033: Select multiple candidate sequences according to the multiple comprehensive probabilities.

[0121] In an embodiment of the present invention, for each generation unit (corresponding to candidate words at different positions) in the candidate word list, gradually increasing weight values are dynamically assigned according to its position index to form a weight gradient distribution.

[0122] Multiply the generation probability corresponding to each candidate word unit by its assigned weight value, and accumulate the weighted probabilities of all units to generate a comprehensive probability score for multiple candidate sequences.

[0123] Sort the candidate sequences based on the comprehensive probability score, and select multiple candidate sequences with the highest scores as the candidate set for the final response, ensuring that the generation results have both diversity and quality.

[0124] In the embodiments of the present invention, through a gradually increasing weight assignment mechanism, the importance of the words generated in the later stage of the sequence is strengthened, effectively avoiding the local optimum problem caused by uniform weight assignment in traditional methods. According to the application scenario requirements, by adjusting the weight growth parameter (such as the base of the exponential function), the attention of the model to different parts of the sequence can be flexibly controlled. Through weight regularization processing (such as sum normalization), the gradient update intensity of the generation units at different positions is balanced, effectively avoiding the problem of difficult model convergence caused by too large weight differences during the training process.

[0125] In traditional technologies, generative models usually adopt fixed weight assignment (such as equal weights or linearly decaying weights), which are difficult to dynamically adapt to the reply generation requirements of different lengths or scenarios. For example, when generating a long response, traditional methods may ignore key semantic information (such as conclusive statements) due to insufficient weights of later tokens, resulting in incomplete reply logic. Based on this, an alternative embodiment of the present invention is proposed.

[0126] Step S103 also includes the following specific implementation manners.

[0127] Step S1034, according to the position index of each unit in the candidate word list and a preset positive real number parameter, calculate the initial weight value through an exponential function, and the greater the position index, the greater the growth amplitude of the corresponding initial weight value.

[0128] Step S1035, divide each initial weight value by the sum of all initial weight values to obtain the weight value corresponding to each unit.

[0129] In the embodiments of the present invention, according to the position index of each unit in the candidate word list (such as the 0th, 1st, etc. in the generation order) and a preset positive real number parameter, calculate the initial weight value through an exponential function, so that the weight growth amplitude of the unit with a larger position index (corresponding to the words generated later in the sequence) is significantly increased.

[0130] Optionally, according to the position index of each unit in the candidate word list and a preset positive real parameter, calculate the initial weight value through an exponential function, and the greater the position index of each unit, the greater the growth amplitude of the corresponding initial weight value; divide each initial weight value by the sum of all initial weight values to obtain the weight value corresponding to each unit. Specifically, sum the initial weight values of all units, and divide each initial weight value by this sum to obtain the normalized weight value, ensuring that the sum of all weights is 1 while retaining the exponential growth trend.

[0131] In the embodiments of the present invention, through the combination of the exponential function and normalization processing, not only the high attention to the second half of the sequence is retained, but also the problem of unbalanced training gradients caused by excessive weight differences is avoided, improving the convergence stability of the model. By adjusting the preset positive real parameter (such as the base of the exponent), the weight growth rate can be flexibly controlled to meet the requirements of different scenarios. Assigning higher weights to the second half of the sequence effectively alleviates the semantic break problem caused by insufficient long-distance dependence modeling in traditional methods, improving the logical coherence of generating long responses.

[0132] Traditional dialogue systems usually cannot handle newly input content by the user during the process of generating responses, resulting in the generated responses not matching the user's latest intention. For example, in scenarios where the user quickly asks follow-up questions or corrects the problem, the traditional model will continue to complete the original generation process and output responses that are disjointed from the current conversation. In addition, the prior art lacks the ability to extract semantic fragments of uncompleted responses, and directly discarding the generated content will cause information loss. Based on this, the present invention proposes an optional embodiment.

[0133] After step S103, the following specific implementation manners are further included.

[0134] Step S201, if a response operation is being executed and new content input by the user is detected, stop the response operation, and update the dialogue context according to the complete semantic fragments in the currently replied content, where the currently replied content is the content that is being replied but not completed;

[0135] Step S202, return to execute the step of generating the global vector corresponding to the dialogue context.

[0136] In the embodiments of the present invention, during the process of generating response content, continuously monitor the user input status. If new content input by the user is detected, immediately terminate the currently executing response generation operation. Perform semantic integrity analysis on the generated but uncompleted response content to identify complete semantic fragments that conform to grammar and logic (such as complete sentences or key phrases).

[0137] Merge the extracted complete semantic fragments with the new content input by the user, update the dialogue context, and form a new dialogue history record. Based on the updated dialogue context, re-execute the global vector generation process to ensure that subsequent responses are highly relevant to the latest dialogue state.

[0138] In the embodiments of the present invention, through an immediate interruption mechanism, redundant responses unrelated to the new input by the user are avoided, significantly improving the dialogue fluency. When dynamically updating the dialogue context, the effective semantic fragments of the generated content are retained to prevent logical discontinuities caused by interruptions, ensuring the coherent connection of multi-turn dialogues. The invalid generation process is terminated in a timely manner to reduce waste of computing resources, especially in high-concurrency scenarios, which can effectively reduce the system load.

[0139] Traditional dialogue systems usually fixedly retain all historical dialogues or adopt a simple time window truncation strategy, resulting in two problems: one is that in long dialogue scenarios, the model responds slowly due to processing too much irrelevant historical information; the other is that key contexts are wrongly truncated (for example, when the user corrects their requirements multiple times). For example, in a medical consultation scenario, the traditional method may generate inaccurate suggestions due to not retaining the early symptom description. Based on this, an alternative embodiment of the present invention is proposed.

[0140] Step S1011 further includes the following specific embodiments.

[0141] Step S1016, if a reply operation is being executed and new content input by the user is detected, stop the reply operation, and update the dialogue context according to the complete semantic fragments in the currently replied content, where the currently replied content is the content that is being replied but not yet completed;

[0142] Step S1017, return to execute the step of generating the global vector corresponding to the dialogue context.

[0143] In the embodiments of the present invention, when new information input by the user is detected, the complete dialogue history record is extracted from the storage medium, including the current input and the multi-turn interaction content before it.

[0144] According to the semantic focus of the current input information (such as keywords, intention classification results), the most relevant continuous N rounds of dialogues (for example, retain the last 3 rounds of dialogues) are screened out from the historical dialogues to form a refined dialogue context.

[0145] Based on the screened N-round dialogue context, an encoder is used to generate a global vector representing the core semantics of the current dialogue, which serves as the basis for subsequent reply generation.

[0146] In the embodiments of the present invention, by dynamically retaining N rounds of key conversations, interference from redundant historical information is avoided. The problem of wasted computing power caused by processing all historical conversations in traditional methods is solved, and the computational complexity is significantly reduced while ensuring semantic coherence. The configurability of the N value supports flexible adaptation to different application requirements.

[0147] As Figure 5 shown, it is a schematic diagram of a terminal device provided by an embodiment of the present invention. The terminal device 500 may include: a processor 501, a memory 502, and a computer program 503 stored in the memory 502 and executable on the processor 501, such as a parallel generation program for reply text. When the processor 501 executes the computer program 503, the steps in the above-mentioned parallel generation embodiments of each reply text are implemented.

[0148] The computer program may be divided into one or more modules / units, and one or more modules / units are stored in the memory 502 and executed by the processor 501 to complete the present invention. One or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the terminal device.

[0149] The terminal device may include, but is not limited to, the processor 501 and the memory 502. Those skilled in the art can understand that Figure 5 merely examples of the terminal device do not constitute a limitation to the terminal device, and it may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the terminal device may further include input / output devices, network access devices, buses, etc.

[0150] The so-called processor 501 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0151] The memory 502 can be an internal storage unit of the terminal device, such as the hard disk or memory of the terminal device. The memory 502 can also be an external storage device of the terminal device, such as a plug-in hard disk equipped on the terminal device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 502 can also include both the internal storage unit and the external storage device of the terminal device. The memory 502 is used to store computer programs and other programs and data required by the terminal device. The memory 502 can also be used to temporarily store the data that has been output or will be output.

[0152] It should be noted that for the convenience and conciseness of description, the structure of the above terminal device can also refer to the specific description of the structure in the method embodiment, which will not be elaborated here.

[0153] The embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above method for parallel generation of reply text can be implemented.

[0154] The embodiment of the present invention provides a computer program product. When the computer program product runs on a mobile terminal, the mobile terminal can execute the steps in the above method for parallel generation of reply text.

[0155] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0156] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0157] In the embodiments provided by the present invention, it should be understood that the disclosed terminal device and method can be implemented in other ways. For example, the above-described terminal device embodiments are merely illustrative. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0158] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0159] In addition, each functional unit in various embodiments of the present invention may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0160] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0161] The above-mentioned embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A parallel generation method for reply text, characterized in that, Including: When input information is detected, generate a global vector corresponding to the dialogue context, and generate an initial response sequence corresponding to the global vector; Input the global vector and the initial response sequence into the A-G module for parallel processing to obtain multiple candidate sequences; Generate a response content according to the candidate sequences; Among them, the step of inputting the global vector and the initial response sequence into the A-G module for parallel processing to obtain multiple candidate sequences includes: Calculate the degree of association between the initial response sequence and the global vector to obtain an intermediate result integrating context information; Input the intermediate result into a recurrent neural network unit with a memory function to obtain a hidden representation of the word to be generated; According to the hidden representation, calculate the generation probability of each candidate word to form a candidate word list sorted by probability; Generate multiple candidate sequences according to the candidate word list; Among them, the step of generating a response content according to the candidate sequences includes: Assign gradually increasing weight values to each unit in the candidate word list; Multiply the word probability corresponding to each unit in the candidate word list by the weight value corresponding to each unit and accumulate to obtain multiple comprehensive probabilities; Select multiple candidate sequences according to the multiple comprehensive probabilities; Among them, the step of assigning gradually increasing weight values to each unit in the candidate word list includes: According to the position index of each unit in the candidate word list and a preset positive real number parameter, calculate the initial weight value through an exponential function, and the greater the position index, the greater the growth amplitude of the corresponding initial weight value; Divide each initial weight value by the sum of all initial weight values to obtain the weight value corresponding to each unit.

2. The parallel generation method of the response text according to claim 1, wherein, When input information is detected, generating a global vector corresponding to the dialogue context further includes: When input information is detected, obtain the dialogue context; Encode each sentence in the dialogue context through a bidirectional gated recurrent unit to obtain a discourse feature vector, and assign a continuous position number to each sentence to obtain a position number; According to each position number, calculate the position embedding PE of each position through alternating sine waves and cosine waves; Concatenate each position embedding PE with each discourse feature vector to obtain a context representation including the dialogue order; Generate the global vector corresponding to the dialogue context according to each context representation.

3. The parallel generation method of the response text according to claim 2, wherein The step of generating the global vector corresponding to the dialogue context according to each context representation includes: Input each concatenated context representation into a multi-head self-attention module to generate the global vector corresponding to the dialogue context including global semantic information.

4. The parallel generation method of the response text according to claim 1, wherein After the step of generating a response content according to the candidate sequences, the method further includes: If a response operation is being executed and new content input by the user is detected, stop the response operation, and update the dialogue context according to the complete semantic segment in the currently replied content, where the currently replied content is the content that is being replied but not yet fully replied; Return to execute the step of generating a global vector corresponding to the dialogue context.

5. The parallel generation method of the response text according to claim 1, wherein When the input information is detected, the steps of generating a global vector corresponding to the dialogue context include: When the input information is detected, obtain the historical dialogue context; According to the input information, retain N rounds of conversations in the historical dialogue context as the dialogue context; Generate the global vector corresponding to the dialogue context.

6. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the method for parallel generation of reply texts as described in any one of claims 1 to 5 are implemented.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method for parallel generation of reply texts as described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Intelligent dialogue method and system

    CN114860910A