Reasoning method and model training method, device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202611215602.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-11
- Publication Date
- 2026-09-25
AI Technical Summary
[0010]根据本公开的另一方面,提供了一种计算机程序产品,包括计算机程序,计算机程序在被处理器执行时实现根据本公开提供的方法。
Smart Images

Figure CN122817880A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, particularly to the fields of Large Language Model (LLM), Multi-Token Prediction (MTP), and reasoning, and can be applied to scenarios such as intelligent question answering, text-to-image generation, and text-to-video generation. More specifically, this disclosure provides a reasoning method, a model training method, an apparatus, an electronic device, and a storage medium. Background Technology
[0002] With the development of artificial intelligence technology, the application of large language models is constantly increasing. In the reasoning stage, large language models can generate output token by token. To improve token generation efficiency, multi-token prediction technology can be used to predict tokens at multiple future positions simultaneously. Summary of the Invention
[0003] This disclosure provides an inference method, a model training method, an apparatus, a device, and a storage medium.
[0004] According to one aspect of this disclosure, a reasoning method is provided, comprising: determining the unprocessed representation features of the lexical to be processed, wherein the lexical to be processed is a lexical determined by a large model based on a preceding lexical sequence, the preceding lexical sequence including the lexical sequence of the input text; fusing the preceding latent state features and the unprocessed representation features to obtain fused features, wherein the preceding latent state features are the latent state features of the preceding lexical in the preceding lexical sequence, the preceding latent state features being determined by the target processing layer of the large model; inputting the fused features into the target processing layer to obtain the unprocessed latent state features; and determining the subsequent candidate lexicals of the lexical to be processed based on the unprocessed latent state features.
[0005] According to one aspect of this disclosure, a model training method is provided, comprising: inputting sample lexical units to be processed into the embedding layer of a large model to obtain sample representation features of the sample lexical units to be processed, wherein the sample lexical units to be processed are determined based on sample response lexical units, the sample response lexical units are lexical units in the sample response sequence, the sample response sequence is the lexical sequence of the sample response result, and the sample response result is the label of the input sample text; fusing the pre-hidden state features and the representation features to be processed to obtain fused features, wherein the pre-hidden state features are the hidden state features for the preceding lexical units, the hidden state features for the preceding lexical units are determined by the target processing layer of the large model, and the preceding lexical units are the lexical units preceding the sample lexical units to be processed; inputting the fused features into the target processing layer to obtain the hidden state features to be processed; determining the subsequent candidate lexical units of the sample lexical units to be processed based on the hidden state features to be processed; determining the loss information for the subsequent candidate lexical units based on the subsequent candidate lexical units and the sample response lexical units for the subsequent candidate lexical units in the sample response sequence; and training the large model based on the loss information for the subsequent candidate lexical units.
[0006] According to another aspect of this disclosure, an inference apparatus is provided, comprising: a first determining module for determining unprocessed representation features of a word to be processed, wherein the word to be processed is a word determined by a large model based on a preceding word sequence, the preceding word sequence including a word sequence of the input text; a first fusion module for fusing the preceding hidden state features and the unprocessed representation features to obtain fused features, wherein the preceding hidden state features are the hidden state features of preceding words in the preceding word sequence, the preceding word sequence being determined by a target processing layer of the large model; a first processing module for inputting the fused features into the target processing layer to obtain unprocessed hidden state features; and a second determining module for determining subsequent candidate words of the word to be processed based on the unprocessed hidden state features.
[0007] According to another aspect of this disclosure, a model training apparatus is provided, comprising: an acquisition module for inputting sample lexical units to be processed into the embedding layer of a large model to obtain sample representation features of the sample lexical units to be processed, wherein the sample lexical units to be processed are determined based on sample response lexical units, the sample response lexical units are lexical units in the sample response sequence, the sample response sequence is the lexical sequence of the sample response result, and the sample response result is the label of the input sample text; and a second fusion module for fusing the pre-hidden state features and the representation features to be processed to obtain fused features, wherein the pre-hidden state features are for the hidden state of the pre-lexical units. The features are defined as follows: the latent state features of the preceding word units are determined by the target processing layer of the large model; the preceding word units are the word units before the word units of the sample to be processed; the second processing module is used to input the fused features into the target processing layer to obtain the latent state features to be processed; the fifth determination module is used to determine the subsequent candidate word units of the sample to be processed based on the latent state features to be processed; the sixth determination module is used to determine the loss information for the subsequent candidate word units based on the subsequent candidate word units and the sample response word units in the sample response sequence for the subsequent candidate word units; the training module is used to train the large model based on the loss information for the subsequent candidate word units.
[0008] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method provided according to this disclosure.
[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing a computer to perform the methods provided according to this disclosure.
[0010] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method provided according to this disclosure.
[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0012] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0013] Figure 1 This is a flowchart of a reasoning method according to an embodiment of the present disclosure;
[0014] Figure 2This is a schematic diagram of a reasoning method according to an embodiment of the present disclosure;
[0015] Figure 3 This is a flowchart of a model training method according to an embodiment of the present disclosure;
[0016] Figure 4 This is a block diagram of a reasoning apparatus according to an embodiment of the present disclosure;
[0017] Figure 5 This is a block diagram of a model training apparatus according to an embodiment of the present disclosure; and
[0018] Figure 6 This is a block diagram of an electronic device according to an embodiment of the present disclosure, which can be applied to at least one of inference methods and model training methods. Detailed Implementation
[0019] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0020] Large language models widely adopt Next Token Prediction (NTP) as the basic training objective. This training method performs self-supervised learning by predicting the next token in the sequence, and has achieved good training results on large-scale corpora.
[0021] In some next-word prediction techniques, for the input sequence Large models can predict the next lexical unit: This technique boasts a simple structure and stable training, making it a mainstream paradigm for pre-training large language models. Based on next-word prediction, a supervisory signal can be generated at each position to monitor the next future position. However, with only one supervisory signal generated at each position, the future information that large models can utilize is limited, making it difficult to further mine supervisory signals from the training data and restricting their ability to model long-distance future dependencies. Furthermore, the need to generate output per word during the inference phase results in low generation efficiency.
[0022] To enhance a model's ability to utilize future information, multi-term prediction technology has emerged as an auxiliary extension to the next term prediction technique. Multi-term prediction can simultaneously predict terms at multiple future positions, providing the model with denser supervision signals. It not only provides draft sequences for speculative decoding, accelerating inference, but also enhances the model's ability to model future information by predicting targets at multiple future positions. In particular, supervision signals from multiple future positions help the model plan its internal representation in advance, enabling it to better predict subsequent content.
[0023] In some multi-lexical prediction techniques, a parallel prediction structure can be employed to simultaneously predict lexical units at multiple future positions. A large language model is used as the main model, and a future lexical prediction head is introduced based on the main model representation. Lexical units at multiple future positions are then predicted simultaneously based on the same historical hidden state. N can be an integer greater than 1. Compared to single-word prediction, multi-word prediction can introduce denser supervision signals into a single training sample, increasing the supervision signal density and improving the efficiency of training data utilization. Simultaneously, the predicted future word sequence can also serve as a draft sequence in the speculative decoding process during the inference stage, thereby improving inference speed. However, this type of method performs parallel predictions for multiple future positions, which lack causal relationships between these prediction positions. Each prediction target is independent, making it difficult to accurately characterize the sequence dependencies in the natural language generation process, model the stepwise conditional dependencies, and potentially lead to a lack of consistent constraints on the learning signals between different prediction targets, thus weakening the effectiveness of the supervision signal.
[0024] To improve the quality of future word predictions and the acceptance rate of speculative decoding, a causal chain-based multi-word prediction technique can be adopted. By preserving the causal dependencies between multiple future prediction positions, subsequent predictions can utilize information from preceding positions, and the future word prediction process can have a recursive dependency structure, thereby enhancing the coherence and consistency of future predictions and thus improving the quality of future word predictions and the acceptance rate of speculative decoding.
[0025] In some causal chain-based multi-term prediction techniques, predictions for multiple future terms can be achieved through a future term prediction module independent of the main model. Unlike methods that predict multiple future terms in parallel, this type of method makes the prediction of the (k+1)th future term dependent on the prediction state corresponding to the kth future term, thereby establishing a causal relationship between multiple future prediction terms. By preserving the causal chain structure in the future prediction process, this method can improve the accuracy of future term predictions and further enhance the acceptance rate and inference performance of speculative decoding.
[0026] However, causal chain-based multi-term prediction technology can achieve prediction of multiple future positions through a future term prediction module attached to the main model. The future term prediction process is separated from the main model decoding process. On the one hand, differences in parameter space and representation space exist between the future term prediction module and the main model, making it difficult to fully align the future term prediction distribution with the main model's prediction distribution, thus reducing the draft acceptance rate of speculative decoding during the inference stage. On the other hand, the supervisory signals generated at multiple future positions first act on the future term prediction module, and then indirectly affect the main model parameters through backpropagation. They fail to directly affect the autoregressive generation path of the main model, making it difficult to fully leverage the role of future supervisory signals in improving the main model's representation learning and long-range planning capabilities. This weakens the constraint effect of future information on the main model's representation learning, limiting the model's ability to model long-range dependencies and effectively utilize future information.
[0027] Therefore, in order to fully leverage the advantages of future lexical prediction in supervised signal enhancement and long-range planning learning, achieve unified optimization of training and inference, improve model generation quality and inference decoding efficiency, this disclosure provides an inference method and a model training method. The inference method of this disclosure will be described below.
[0028] Figure 1 This is a flowchart of a reasoning method according to an embodiment of the present disclosure.
[0029] like Figure 1 As shown, the method 100 may include operations S110 to S140.
[0030] In operation S110, the representation features of the lexical units to be processed are determined.
[0031] In embodiments of this disclosure, the inference method can be performed by a large model. The large model can receive input data in one or more modalities, including input text.
[0032] In this embodiment, the lexical unit to be processed can be a lexical unit determined by the large model based on the preceding lexical unit sequence. Taking input data as input text as an example, the lexical unit to be processed can be a lexical unit generated during the reasoning process of the large model based on the input text. The lexical unit to be processed can be embedded to obtain the representation features to be processed.
[0033] The preceding lexical sequence can include the lexical sequence of the input text. For example, the input text could be "The weather is nice today". This input text is then converted into a lexical sequence. . It can be a word group for "today". It can be a lexical unit for "weather". It can be a lexical unit for "true". which can be a token "hao". The large model performs reasoning based on the token sequence of the input text, for example, to obtain the token as the token to be processed. The token can be a token "a". Embedding processing is performed on the token to obtain the representation feature of the token , which is used as the to-be-processed representation feature. The to-be-processed token and the token sequence of the input text can be concatenated to obtain the current token sequence .
[0034] In operation S120, the previous hidden state feature and the to-be-processed representation feature are fused to obtain a fused feature.
[0035] In the embodiment of the present disclosure, the previous hidden state feature is a hidden state feature for a previous token, and the hidden state feature of the previous token is determined by a target processing layer of the large model.
[0036] The previous token may be the last token in the current token sequence. Concatenating the previous token sequence and the to-be-processed token can obtain the current token sequence. The previous token may also be a token before the to-be-processed token in the current token sequence. For example, the large model may include a processing network and an output network. The processing network may include multiple processing layers. The output network may also be referred to as an output head. The processing layer may be a Transformer layer. Any one of the multiple processing layers may serve as the target processing layer. As mentioned above, after the large model decodes the token the current token sequence may be . The previous hidden state feature may be a hidden state feature for the token determined by the target processing layer. Fusing the hidden state feature for the token and the representation feature of the token can obtain the fused feature. The fusion method may be various fusion methods such as concatenation or addition.
[0037] In operation S130, the fused feature is input to the target processing layer to obtain a to-be-processed hidden state feature.
[0038] For example, after fusing the hidden state feature for the token and the representation feature of the token the obtained fused feature can be input to the target processing layer to obtain the to-be-processed hidden state feature. It can be understood that the to-be-processed hidden state feature can be used as the hidden state feature of the token .
[0039] In operation S140, a subsequent candidate token of the to-be-processed token is determined according to the to-be-processed hidden state feature.
[0040] In this embodiment of the disclosure, the latent state features to be processed are input into the output network of a large model to obtain subsequent candidate words. For example, when targeting words... Hidden state features and lexical units After feature fusion, the fused features are input into the target processing layer to obtain the latent state features to be processed. These latent state features are then input into the output head of the large model to obtain the lexical units. , as a candidate word element.
[0041] In this embodiment, candidate lexical generation is primarily based on the latent state features output by the target processing layer of the large model. A Transformer layer is reused within the unified parameter space of the large model for recursive multi-lexical generation. This fully leverages the causal relationships between characters in the input text, as well as the relationships and semantic connections between generated lexical units and their predecessors. This ensures that the multi-lexical prediction process and the large model's autoregressive generation process share the same latent state evolution path. The predicted distribution of future lexical units always depends on the consistent representation space within the large model, reducing the structural offset between the predicted and generated distributions. This improves the acceptance rate and stability of candidate lexical units generated during speculative decoding in the inference stage, effectively enhancing inference efficiency.
[0042] It is understood that the reasoning method of this disclosure has been explained above, and the reasoning method of this disclosure will be further explained below.
[0043] Figure 2 This is a schematic diagram of a reasoning method according to an embodiment of the present disclosure.
[0044] like Figure 2 As shown, the large model 20 includes an embedding layer 201, L decoder layers, and an output head 203. The L decoder layers include the first L-1 decoder layers 2021 and the Lth decoder layer 2022. The large model 20 may also include a fusion layer 204. L can be an integer greater than 1.
[0045] In some embodiments, the target processing layer is the last processing layer among multiple processing layers of a large model. For example, the processing layer of the large model could be a decoding layer. Figure 2 As shown, the Lth decoding layer can be used as the target processing layer.
[0046] The word sequence of the input text can be fed into a large model. The word sequence of the input text can be... .like Figure 2 As shown, the large model 20 can encode the word sequence of the input text to obtain words. Hidden state features For example, lexical units can be determined using the following formula. Hidden state features :
[0047] (Formula 1)
[0048] It can represent L decoding layers. Representing the above lexical sequence Embedded features.
[0049] Lexicons Hidden state features Input to output header 203 to obtain the decoded tokens. For example, lexical units can be determined using the following formula. :
[0050] (Formula 2)
[0051] Indicates: based on the current hidden state features Predict the next word element. The probability. For example, the probability of multiple lexical units can be determined, and based on the probabilities of multiple lexical units, a lexical unit can be selected from among them. Lexical units These can be used as lexical units to be processed. Lexical unit Hidden state features This can be used as a feature in the pre-hidden state. It can be used to compare the word sequence of the input text with the word... By concatenating the sequences, we obtain a word sequence. , which serves as the current word sequence.
[0052] Next, the above operations S110~S130 can be performed to convert the word elements. As the lexical units to be processed, they are input into the embedding layer 201 to obtain the lexical units. The representation features are used as the representation features to be processed. Lexical units can be used as representation features. Representational features and lexical units Hidden state features The inputs are fed into the fusion layer 204 to combine the tokens. Representational features and lexical units Hidden state features The features are then fused together to obtain the fused features. These fused features can be input into the Lth decoding layer (2022) to obtain the hidden state features. These are the latent state features to be processed. For example, the latent state features can be obtained using the following formula. :
[0053] (Formula 3)
[0054] represents the L-th decoding layer. represents a fused feature.
[0055] Next, operation S140 may be performed to input the hidden state feature into the output head 203 to obtain a token subsequent to token , which is used as a subsequent candidate token. For example, the token may be determined by the following formula: :
[0056] (Formula 4)
[0057] represents: based on the current hidden state feature , predict the probability that the next token is .
[0058] According to the embodiments of the present disclosure, the last processing layer of a large model is used as the target processing layer, and no independent multi-token prediction branch network or external prediction module is introduced, which avoids decoupling the future token prediction process from the autoregressive generation process of the large model. The multi-token prediction process is internalized into a recursive expansion form of the autoregressive generation process of the large model, enabling the model to complete step-by-step prediction and state update in a unified generation link, thereby realizing continuous expansion of multi-step generation. In each state update step, joint modeling is performed on the current hidden state feature and the embedding representation of the corresponding token, and state update is performed through the last Transformer layer of the large model and the fusion layer connected thereto, thereby implementing a recursive state update mechanism based on token feedback, which can effectively reduce inference complexity, reduce computing resource overhead, and improve the acceptance rate and stability of candidate tokens during the speculative decoding process in the inference phase.
[0059] It can be understood that the model principle of the present disclosure has been described above, and the fusion layer of the present disclosure will be described below.
[0060] As shown in Figure 2 , the fusion layer 204 may include a normalization module 2041, a normalization module 2042, a concatenation module 2043, and a linear projection module 2044. The normalization module 2041 and the normalization module 2042 may implement one of various normalization processing. For example, the normalization module may implement layer normalization processing. The layer normalization processing may be Root Mean Square Layer Normalization (RMSNorm).
[0061] In some embodiments, fusing the pre-hidden state features and the representation features to be processed to obtain fused features includes: normalizing the pre-hidden state features to obtain pre-normalized features; normalizing the representation features to be processed to obtain the representation features to be processed; and concatenating the pre-normalized features and the representation features to be processed to obtain fused features.
[0062] With the above-mentioned morphemes For example, such as Figure 2 As shown, the word elements Input embedding layer 201 to obtain lexical units The representational features of lexical units. The representation features are input into the normalization module 2041 to obtain the word elements. The normalized features are used as the normalized features to be processed. The lexical units... Hidden state features Inputting the data into the normalization module 2042 yields the hidden state features. The normalized features are used as the pre-normalized features. Lexical units are... Normalized features and latent state features The normalized features are input into the stitching module 2043 to obtain the stitched features. The stitched features are then input into the linear projection module 2044 to obtain the linear projection features, which are used as the fusion features.
[0063] Through the embodiments disclosed herein, the hidden state is updated by using the last Transformer layer of the large model and its connected normalization and linear mapping modules, thereby effectively realizing a recursive state update mechanism based on lexical feedback, which can further improve the acceptance rate of candidate lexical units.
[0064] In some embodiments, the subsequent candidate lexical units can be treated as lexical units to be processed, and the process can return to the operation of determining the representation features to be processed of the lexical units to be processed, until the subsequent candidate lexical unit is the final lexical unit. For example, when obtaining lexical units... Then, the word elements can be... As a word element to be processed, latent state features As a feature of the preceding hidden state, return to operation S110 above, and repeat operations S110 to S140 until the end word is decoded. This is to determine the word. Subsequent morphemes For example, determine the word elements. The method of determining lexical units The methods used are the same or similar, and will not be elaborated further in this disclosure. Furthermore, in determining the lexical units... During the process, lexical units can be obtained. Representational features and latent state features The fusion feature can be input into the Lth decoding layer to obtain the hidden state feature. , as a word element Hidden state features To determine the lexical units Subsequent morphemes For example, determine the word elements. The method of determining lexical units The methods used are the same or similar, and will not be elaborated further in this disclosure. Furthermore, in determining the lexical units... During the process, lexical units can be obtained. Representational features and latent state features The fusion feature can be input into the Lth decoding layer to obtain the hidden state feature. , as a word element Hidden state features Hidden state features You can input it into the output head of a large model to obtain lexical units. Therefore, inference can be completed with significantly reduced computing resources.
[0065] After decoding the end word, the resulting M words can include: M can be an integer greater than or equal to 3. Based on lexical sequences. This allows us to determine the response to the input text.
[0066] It is understood that the above provides a method for inference with significantly reduced computing resources. In other embodiments, one or more inference cycles can be set, and after the end of an inference cycle, the candidate tokens decoded in that inference cycle can be verified to improve inference accuracy and thus improve user experience. The verification method will be described below.
[0067] In some embodiments, the preceding lexical sequence includes the baseline lexical sequence for the current inference cycle, which includes the lexical sequence of the input text. For example, at the beginning of the first inference cycle, the lexical sequence of the input text can be used as the baseline lexical sequence for the current inference cycle. Using the aforementioned lexical sequence of the input text... As an example, based on this lexical sequence, the lexical units can be determined. This is understandable; the above text regarding the identification of word elements... The same description applies to this embodiment, and will not be repeated here.
[0068] Next, we can use word elements. As the first word to be processed in the first inference cycle, it is concatenated with the baseline word sequence of the first inference cycle to obtain the word sequence. As a target word element The current word sequence. Operations S110 to S140 can be performed to obtain the word sequence. This is understandable; the above text regarding obtaining lexical units... The same description applies to this embodiment, and will not be repeated here.
[0069] The above text will discuss word elements. Unlike the next candidate lexical unit to be processed, in some embodiments, the above method may further include: fusing the current lexical unit sequence and the subsequent candidate lexical units to obtain the subsequent lexical unit sequence. For example, lexical units... With lexical sequence This allows us to obtain the word sequence. , as the subsequent word sequence.
[0070] Next, it can be determined whether the conditions for continuing inference are met in the subsequent lexical sequence. The conditions for continuing inference can include at least one of the following: the number of candidates is less than a preset threshold for the number of lexical units; or a terminator exists in the subsequent lexical sequence. The number of candidates is the number of candidate lexical units in the subsequent lexical sequence. For example, a lexical sequence... in, word element As a candidate word, the number of candidates in this word sequence can be determined to be 1. With a preset word quantity threshold of 3 and word... Taking a non-ending lexical as an example, the lexical sequence can be determined. The conditions for continuing the reasoning are met.
[0071] In response to determining that the subsequent word sequence meets the conditions for continuing inference, the subsequent candidate word is treated as a word to be processed, and the process returns to the operation of determining the unprocessed representation features of the word to be processed, so that inference can continue in the current inference cycle until the subsequent word sequence no longer meets the conditions for continuing inference. For example, in the word sequence If the conditions for continued reasoning are met, then this lexical sequence is used as a target lexical. The current word sequence, word As the second unprocessed word in the first inference cycle, the above operations S110 to S140 are executed again to obtain the word. , as a word element The candidate word group is the one following the word group. With lexical sequence By concatenating the sequences, we obtain a word sequence. . with word elements Taking the terminator as an example, we can determine the terminator sequence. The conditions for continued reasoning are met. This can be understood as determining the lexical units. The method of determining lexical units The methods used are the same or similar, and will not be elaborated further in this disclosure.
[0072] In lexical sequence If the conditions for continued reasoning are met, then this lexical sequence is used as a target lexical. The current word sequence, word As the third unprocessed word in the first inference cycle, the above operations S110 to S140 are executed again to obtain the word. , as a word element The candidate word group is the one following the word group. With lexical sequence By concatenating the sequences, we obtain a word sequence. It's understandable that determining the lexical units is necessary. The method of determining lexical units The methods used are the same or similar, and will not be elaborated further in this disclosure.
[0073] In some embodiments, the above method may further include: in response to determining that the subsequent word sequence does not meet the conditions for continuing inference, inputting the subsequent word sequence into a large model to obtain at least one verification result, the verification result indicating whether the candidate word is determined as an output word by the large model. For example, after obtaining the word sequence... Subsequently, it can be determined that the word sequence contains 3 candidate words, and the number of candidates is equal to the preset word number threshold. Therefore, the word sequence can be determined. The conditions for continuing reasoning are not met. Next, the word sequence can be input into the larger model to determine the words. Validation results, lexical units The verification results and lexical units The verification results.
[0074] In some embodiments, the method may further include: determining an intermediate lexical sequence based on the subsequent lexical sequence and at least one verification result. During the process of determining the intermediate lexical sequence, it may be determined whether there are any lexical units to be deleted in the subsequent lexical sequence. For example, based on lexical units... Validation results, lexical units The verification results and lexical units The validation results can determine whether there are any words to be deleted in the subsequent word sequence. Words to be deleted can serve as candidate words that the large model has not identified as output words, indicating the validation results.
[0075] Regarding the verification results, lexical units To word element It is mainly determined based on the hidden state features provided by the target processing layer. The word sequence... After being input into a large model and processed by multiple processing layers, lexical units can be obtained. Global hidden state features. Lexical units The global hidden state features are input into the output head of the large model to determine the lexical units. The subsequent output words. If the word with lexical elements The subsequent output terms are consistent, which confirms that the validation result represents terms. The term is identified as an output term by the large model. It's understandable that the verification methods for candidate terms are not limited to this. After inputting the global hidden state features into the output head, the probabilities of multiple terms can be obtained. Based on these probabilities and the terms mentioned above... As a candidate lexical element The probability of each candidate lexical unit can also be used to determine the verification result. By determining the verification result of each candidate lexical unit through the embodiments of this disclosure, the accuracy of inference can be improved, model illusions reduced, and user experience effectively enhanced. Furthermore, the verification results of multiple candidate lexical units can be determined at once, effectively improving inference efficiency.
[0076] In lexical With regard to lexical units If the output lexical units are the same, the lexical units can be determined. The verification results. If the word element The verification result represents the word element. Not identified as an output term by the large model; terminology can be determined. Validation failed and was not accepted by the large model. (Word units can be...) As a term to be deleted.
[0077] In response to determining that at least one candidate lexical unit contains a lexical unit to be deleted, the lexical unit to be deleted is deleted from the subsequent lexical unit sequence to obtain an intermediate lexical unit sequence, including: if at least one lexical unit exists after the lexical unit to be deleted, the lexical unit to be deleted and at least one lexical unit following the lexical unit to be deleted can be deleted from the subsequent lexical unit sequence to obtain the intermediate lexical unit sequence. For example, the above-mentioned lexical unit sequence There are words to be deleted in the text. And word elements There are still word elements afterwards. In this case, the lexical unit can be deleted. and word elements , thus obtaining the word sequence This serves as an intermediate word sequence. By removing words that failed verification through this embodiment, the large model can continue reasoning on the correct sequence, further improving the accuracy of the model's reasoning and obtaining response results that better match the input text.
[0078] Next, based on the intermediate lexical sequence, the response to the input text can be determined. Determining this response may include: in response to the determination that there is no ending lexical in the intermediate lexical sequence, obtaining new lexicals determined by the large model based on the intermediate lexical sequence. For example, as mentioned above, lexical... Not the end of a lexical sequence. After being input into a large model, and processed through multiple layers, lexical units can be obtained. Global hidden state features. (This refers to the use of lexical units.) Global latent state features, when input into the output head of a large model, can yield word units. As a new lexical unit, it can be determined whether the new lexical unit is a terminator.
[0079] Determining the response result may further include: in response to determining that the new lexical is not the end lexical, determining the baseline lexical sequence for the subsequent inference cycle based on the intermediate lexical sequence and the new lexical; taking the new lexical as the first lexical to be processed in the subsequent inference cycle, returning to the operation of determining the unprocessed representation features of the lexical to be processed, in order to perform inference in the subsequent inference cycle. For example, using lexical... Taking the non-ending lexical as an example, the lexical sequence with lexical elements By concatenating the components, the baseline word sequence for the second inference cycle is obtained. And will the word elements As the first word to be processed in the second inference cycle, operations S110 to S140 are repeated until the condition for continuing inference is no longer met, at which point the verification process is restarted. Assume that in the second inference cycle, no candidate word is a terminal word. Therefore, the sequence of words that do not meet the condition for continuing inference in the second inference cycle can be... It's understandable that determining the lexical units is necessary. , word elements and word elements The method of determining lexical units , word elements and word elements The methods used are the same or similar, and will not be elaborated further in this disclosure.
[0080] It is understood that the above example of the existence of a word to be deleted in the intermediate word sequence is used to illustrate this disclosure, but this disclosure is not limited to this, as will be explained below.
[0081] In some embodiments, determining the response result for the input text based on the intermediate lexical sequence may further include: in response to determining that there is no lexical to be deleted among at least one candidate lexical, determining the subsequent lexical sequence as the intermediate lexical sequence. For example, if the lexical is determined... , word elements and word elements The validation results for each term indicate that they were identified as output terms by the large model, and the term sequence can be... , as an intermediate word sequence.
[0082] In some embodiments, determining the response result for the input text based on the intermediate lexical sequence may further include: in response to determining that the new lexical is a terminator, fusing the intermediate lexical sequence with the new lexical to obtain a fused lexical sequence; and determining the response result for the input text based on the fused lexical sequence. For example, a large model based on the lexical sequence can be obtained. A newly identified lexical unit. A new lexical unit is a lexical unit. And word elements Taking the ending morpheme as an example, the morpheme can be... By concatenating this lexical sequence, the resulting fused lexical sequence can be... .
[0083] In this fused lexical sequence, t can be 4, and the lexical units... It can mean "today", word element It can mean "weather", word element It can mean "true", a word element It can mean "good". That is, the input text is "The weather is really nice today". Furthermore, in the fused lexical sequence, the lexical... It can mean "ah", a word element. It can represent ",", word element It can mean "we", a word element. It can mean "together", word element It can mean "to go out", a word element. It can mean "ba". (Word element) Terminology marks the end of a sentence and do not represent actual characters. Therefore, for the input text "The weather is so nice today," the response could be "The weather is so nice today, let's go out together." It's important to understand that the number of terminology marks shown above is merely an example. In different inference tasks, the number of terminology marks can be more or less.
[0084] It is understood that the above description uses the example of the absence of a terminator in the intermediate word sequence to illustrate this disclosure, but this disclosure is not limited thereto. In some embodiments, in response to determining that a terminator exists in the intermediate word sequence, a response result for the input text is determined based on the intermediate word sequence. For example, in an inference cycle, if the decoded candidate word is a terminator, and none of the candidate word decoded in that inference cycle is a terminator to be deleted, a response result can be determined based on the intermediate word sequence.
[0085] It is understood that the input data for a large model is not limited to text-based data, but may also include data from at least one of other modalities, such as images, audio, and video. Using the method provided in this disclosure, when the input data is multimodal, multi-lexical prediction can be achieved based on the causal logical relationships between the input text and other modalities. It is also understood that the lexical terms in this disclosure include text lexical terms, and may also include at least one of image lexical terms, audio lexical terms, and video lexical terms.
[0086] As can be understood, the reasoning method of this disclosure has been explained above, and the model training method of this disclosure will be explained below.
[0087] Figure 3 This is a flowchart of a model training method according to an embodiment of the present disclosure.
[0088] like Figure 3 As shown, the method 300 may include operations S310 to S360.
[0089] In operation S310, the word units of the sample to be processed are input into the embedding layer of the large model to obtain the representation features of the word units of the sample to be processed.
[0090] In this embodiment of the disclosure, the large model can receive input sample data including one or more modalities of input text. The embedding layer can embed the tokens of the sample to be processed to obtain the representation features of the sample to be processed.
[0091] In this embodiment of the disclosure, the sample lexical units to be processed are determined based on the sample response lexical units. The input sample lexical units can be lexical units in the input sample text, the sample response lexical units are lexical units in the sample response sequence, and the sample response sequence can be a lexical sequence of the sample response result. The sample response result is the label of the input sample text.
[0092] For example, the input sample text could be "The weather is really nice today". This input sample text is then converted into a sequence of words. , as the input sample sequence. It can be a word group for "today". It can be a lexical unit for "weather". It can be a lexical unit for "true". This can be a "good" lexical unit. Large models perform forward computation based on the lexical sequence of the input text, for example, to obtain lexical units. .
[0093] The tags for the input sample text "The weather is so nice today" can be used to complete the input sample text. The sample response can be "The weather is so nice today, let's go out together." The word sequence of the sample response can be... Lexical units It can mean "ah", a word element. It can represent ",", word element It can mean "we", a word element. It can mean "together", word element It can mean "to go out", a word element. It can mean "ba". (Word element) This is the terminator, not an actual character. Terminators can be... As sample words to be processed, they are input into the embedding layer to obtain words. The representation features are used as the representation features of the sample to be processed. It can be understood that if the word sequence of the sample response result contains duplicate parts of the word sequence of the input sample text, deduplication can be performed, and the deduplicated sequence can be used as the sample response sequence. Alternatively, it can be understood that if the word sequence of the sample response result does not contain duplicate parts of the word sequence of the input sample, the word sequence of the sample response result can be used as the sample response sequence.
[0094] In operation S320, the features of the preceding hidden state and the representation features of the sample to be processed are fused to obtain the fused features.
[0095] In this embodiment, the preceding latent state feature refers to the latent state feature of the preceding lexical unit. This latent state feature is determined by the target processing layer of the large model. The preceding lexical unit is the lexical unit preceding the lexical unit of the sample to be processed. If the sample response lexical unit is the first lexical unit in the sample response sequence, its preceding lexical unit can be the last lexical unit in the input sample sequence. The input sample sequence can be a sequence of lexical units in the input text. If the sample response lexical unit is not the first lexical unit in the sample response sequence, its preceding lexical unit can be the lexical unit preceding that sample response lexical unit in the sample response sequence.
[0096] For example, a large model may include a processing network and an output network. The processing network may include multiple processing layers. The output network can also be called the output head. Processing layers can be Transformer layers. Any of the multiple processing layers can serve as the target processing layer. As mentioned above, in a large model, words are decoded... Then, the current word sequence can be In the pre-hidden state features, the features can be determined by the target processing layer for specific lexical units. The hidden state features. This will target lexical units. Hidden state features and lexical units By fusing the representation features, a fused feature can be obtained. The fusion method can be concatenation or addition, etc. It can be understood that after training is complete, the word sequence... As an input sequence, a large model can, for example, decode words. .
[0097] In operation S330, the fused features are input into the target processing layer to obtain the hidden state features to be processed.
[0098] For example, when targeting word elements Hidden state features and lexical units After the representation features are fused, the fused features can be input into the target processing layer to obtain the hidden state features to be processed.
[0099] In operation S340, the candidate lexical units of the sample lexical unit to be processed are determined based on the features of the hidden state to be processed.
[0100] In this embodiment of the disclosure, the latent state features to be processed are input into the output network of a large model to obtain subsequent candidate words. For example, when targeting words... Hidden state features and lexical units After feature fusion, the fused features are input into the target processing layer to obtain the latent state features to be processed. These latent state features are then input into the output head of the large model to obtain the lexical units. , as a candidate word element.
[0101] In operation S350, loss information for the subsequent candidate word is determined based on the subsequent candidate word and the sample response word for the subsequent candidate word in the sample response sequence.
[0102] For example, according to lexical units and word elements It can be determined that it targets word elements. The loss information. It's understandable that loss information can be determined based on various loss functions. These various loss functions include, for example, the cross-entropy loss function.
[0103] In the S360 operation, a large model is trained based on the loss information for subsequent candidate words.
[0104] For example, based on the target word The loss information can be used to adjust the parameters of a large model in order to train a large model.
[0105] In this embodiment, candidate lexical generation is primarily based on the hidden state features output from the target processing layer of the large model. A Transformer layer is reused within the unified parameter space of the large model for recursive multi-lexical generation. This fully utilizes the causal logic and semantic relationships between characters in the input text, allowing the multi-lexical prediction process and the large model's autoregressive generation process to share the same hidden state evolution path. Therefore, all future lexical prediction losses can directly affect the large model parameters and its last Transformer layer, enabling future supervision signals to directly constrain the large model's representation learning process. This allows for more efficient end-to-end optimization of the model using future information, thereby enhancing the model's ability to model long-range dependencies and improving the efficiency of future information utilization.
[0106] As can be understood, the model training method of this disclosure has been explained above, and the model training method of this disclosure will be further explained below.
[0107] The large model can include an embedding layer, L decoding layers, and an output header. The L decoding layers include the first L-1 decoding layers and the Lth decoding layer. The large model may also include a fusion layer.
[0108] In some embodiments, the target processing layer is the last processing layer among multiple processing layers of a large model. For example, the processing layer of the large model can be a decoding layer. The Lth decoding layer can serve as the target processing layer.
[0109] The word sequence of the input sample text can be fed into the large model. The word sequence of the input sample text can be... The labels of the input sample text can be used as the sample response results for that input sample text. The word sequence of the sample response results can be... Large models can encode the word sequence of the input sample text to obtain words. Hidden state features For example, lexical units can be determined using Formula 1 above. Hidden state features .
[0110] Lexicons Hidden state features Input to the output head to obtain the decoded tokens. For example, lexical units can be determined using the following formula. :
[0111] (Formula 5)
[0112] Lexicon Hidden state features This can be used as a feature in the pre-hidden state. Lexical units can be used as the to-be-processed sample token. The token sequence of the input text can be spliced with the token to obtain a token sequence , which is used as the current sample token sequence.
[0113] Next, the foregoing operations S310 to S330 can be performed, and the token is used as the to-be-processed sample token and input into the embedding layer to obtain the representation feature of the token , which is used as the to-be-processed representation feature. The representation feature of the token and the hidden state feature of the token are respectively input into the fusion layer, so as to fuse the representation feature of the token and the hidden state feature of the token to obtain a fusion feature. The fusion feature can be input into the L-th decoding layer to obtain a hidden state feature , which is used as the to-be-processed hidden state feature. For example, the hidden state feature can be obtained by the following formula :
[0114] (Formula 6)
[0115] Next, operation S340 can be performed, and the hidden state feature is input into the output head to obtain the subsequent token of the token , which is used as the subsequent candidate token. For example, the token can be determined by the following formula 4 :
[0116] (Formula 7)
[0117] Next, different from the inference stage, in some embodiments, the foregoing method 300 may further include: using the sample response token for the subsequent candidate token as the to-be-processed sample token, and returning to the operation of inputting the to-be-processed sample token into the embedding layer of the large model until each token in the sample response result is used as the to-be-processed token. For example, after obtaining the token , the token in the token sequence of the sample response result can be used as the to-be-processed sample token, returned to the foregoing operation S310, and operations S310 to S340 are repeatedly performed until operations S310 to S340 are performed based on the last token in the sample response sequence. For another example, taking determining the token after the token as an example, the method for determining the token is the same as the method for determining the token The methods used are the same or similar, and will not be elaborated further in this disclosure.
[0118] After decoding the end word, the resulting M candidate words can include: M can be an integer greater than or equal to 3.
[0119] In some embodiments, determining loss information for a subsequent candidate word based on the subsequent candidate word and the sample response word in the sample response sequence for the subsequent candidate word includes: determining multiple loss information based on multiple subsequent candidate words and multiple sample response words for the multiple subsequent candidate words respectively.
[0120] In some embodiments, training a large model based on loss information for subsequent candidate words includes: determining fusion loss information based on multiple loss information; and training a large model based on the fusion loss information.
[0121] For example, the word sequence of the sample response result can be Based on the word sequence of the sample response and the aforementioned M candidate words, multiple loss parameters can be determined. Based on these multiple loss parameters, the fusion loss parameter can be determined. For example, the fusion loss parameter (loss) can be determined using the following formula:
[0122] (Formula 8)
[0123] It can be the cross-entropy loss function. There can be M candidate lexical units The k-th candidate word element, It can be the k-th sample response word in the word sequence of the sample response results. k can be an integer greater than or equal to 1 and less than or equal to M.
[0124] It is understood that the method of this disclosure has been described above, and the apparatus of this disclosure will be described below.
[0125] Figure 4 This is a block diagram of a reasoning apparatus according to an embodiment of the present disclosure.
[0126] like Figure 4 As shown, the device 400 may include a first determining module 410, a first fusion module 420, a first processing module 430, and a second determining module 440.
[0127] The first determining module 410 is used to determine the representation features to be processed for the lexical units to be processed. The lexical units to be processed are those determined by the large model based on the preceding lexical sequence. The preceding lexical sequence includes the lexical sequence of the input text.
[0128] The first fusion module 420 is used to fuse the pre-hidden state features and the representation features to be processed to obtain fused features. The pre-hidden state features are the hidden state features of the preceding words in the preceding word sequence, and the hidden state features of the preceding words are determined by the target processing layer of the large model.
[0129] The first processing module 430 is used to input the fused features into the target processing layer to obtain the hidden state features to be processed.
[0130] The second determining module 440 is used to determine the subsequent candidate lexical units of the lexical unit to be processed based on the hidden state features of the lexical unit to be processed.
[0131] In some embodiments, the target processing layer is the last processing layer among multiple processing layers of a large model.
[0132] In some embodiments, the preceding lexical sequence includes the baseline lexical sequence of the current inference cycle, which includes the lexical sequence of the input text.
[0133] In some embodiments, the apparatus 400 further includes: a second fusion module, configured to fuse the current lexical sequence and subsequent candidate lexical units to obtain a subsequent lexical sequence. The current lexical sequence includes the preceding lexical sequence and the lexical unit to be processed. A first return module, configured to, in response to determining that the subsequent lexical sequence satisfies the conditions for continuing inference, return to the operation of determining the unprocessed representation features of the unprocessed lexical unit as the subsequent candidate lexical unit to be processed, so as to continue inference in the current inference cycle until the subsequent lexical sequence no longer satisfies the conditions for continuing inference.
[0134] In some embodiments, the conditions for continuing reasoning include at least one of the following: the number of candidates is less than a preset threshold for the number of lexical units, where the number of candidates is the number of candidate lexical units in the subsequent lexical unit sequence; or there is an end lexical unit in the subsequent lexical unit sequence.
[0135] In some embodiments, the apparatus 400 further includes: a first obtaining module, configured to, in response to determining that the subsequent word sequence does not meet the conditions for continuing inference, input the subsequent word sequence into a large model to obtain at least one verification result, the verification result indicating whether the candidate word is determined as an output word by the large model; a third determining module, configured to determine an intermediate word sequence based on the subsequent word sequence and at least one verification result; and a fourth determining module, configured to determine a response result for the input text based on the intermediate word sequence.
[0136] In some embodiments, the third determining module includes one of the following: a deletion submodule, configured to delete the word to be deleted from the subsequent word sequence in response to determining that at least one candidate word contains a word to be deleted, thereby obtaining an intermediate word sequence. The word to be deleted is a candidate word that, according to the verification results, was not determined as an output word by the large model. A first determining submodule, configured to determine the subsequent word sequence as an intermediate word sequence in response to determining that at least one candidate word does not contain a word to be deleted.
[0137] In some embodiments, the deletion submodule includes a deletion unit, configured to delete the word to be deleted and at least one word following the word to be deleted from the subsequent word sequence, in the case where at least one word exists after the word to be deleted, to obtain an intermediate word sequence.
[0138] In some embodiments, the fourth determining module includes: an acquisition submodule, configured to acquire a new lexical unit determined by the large model based on the intermediate lexical unit sequence in response to determining that there is no end lexical unit in the intermediate lexical unit sequence; a second determining submodule, configured to determine a baseline lexical unit sequence for the subsequent inference cycle based on the intermediate lexical unit sequence in response to determining that the new lexical unit is not an end lexical unit; and a return submodule, configured to return to the operation of determining the unprocessed representation features of the unprocessed lexical unit as the first unprocessed lexical unit in the subsequent inference cycle for inference in the subsequent inference cycle.
[0139] In some embodiments, the fourth determining module further includes: a fusion submodule, configured to, in response to determining that the new lexical is an end lexical, fuse the intermediate lexical sequence with the new lexical to obtain a fused lexical sequence; and a third determining submodule, configured to, based on the fused lexical sequence, determine the response result for the input text.
[0140] In some embodiments, the fourth determining module further includes a fourth determining submodule, configured to determine a response result for the input text based on the intermediate word sequence in response to determining that an end word exists in the intermediate word sequence.
[0141] In some embodiments, the first fusion module 420 includes: a first normalization submodule, used to normalize the previous hidden state features to obtain previous normalized features; a second normalization submodule, used to normalize the representation features to be processed to obtain the normalized features to be processed; and a splicing submodule, used to splice the previous normalized features and the normalized features to be processed to obtain fused features.
[0142] Figure 5 This is a block diagram of a model training apparatus according to an embodiment of the present disclosure.
[0143] like Figure 5As shown, the device 500 may include an acquisition module 510, a second fusion module 520, a second processing module 530, a fifth determination module 540, a sixth determination module 550, and a training module 560.
[0144] The module 510 is used to input the unprocessed sample lexical units into the embedding layer of the large model to obtain the unprocessed sample representation features of the unprocessed sample lexical units. The unprocessed sample lexical units are determined based on the sample response lexical units, which are the lexical units in the sample response sequence. The sample response sequence is the lexical sequence of the sample response result, and the sample response result is the label of the input sample text.
[0145] The second fusion module 520 is used to fuse the pre-hidden state features and the representation features to be processed to obtain fused features. The pre-hidden state features are the hidden state features of the preceding lexical units. The hidden state features of the preceding lexical units are determined by the target processing layer of the large model. The preceding lexical units are the lexical units before the lexical units of the sample to be processed.
[0146] The second processing module 530 is used to input the fused features into the target processing layer to obtain the hidden state features to be processed.
[0147] The fifth determining module 540 is used to determine the subsequent candidate words of the sample words to be processed based on the hidden state features to be processed.
[0148] The sixth determining module 550 is used to determine the loss information for the subsequent candidate word based on the subsequent candidate word and the sample response word in the sample response sequence for the subsequent candidate word.
[0149] Training module 560 is used to train a large model based on loss information for subsequent candidate words.
[0150] In some embodiments, the target processing layer is the last processing layer among multiple processing layers of a large model.
[0151] In some embodiments, the apparatus 500 further includes a second return module, configured to return to the operation of inputting the sample response lexicon for the subsequent candidate lexicon as a new sample lexicon to be processed, until any lexicon in the sample response result is used as a sample lexicon to be processed.
[0152] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0153] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0154] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0155] like Figure 6 As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 602 or a computer program loaded from storage unit 608 into random access memory (RAM) 603. RAM 603 may also store various programs and data required for the operation of device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0156] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of monitors, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0157] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as at least one of the inference methods and model training methods. For example, in some embodiments, at least one of the inference methods and model training methods may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of at least one of the inference methods and model training methods described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured by any other suitable means (e.g., by means of firmware) to perform at least one of the inference method and the model training method.
[0158] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), systems-on-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0159] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0160] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (EPROM) or flash memory, optical fiber, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0161] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a cathode ray tube (CRT) monitor or a liquid crystal display (LCD)); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0162] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0163] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
[0164] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0165] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A reasoning method, comprising: Determine the representation features of the lexical units to be processed, wherein the lexical units to be processed are lexical units determined by the large model based on the preceding lexical sequence, and the preceding lexical sequence includes the lexical sequence of the input text; The preceding hidden state features and the representation features to be processed are fused to obtain fused features. The preceding hidden state features are the hidden state features of the preceding words in the preceding word sequence. The hidden state features of the preceding words are determined by the target processing layer of the large model. The fused features are input into the target processing layer to obtain the hidden state features to be processed; Based on the hidden state features to be processed, the subsequent candidate lexical units of the lexical unit to be processed are determined.
2. The method according to claim 1, wherein, The target processing layer is the last processing layer among the multiple processing layers of the large model.
3. The method according to claim 1, wherein, The preceding lexical sequence includes the baseline lexical sequence of the current inference cycle, and the baseline lexical sequence of the current inference cycle includes the lexical sequence of the input text. Also includes: The current word sequence and the subsequent candidate word sequence are merged to obtain the subsequent word sequence, wherein the current word sequence includes the preceding word sequence and the word to be processed. In response to determining that the subsequent word sequence satisfies the condition for continuing inference, the subsequent candidate word is treated as a word to be processed, and the process returns to the operation of determining the unprocessed representation features of the word to be processed, so as to continue inference in the current inference cycle until the subsequent word sequence no longer satisfies the condition for continuing inference. The conditions for continued reasoning include at least one of the following: The number of candidates is less than a preset threshold for the number of lexical units, where the number of candidates is the number of candidate lexical units in the subsequent lexical unit sequence. The terminology states that there is an ending terminology in the subsequent terminology sequence.
4. The method according to claim 3, further comprising: In response to determining that the subsequent word sequence does not meet the conditions for continuing reasoning, the subsequent word sequence is input into the large model to obtain at least one verification result, the verification result indicating whether the candidate word is determined as an output word by the large model; The intermediate word sequence is determined based on the subsequent word sequence and the at least one verification result; Based on the intermediate word sequence, the response result for the input text is determined.
5. The method according to claim 4, wherein, Determining the intermediate lexical sequence based on the subsequent lexical sequence and the at least one verification result includes one of the following operations: In response to determining that at least one of the candidate lexical units contains a lexical unit to be deleted, the lexical unit to be deleted is deleted from the subsequent lexical unit sequence to obtain the intermediate lexical unit sequence, wherein the lexical unit to be deleted is a candidate lexical unit that is not determined as an output lexical unit by the verification result; In response to determining that there is no word to be deleted among at least one of the candidate words, the subsequent word sequence is determined as the intermediate word sequence.
6. The method according to claim 5, wherein, The step of deleting the word to be deleted from the subsequent word sequence to obtain the intermediate word sequence includes: If at least one word exists after the word to be deleted, delete the word to be deleted and at least one word after the word to be deleted from the subsequent word sequence to obtain the intermediate word sequence.
7. The method according to claim 4, wherein, Determining the response result for the input text based on the intermediate word sequence includes: In response to determining that there is no ending word in the intermediate word sequence, new words determined by the large model based on the intermediate word sequence are obtained; In response to determining that the new lexical is not an ending lexical, a baseline lexical sequence for the subsequent inference cycle is determined based on the intermediate lexical sequence; and The new lexical is used as the first lexical to be processed in the subsequent inference cycle, and the process returns to the operation of determining the representation features of the lexical to be processed, so as to perform inference in the subsequent inference cycle.
8. The method according to claim 7, wherein, The step of determining the response result for the input text based on the intermediate word sequence further includes: In response to determining that the new lexical unit is an end lexical unit, the intermediate lexical unit sequence is merged with the new lexical unit to obtain a merged lexical unit sequence; Based on the fused word sequence, the response result for the input text is determined.
9. The method according to claim 7, wherein, The step of determining the response result for the input text based on the intermediate word sequence further includes: In response to determining that there is an end word in the intermediate word sequence, a response result for the input text is determined based on the intermediate word sequence.
10. The method according to claim 1, wherein, The process of fusing the pre-hidden state features and the representation features to be processed to obtain fused features includes: The features of the previously hidden states are normalized to obtain the previously normalized features; The features to be processed are normalized to obtain the normalized features to be processed. The previously normalized features and the normalized features to be processed are concatenated to obtain the fused features.
11. A model training method, comprising: The unprocessed sample lexical units are input into the embedding layer of the large model to obtain the unprocessed sample representation features of the unprocessed sample lexical units. The unprocessed sample lexical units are determined based on the sample response lexical units. The sample response lexical units are the lexical units in the sample response sequence. The sample response sequence is determined according to the lexical sequence of the sample response result. The sample response result is the label of the input sample text. The preceding hidden state features and the representation features of the sample to be processed are fused to obtain fused features. The preceding hidden state features are the hidden state features of the preceding word units. The hidden state features of the preceding word units are determined by the target processing layer of the large model. The preceding word units are the word units before the word units of the sample to be processed. The fused features are input into the target processing layer to obtain the hidden state features to be processed; Based on the hidden state features to be processed, the subsequent candidate words of the sample word to be processed are determined; Based on the candidate lexical unit and the sample response lexical unit in the sample response sequence for the candidate lexical unit, the loss information for the candidate lexical unit is determined; The large model is trained based on the loss information for the subsequent candidate words.
12. The method according to claim 11, wherein, The target processing layer is the last processing layer among the multiple processing layers of the large model. Also includes: The sample response lexical for the candidate lexical is taken as a new sample lexical to be processed, and the process returns to the operation of inputting the sample lexical to be processed into the embedding layer of the large model, until any lexical in the sample response result is taken as a sample lexical to be processed.
13. A reasoning device, comprising: The first determining module is used to determine the representation features to be processed of the word to be processed, wherein the word to be processed is a word determined by the large model based on the preceding word sequence, and the preceding word sequence includes the word sequence of the input text; The first fusion module is used to fuse the pre-hidden state features and the representation features to be processed to obtain fused features. The pre-hidden state features are the hidden state features of the pre-words in the pre-word sequence. The hidden state features of the pre-words are determined by the target processing layer of the large model. The first processing module is used to input the fused features into the target processing layer to obtain the hidden state features to be processed; The second determining module is used to determine the subsequent candidate lexical units of the lexical unit to be processed based on the hidden state features to be processed.
14. A model training device, comprising: The module is used to input the unprocessed sample lexical units into the embedding layer of the large model to obtain the unprocessed sample representation features of the unprocessed sample lexical units. The unprocessed sample lexical units are determined based on the sample response lexical units. The sample response lexical units are the lexical units in the sample response sequence. The sample response sequence is determined according to the lexical sequence of the sample response result. The sample response result is the label of the input sample text. The second fusion module is used to fuse the preceding hidden state features and the representation features to be processed to obtain fused features. The preceding hidden state features are the hidden state features of the preceding word units. The hidden state features of the preceding word units are determined by the target processing layer of the large model. The preceding word units are the word units before the word units of the sample to be processed. The second processing module is used to input the fused features into the target processing layer to obtain the hidden state features to be processed; The fifth determining module is used to determine the subsequent candidate word of the sample word to be processed based on the hidden state features to be processed; The sixth determining module is used to determine loss information for the subsequent candidate word based on the subsequent candidate word and the sample response word for the subsequent candidate word in the sample response sequence; The training module is used to train the large model based on the loss information for the subsequent candidate lexical units.
15. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 12.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 12.
17. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 12.