Training method and device for model explicit learning position information, equipment and medium
The vector features of the training number are obtained through word participle transformation and deep learning model inference, and the activation function and autoregressive loss function are used to optimize the position information prediction, which solves the difficulties of large models in position information modeling and improves the model's learning ability and generation effect.
Patent Information
- Application Number
- CN202510057056.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing Decoding Model based on Transformers encounters difficulties in modeling location information, especially the lack of processing mechanisms at the output, which makes it difficult for large models to effectively learn location features and limited performance.
Get the training numbers by performing word-partial transformation of training samples and inputting these numbers into the deep learning model for model inference to obtain vector features. Then, the prediction process is performed by the activation function to obtain the prediction probability of the absolute position, the relative position and the next training number. Finally, the autoregressive loss function is used to optimize the model to enhance the learning of position information.
By actively learning position information, the learning ability of large models for absolute and relative positions is improved, and the generation effect and prediction accuracy of model are improved.
Smart Images

Figure CN119990365A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a training method, device, equipment and medium for explicitly learning position information of a model. Background Art
[0002] Natural language processing (NLP) combines the essence of computers, artificial intelligence, and linguistics, aiming to achieve comprehensive computer processing of human language. With technological innovation, NLP research methods have shifted from traditional rule-based and statistical methods to modern methods that combine machine learning and deep learning. As a key technology of NLP, text generation has undergone significant progress from early ngrams to current large models. However, although the Transformers-based decoding model has improved the parallelism of training, it has encountered challenges in modeling position information. The existing solution adds position encoding to the embedding layer. Although position information is introduced at the input end, the output end lacks a processing mechanism, which makes it difficult for large models to effectively learn position features and limits performance. Summary of the invention
[0003] The embodiments of the present invention provide a training method, device, equipment and medium for explicitly learning position information by a model, aiming to solve the problem in the prior art that large models cannot effectively learn position information.
[0004] In the first aspect, an embodiment of the present invention provides a training method for explicitly learning position information of a model, which is applied to a large model and includes: performing word segmentation conversion on the training samples to obtain the training number corresponding to each minimum training unit; inputting the training number into a preset deep learning model for model inference to obtain the vector features of each training number; predicting the vector features through a preset activation function to obtain the absolute position prediction probability, the relative position prediction probability and the prediction probability of the next training number for each preset position; and training and optimizing the prediction results of the absolute position prediction probability, the relative position prediction probability and the prediction probability of the next training number according to the autoregressive loss function.
[0005] In the second aspect, an embodiment of the present invention also provides a training device for explicitly learning position information of a model, which is applied to a large model and includes: a word segmentation unit, which is used to perform word segmentation conversion on the training samples to obtain the training number corresponding to each minimum training unit; an input unit, which is used to input the training number into a preset deep learning model for model inference to obtain the vector features of each training number; a prediction unit, which is used to predict the vector features through a preset activation function to obtain the absolute position prediction probability, the relative position prediction probability and the prediction probability of the next training number for each preset position; a training unit, which is used to train and optimize the prediction results of the absolute position prediction probability, the relative position prediction probability and the prediction probability of the next training number according to the autoregressive loss function.
[0006] In a third aspect, an embodiment of the present invention further provides a computer device, which includes a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above method when executing the computer program.
[0007] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores a computer program, wherein the computer program includes program instructions, and the program instructions can implement the above method when executed by a processor.
[0008] The embodiment of the present invention provides a training method, device, equipment and medium for explicitly learning location information of a model. The method is applied to a large model, and includes: performing word segmentation conversion on the training sample to obtain the training number corresponding to each minimum training unit; inputting the training number into a preset deep learning model for model inference to obtain the vector feature of each training number; performing prediction processing on the vector feature through a preset activation function to obtain the absolute position prediction probability, relative position prediction probability and prediction probability of the next training number of each preset position; and training and optimizing the prediction results of the absolute position prediction probability, the relative position prediction probability and the prediction probability of the next training number according to the autoregressive loss function. The embodiment of the present invention performs word segmentation conversion on the training samples to obtain the corresponding training numbers so that the model can understand and process text data more accurately, and inputs the training numbers into the preset deep learning model for model reasoning to obtain the corresponding vector features, and predicts multiple position probabilities based on the vector features, so as to add predictions of relative positions and absolute positions, enhance the learning of relative positions and absolute positions by the large model, and further improve the ability of the large model to learn relative position information through predictions between relative positions, and finally optimize the results of the predicted probabilities through training through an autoregressive loss function, so as to improve the prediction accuracy and generalization ability of the model while accelerating the training efficiency and realizing multi-task learning. By predicting the absolute position, relative position and the next training number, the model is actively promoted to learn position information, thereby continuously improving the ability of the large model to learn position information, so that it can learn position information efficiently. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying any creative work.
[0010] Figure 1 A schematic diagram of a flow chart of a training method for explicitly learning position information of a model provided in an embodiment of the present invention;
[0011] Figure 2 A schematic diagram of a sub-process of a training method for explicitly learning position information of a model provided in an embodiment of the present invention;
[0012] Figure 3 A schematic diagram of a sub-process of a training method for explicitly learning position information of a model provided in an embodiment of the present invention;
[0013] Figure 4 A schematic diagram of a sub-process of a training method for explicitly learning position information of a model provided in an embodiment of the present invention;
[0014] Figure 5 A schematic diagram of a sub-process of a training method for explicitly learning position information of a model provided in an embodiment of the present invention;
[0015] Figure 6 A schematic diagram of a sub-process of a training method for explicitly learning position information of a model provided in an embodiment of the present invention;
[0016] Figure 7 A schematic block diagram of a training device for explicitly learning position information of a model provided in an embodiment of the present invention;
[0017] Figure 8 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0018] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0019] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0020] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0021] It should be further understood that the term "and / or" used in the present description and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0022] See also Figure 1 , Figure 1A flow chart of a training method for explicitly learning location information of a model provided in an embodiment of the present invention. The training method for explicitly learning location information of a model in this embodiment can be applied to the learning and training of a large model. Specifically, this method is applied in a large model to obtain vector features by performing vector processing through a deep learning model in the large model, thereby helping the large model to actively perform location learning based on the vector features to obtain location prediction probability, so as to enhance the learning of absolute and relative positions by the large model, thereby improving the model generation effect.
[0023] Figure 1 1 is a flow chart of a training method for explicitly learning location information of a model provided by an embodiment of the present invention. As shown in the figure, the method includes the following steps S110-S140.
[0024] S110, performing word segmentation conversion on the training samples to obtain the training number corresponding to each minimum training unit.
[0025] In this embodiment, the training sample is a sample in the data set that the model is expected to learn. The minimum training unit is an independent unit obtained after word segmentation, which constitutes the input of model training. These units can be words, characters, or phrases, etc., depending on the method and purpose of word segmentation. The training number is the result of mapping each minimum training unit to a unique identifier. The training sample is subjected to word segmentation conversion to obtain the training number corresponding to each minimum training unit. Specifically, training samples containing the target language pattern can be collected in advance, and the text data can be cleaned to remove noise (such as HTML tags, special characters, etc.). The text is divided into minimum training units by a word segmentation tool (such as Tokenizer), and each unique minimum training unit is mapped to a unique number according to the vocabulary to obtain the corresponding training number. For example, the vocabulary V = [a, b, c, d, e], and the training number of a is 0, the training number of b is 1,..., and the training number of e is 4 in order. There is a sentence s = "aecdbd", then the minimum training unit obtained by word segmentation according to the vocabulary is: [a, e, c, d, b, d], and each minimum training unit is set as a training number, and the result is: [0, 4, 2, 3, 1, 3]. Among them, the training number can be represented by token in this embodiment. By converting the training sample through word segmentation to obtain the training number, the training sample is converted into a format that can be processed by the model, which lays the foundation for subsequent model training.
[0026] In one embodiment, if Figure 2 As shown, the step S110 also includes steps S111-S112.
[0027] S111, performing word segmentation processing on the training sample by using a preset segmentation tool to obtain a minimum training unit;
[0028] S112: Map the minimum training unit according to a preset word segmentation table to obtain a corresponding training number.
[0029] In this embodiment, the preset word segmentation tool is a software or algorithm for segmenting continuous text strings into smaller, meaningful units. In this embodiment, the preset word segmentation tool is a Tokenizer, which is a tool or algorithm for segmenting text data into smaller units (usually called tokens). These tokens can be words, subwords, characters, or other text fragments defined according to specific tasks. It can be understood that the main purpose of Tokenizer is to convert continuous character sequences in natural language into discrete units that can be understood and processed by the model, so as to facilitate subsequent analysis and processing. The preset word segmentation table is a data structure table (usually a hash table or a dictionary) for storing the mapping relationship between the minimum training unit and the training number. The training sample is segmented by the preset segmentation tool to obtain the minimum training unit. Specifically, according to the language characteristics and model requirements, an appropriate Tokenization strategy is selected. For English, space segmentation may be used; while for languages such as Chinese that do not have obvious word separations, character-based Tokenization or subword-based Tokenization may be used. Segmentation is performed according to the selected Tokenization strategy to obtain the minimum training unit. According to the selected Tokenization strategy, a preset vocabulary is constructed. Map each minimum training unit to a unique training number according to the vocabulary. Use the preset segmentation tool to segment the training samples, and map the obtained minimum training unit to the preset segmentation table to obtain the training number, which provides a basis for subsequent feature extraction and model training, and indirectly improves the performance and accuracy of the model.
[0030] S120: Input the training number into a preset deep learning model for model inference to obtain vector features of each training number.
[0031] In this embodiment, the deep learning model is an algorithm that can automatically learn data representation and patterns. In large models, deep learning models are usually used to process text data, learn the semantic and grammatical structure of text, and perform various NLP tasks, such as text classification, named entity recognition, machine translation, etc. In this embodiment, the preset deep learning model is a Transformers model, wherein the Transformers model is a deep learning model based on a self-attention mechanism. The core of the Transformers model is the encoder and decoder structure, which are composed of multiple self-attention layers and feedforward neural network layers. These layers can capture contextual information in the text and generate vector features for each token. The training number is input into the preset deep learning model for model reasoning to obtain the vector features of each training number. Specifically, the training number is input into the preset deep learning model, and the multi-layer self-attention mechanism and feedforward neural network of the decoder in the model process these vectors to generate the vector features of each token, wherein the vector feature refers to a vector representing a token. Generally, the vector length L is set to 768. By inputting the training number into the Transformers model for model inference, the vector features of each token are obtained, providing a basis for the subsequent model's explicit learning.
[0032] S130, predicting the vector features through a preset activation function to obtain an absolute position prediction probability, a relative position prediction probability, and a prediction probability of the next training number for each preset position.
[0033] In this embodiment, the preset activation function is a function added to the neural network to help the network learn complex patterns in the data. In this embodiment, the preset activation function is a Softmax function, which can convert a vector into a probability distribution, that is, each element in the vector is converted to a value between 0 and 1, and the sum of these values is 1. In NLP tasks, the softmax function is used in the output layer of the model to convert the predicted output of the model into the probability of each possible category. The preset position is a set or randomly selected sentence position. The vector features are predicted and processed by the preset activation function. Specifically, after the vector features are passed through the model reasoning and the softmax activation function, the probability value of predicting the next token at each position can be obtained: as well as the relative position prediction probability and absolute position probability of each training number at the current position. For example, if the training sample is "I am a student", the minimum training unit is "I", "I am", "student", "student", and the corresponding training numbers are "a", "c", "d", "e". If the preset position is the second position of the sentence, then the absolute position prediction probability refers to the probability that "c" is the second position, and the relative position probability refers to the probability of the difference between "c" and other positions when it is the second position. The prediction probability of the next training number refers to the probability of "c" followed by "d". The vector features of the model are predicted and processed by the preset activation function, and the absolute position prediction probability, relative position prediction probability and prediction probability of the next training number of the current training number can be obtained to actively promote the model to learn the absolute position and the relative position, thereby improving the ability of the large model to learn relative position information and improve the model generation effect.
[0034] In one embodiment, if Figure 3 As shown, the step S130 also includes steps S131-S133.
[0035] S131, numbering the training samples by position, and assigning the position number to the training number of the corresponding position;
[0036] S132, performing absolute position prediction on all the vector features through a preset position decoding matrix and the preset activation function, and determining a first probability prediction set of a number of the training numbers at the preset positions;
[0037] S133, determining a corresponding target training number according to the position number of the preset position, and determining the absolute position prediction probability in the first probability prediction set according to the target training number.
[0038] In this embodiment, the position decoding matrix is used in the Transformers model to represent the position information of each word in the sequence. Since the Transformers model itself does not have a built-in mechanism for processing the sequence order, it is necessary to retain the order information of the words in the sequence through position encoding, which is the last layer of the model and is a fully connected layer. Among them, the specific matrix is not limited, and the corresponding technical effect can be achieved. The training samples are numbered by position, and the position number is assigned to the training number of the corresponding position. Specifically, the present invention numbers the training numbers converted by the training samples, and assigns a position number to each training number in turn. For example, if the training sample is "I am a student", the minimum training unit is "I", "I am", "student", "student", and the corresponding training numbers are "a", "c", "d", and "e". Then "a", "c", "d", and "e" are numbered "1", "2", "3", and "4" respectively, and the position number is assigned to the training number of the corresponding position, that is, a corresponds to position number 1, c corresponds to position number 2, and d corresponds to position number 3. All the vector features are predicted for absolute position through the preset position decoding matrix and the preset activation function. Specifically, if the preset position number is 2, the process of performing absolute position prediction is:
[0039]
[0040] Among them, p(Pos L ) is the predicted probability of each training number at the preset position, softmax is the preset activation function, W position A decoding matrix for the preset position, is the feature vector for each training number. Where M is the decoder, is the jth sample in training sample i. By performing absolute position prediction on each vector feature, the probability of all training numbers at the current preset position can be obtained, that is, a first probability prediction set is formed. The corresponding target training number is determined according to the position number of the preset position, and the absolute position prediction probability is determined in the first probability prediction set according to the target training number. For example, if the position number of the preset position is 2, then the training number c corresponding to number 2 is the target training number, and the probability corresponding to c is used as the absolute position prediction probability. By predicting the absolute position probability, the model can learn the position information of words or characters in the text sequence and make accurate predictions, thereby improving the model's ability to understand the text and position learning ability.
[0041] In one embodiment, if Figure 4 As shown, the step S130 also includes steps S134-S136.
[0042] S134, randomly selecting two training numbers from a plurality of training numbers to form a training pair;
[0043] S135, performing relative position prediction on the vector features of the training pair through the preset position decoding matrix and the preset activation function to obtain a second probability prediction set;
[0044] S136. Determine a target relative position value according to the position number corresponding to the training pair, and determine the relative position prediction probability in the second probability prediction set according to the target relative position value.
[0045] In this embodiment, the second probability prediction set is a set consisting of all relative position probabilities between training pairs. Two training numbers are randomly selected from a number of training numbers to construct a training pair. Specifically, two training numbers (tokens) are randomly selected from a training sample, and a training pair is formed based on the two tokens. Among them, which token is used as or No limitation is made. The vector features of the training pair are predicted relative to each other through the preset position decoding matrix and the preset activation function. Specifically, at this time, the position number of the preset position decoding matrix is regarded as the quantity, that is, the position difference between two tokens is predicted. The steps of predicting the relative position of the vector features of the training pair through the preset position decoding matrix and the preset activation function are as follows:
[0046]
[0047] Wherein, the p(Pos right-left ) is the probability value predicted for relative position prediction, softmax is the preset activation function, W position A decoding matrix for the preset position, They are respectively the vector features of the two training numbers in the training pair. Among them, it should be noted that the probability of predicting multiple position differences, for example, the probability that the position difference between the training numbers a and d is 2 is 80%, the probability of 1 is 10%, and the probability of 0 is 10%. The second probability prediction set is composed of the predicted probability values, and the target relative position value is determined according to the position number corresponding to the training pair. The relative position prediction probability is determined in the second probability prediction set according to the target relative position value, that is, the position numbers of a and d are 0 and 2, and the probability of the position difference being 2 is 80% as the relative position prediction probability. By determining the relative position probability, the relative position between tokens is modeled, and the position difference between two tokens is predicted, so as to further enhance the learning of relative position information of the large model and improve the model generation effect.
[0048] In one embodiment, if Figure 5 As shown, the step S130 also includes steps S137-S138.
[0049] S137, performing position prediction on all the vector features through a preset decoding matrix and the preset activation function, and obtaining probability values of all the training numbers being located at the next position of the preset position;
[0050] S138. Determine a target training number according to the training sample, and determine the probability value of the target training number as the predicted probability of the next training number.
[0051] In this embodiment, the preset decoding matrix is the last layer of the preset deep learning model, which is a fully connected layer used to calculate the probability distribution and predict the next token. Among them, there is no limitation on the specific decoding matrix, as long as the corresponding effect can be achieved. All the vector features are predicted in position through the preset decoding matrix and the preset activation function to obtain the probability value of all the training numbers being located at the next position of the preset position. Specifically, after selecting the preset position, the vector features are predicted in position through the preset decoding matrix and the preset activation function to predict the probability value of all feature vectors being at the next position of the preset position. For example, if the position number of the current preset position is 2, then the probability value of other training numbers being at position number 3 after 2 is predicted. The specific process of position prediction is:
[0052]
[0053] in, is the probability value of the token after the current preset position (or current token). Softmax is the preset activation function, W decoder is the preset decoding matrix, After model inference The corresponding vector feature is is the current token. It can be understood that, because each position has its corresponding training number (token), the probability value of the token at the next position can be predicted by the token at that position, which can also be understood as using the token at the previous position to predict the token at the next position, that is, x 1 Predict x 2 , x 2 Predict x 3 ..., then we can get x 2 The probability distribution of can be used to calculate the loss function. It should be noted that It is in the first position, and there is no other token that can predict it, so we start from 2, that is, j is greater than or equal to 2. For example, the training number of position number 2 is "c", so the probability of all training numbers "a", "c", "d", and "e" being located after "c" is predicted, and the target training number is determined according to the training sample, that is, the training number after "c" is determined to be "d" according to the real data, and the probability corresponding to "d" is used as the predicted probability of the next training number. By predicting the predicted probability of the next training number, the model's understanding and learning ability of the position of the text sequence can be improved.
[0054] S140, training and optimizing the prediction results of the absolute position prediction probability, the relative position prediction probability and the prediction probability of the next training number according to an autoregressive loss function.
[0055] In this embodiment, the autoregressive loss function is a special cross-entropy loss function, which optimizes the model by predicting the conditional probability of each word one by one. The cross-entropy loss function is mainly used to measure the difference between the probability distribution of the model output and the true probability distribution. According to the autoregressive loss function, the prediction results of the absolute position prediction probability, the relative position prediction probability and the prediction probability of the next training number are trained and optimized. Specifically, the value of the loss function is calculated according to the prediction result and the true result. This value reflects the accuracy of the model prediction. The parameters of the model are updated according to the gradient of the loss function using gradient descent or other optimization algorithms. This process is iterated continuously until the loss function converges to a smaller value, indicating that the prediction performance of the model has been improved. Among them, the performance of the model on the validation set can be regularly evaluated during the training process to avoid overfitting. Adjust the model structure, loss function or optimization strategy according to the evaluation results. The prediction results of the model are trained and optimized according to the autoregressive loss function to improve the training effect of the model.
[0056] In one embodiment, if Figure 6 As shown, the step S140 also includes steps S141-S142.
[0057] S141, constructing the autoregressive loss functions of the absolute position prediction probability, the relative position prediction probability and the prediction probability of the next training number respectively;
[0058] S142. Add the three corresponding autoregressive loss functions, and perform training optimization on the position prediction result according to the added loss function.
[0059] In this embodiment, autoregressive loss functions are constructed for the absolute position prediction probability, the relative position prediction probability, and the prediction probability of the next training number, respectively, wherein the constructed autoregressive loss functions are:
[0060]
[0061] Among them, loss autogress is the loss function of the predicted probability of the next training number, is the predicted probability of the next training number, and L is the total number of training numbers. Among them, because it is an autoregressive loss function and the position of the previous token is used to predict the token of the next position, It is in the first position and there is no other token that can predict it, so we start counting from j=2.
[0062]
[0063] Among them, loss absolute is the loss function of the absolute position prediction probability, p(Pos L ) is the absolute position prediction probability.
[0064]
[0065] Among them, loss relative is the loss function of the relative position prediction probability, p(Pos right-left ) is the relative position prediction probability, T is the number of randomly selected training pairs, and k is the kth training pair. The autoregressive loss functions corresponding to the three are added together to obtain the added loss function: loss = loss autogress +loss absolute +loss relative . The training and optimization of the position prediction results are performed based on the added loss function. Specifically, gradient descent or other optimization algorithms are used to minimize the total loss function loss. Through iterative training, the model will gradually learn how to accurately predict the absolute position, relative position, and the next training number at the same time. In each iteration, the loss function is calculated based on the prediction results of the current model and the true label, and then the parameters of the model are updated through the back propagation algorithm to reduce the prediction error. By constructing an autoregressive loss function and adding the corresponding autoregressive loss functions of the three to help complete the training optimization, the large model can actively learn the location information and improve the model's ability to learn location information, thereby improving the quality of the data generated by the model.
[0066] Figure 7 is a schematic block diagram of a training device 200 for explicitly learning position information of a model provided by an embodiment of the present invention. Figure 7As shown, corresponding to the above model explicit learning position information training method, the present invention also provides a model explicit learning position information training device. The model explicit learning position information training device includes a unit for executing the above model explicit learning position information training method, and the device can be configured in a desktop computer, tablet computer, laptop computer, etc. Specifically, please refer to Figure 7 The training device for explicitly learning position information of the model includes a word segmentation unit 210, an input unit 220, a prediction unit 230 and a training unit 240.
[0067] The word segmentation unit 210 is used to perform word segmentation conversion on the training samples to obtain the training number corresponding to each minimum training unit.
[0068] In one embodiment, the word segmentation unit 210 includes an acquisition unit and a mapping unit.
[0069] An acquisition unit, used to perform word segmentation processing on the training sample by using a preset segmentation tool to obtain a minimum training unit;
[0070] A mapping unit is used to map the minimum training unit according to a preset word segmentation table to obtain a corresponding training number.
[0071] The input unit 220 is used to input the training number into a preset deep learning model for model inference to obtain the vector features of each training number.
[0072] The prediction unit 230 is used to perform prediction processing on the vector features through a preset activation function to obtain the absolute position prediction probability, the relative position prediction probability and the prediction probability of the next training number of each preset position.
[0073] In one embodiment, the prediction unit 230 includes a value assignment unit, an absolute position prediction unit, and an absolute probability determination unit.
[0074] An assignment unit, used for assigning position numbers to the training samples and assigning the position numbers to the training numbers at the corresponding positions;
[0075] An absolute position prediction unit, used for performing absolute position prediction on all the vector features through a preset position decoding matrix and the preset activation function, and determining a first probability prediction set of a number of the training numbers at the preset positions;
[0076] The absolute probability determination unit is used to determine the corresponding target training number according to the position number of the preset position, and determine the absolute position prediction probability in the first probability prediction set according to the target training number.
[0077] In one embodiment, the prediction unit 230 includes a construction unit, a relative position prediction unit, and a relative probability prediction unit.
[0078] A construction unit, used for randomly selecting two training numbers from a plurality of training numbers to construct a training pair;
[0079] A relative position prediction unit, used to perform relative position prediction on the vector features of the training pair through the preset position decoding matrix and the preset activation function to obtain a second probability prediction set;
[0080] A relative probability prediction unit is used to determine a target relative position value according to a position number corresponding to the training pair, and to determine the relative position prediction probability in the second probability prediction set according to the target relative position value.
[0081] In one embodiment, the prediction unit 230 includes a position prediction unit and a position probability determination unit.
[0082] A position prediction unit, used to perform position prediction on all the vector features through a preset decoding matrix and the preset activation function, and obtain probability values of all the training numbers being located at the next position of the preset position;
[0083] The position probability determination unit is used to determine a target training number according to the training sample, and determine the probability value of the target training number as the predicted probability of the next training number.
[0084] The training unit 240 is used to train and optimize the prediction results of the absolute position prediction probability, the relative position prediction probability and the prediction probability of the next training number according to the autoregressive loss function.
[0085] In one embodiment, the training unit 240 includes a function construction unit and an optimization unit.
[0086] A function construction unit, used to respectively construct an autoregressive loss function of the absolute position prediction probability, the relative position prediction probability and the prediction probability of the next training number;
[0087] The optimization unit is used to add the autoregressive loss functions corresponding to the three, and perform training optimization of the position prediction results according to the added loss function.
[0088] It should be noted that technical personnel in the relevant field can clearly understand that the specific implementation process of the training device 200 and each unit of the above-mentioned model explicitly learning position information can refer to the corresponding description in the aforementioned method embodiment, and for the convenience and conciseness of the description, it will not be repeated here.
[0089] The training device for explicitly learning position information of the above-mentioned model can be implemented in the form of a computer program, which can be used in a computer program such as Figure 8 Runs on the computer device shown.
[0090] See also Figure 8 , Figure 8 5 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 500 may be a terminal or a server, wherein the terminal may be an electronic device with communication functions such as a smart phone, a tablet computer, a laptop computer, a desktop computer, a personal digital assistant, and a wearable device. The server may be an independent server or a server cluster composed of multiple servers.
[0091] See also Figure 8 The computer device 500 includes a processor 502 , a memory and a network interface 505 connected via a system bus 501 , wherein the memory may include a non-volatile storage medium 503 and an internal memory 504 .
[0092] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions, and when the program instructions are executed, the processor 502 can execute a training method for explicitly learning position information of a model.
[0093] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500 .
[0094] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a training method for explicitly learning position information of a model.
[0095] The network interface 505 is used to communicate with other devices over the network. Figure 8 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device 500 to which the solution of the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0096] The processor 502 is used to run a computer program 5032 stored in the memory to implement the steps of the above method.
[0097] It should be understood that in the embodiment of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0098] It can be understood by those skilled in the art that all or part of the processes in the method for implementing the above embodiment can be completed by instructing the relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiment of the above method.
[0099] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor executes the steps of the above method.
[0100] The storage medium may be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk, etc., which are computer-readable storage media that can store program codes.
[0101] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0102] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0103] The steps in the method of the embodiment of the present invention can be adjusted in order, combined and deleted according to actual needs. The units in the device of the embodiment of the present invention can be combined, divided and deleted according to actual needs. In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0104] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, terminal, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention.
[0105] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A training method for explicitly learning position information of a model, characterized in that: The method is applied to a large model, and the method comprises: Perform word segmentation conversion on the training samples to obtain the training number corresponding to each minimum training unit; Input the training number into a preset deep learning model for model inference to obtain vector features of each training number; The vector features are predicted by a preset activation function to obtain the absolute position prediction probability, the relative position prediction probability and the prediction probability of the next training number of each preset position; The prediction results of the absolute position prediction probability, the relative position prediction probability and the prediction probability of the next training number are trained and optimized according to the autoregressive loss function.
2. The method according to claim 1, characterized in that The step of performing word segmentation conversion on the training samples to obtain the training number corresponding to each minimum training unit includes: The training samples are segmented using a preset segmentation tool to obtain a minimum training unit; The minimum training unit is mapped according to a preset word segmentation table to obtain a corresponding training number.
3. The method according to claim 1, characterized in that The step of performing prediction processing on the vector features through a preset activation function to obtain the absolute position prediction probability of each preset position includes: Numbering the training samples by position, and assigning the position number to the training number of the corresponding position; Performing absolute position prediction on all the vector features through a preset position decoding matrix and the preset activation function to determine a first probability prediction set of a number of the training numbers at the preset positions; The corresponding target training number is determined according to the position number of the preset position, and the absolute position prediction probability is determined in the first probability prediction set according to the target training number.
4. The method according to claim 3, characterized in that The step of performing prediction processing on the vector feature through a preset activation function to obtain the relative position prediction probability of each preset position includes: Randomly selecting two training numbers from a number of the training numbers to construct a training pair; Perform relative position prediction on the vector features of the training pair through the preset position decoding matrix and the preset activation function to obtain a second probability prediction set; A target relative position value is determined according to the position number corresponding to the training pair, and the relative position prediction probability is determined in the second probability prediction set according to the target relative position value.
5. The method according to claim 4, characterized in that The step of performing prediction processing on the vector feature by using a preset activation function to obtain the prediction probability of the next training number also includes: Performing position prediction on all the vector features through a preset decoding matrix and the preset activation function to obtain probability values of all the training numbers being located at the next position of the preset position; A target training number is determined according to the training sample, and a probability value of the target training number is determined as a predicted probability of the next training number.
6. The method according to claim 1, characterized in that The step of training and optimizing the prediction results of the absolute position prediction probability, the relative position prediction probability and the prediction probability of the next training number according to the autoregressive loss function comprises: Constructing autoregressive loss functions of the absolute position prediction probability, the relative position prediction probability and the prediction probability of the next training number respectively; The corresponding autoregressive loss functions of the three are added together, and the training and optimization of the position prediction results are performed based on the added loss function.
7. The method according to claim 4, characterized in that The step of performing relative position prediction on the vector features of the training pair through the preset position decoding matrix and the preset activation function is: Among them, p(Pos right-left ) is the probability value predicted for relative position prediction, softmax is the preset activation function, W position A decoding matrix for the preset position, They are respectively the vector features of the two training numbers in the training pair.
8. A training device for explicitly learning position information by a model, characterized in that: The device is applied to a large model and comprises: The word segmentation unit is used to perform word segmentation conversion on the training samples to obtain the training number corresponding to each minimum training unit; An input unit, used to input the training number into a preset deep learning model for model inference to obtain vector features of each training number; A prediction unit, used to perform prediction processing on the vector features through a preset activation function to obtain an absolute position prediction probability, a relative position prediction probability and a prediction probability of the next training number for each preset position; A training unit is used to train and optimize the prediction results of the absolute position prediction probability, the relative position prediction probability and the prediction probability of the next training number according to an autoregressive loss function.
9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the method according to any one of claims 1 to 7 can be implemented.