Large language model training method and device based on trainable residual connection and dual-scale convolution Transformer, computer equipment and readable storage medium

By introducing trainable residual connection and dual-scale convolution modules into the Transformer model, the problem of the difficulty of information transmission and insufficient local information capture of the Transformer model when processing long sequence data is solved, and more powerful expression and generalization capabilities are achieved.

CN119940416AActive Publication Date: 2025-05-06HUNAN KUNLUNYUAN ARTIFICIAL INTELLIGENCE APPLICATION SOFTWARE CO LTD

Patent Information

Application Number
CN202411849409.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-05-06
Estimated Expiration
2044-12-16

AI Technical Summary

Technical Problem

The Transformer model has difficulty in transmitting information when processing long sequence data, and focuses on global information capture and ignores local information capture, resulting in insufficient processing of complex language structures and multi-scale semantic features.

Method used

Using the Transformer architecture based on trainable residual connections and dual-scale convolution, the model's ability to capture local features is greatly enhanced, and semantic features of different scales are captured through the dual-scale convolution module.

Benefits of technology

It significantly improves the model's expression ability and generalization ability, can more effectively handle complex language structures and understand multi-scale semantic relationships, and improves performance in multiple natural language processing tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940416A_ABST
    Figure CN119940416A_ABST
Patent Text Reader

Abstract

The invention discloses a large language model training method and device based on trainable residual connection and dual-scale convolution Transform, computer equipment and a readable storage medium, and the method comprises the steps: firstly obtaining a basic model based on a multi-layer Transform architecture, each layer of the basic model comprises a self-attention and feed-forward network and is embedded into a dual-scale convolution module, and outputting the fused output as the output of the layer, and a trainable weight matrix is configured between input and output of each layer to adjust the residual connection strength. The method comprises the following steps: obtaining a preprocessed sample document to construct a training set, training a basic model to a preset condition based on the training set, and obtaining a large language model fusing trainable residual connection and dual-scale convolution, so that the model performance and generalization ability can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a training method, device, computer equipment and readable storage medium for a large language model based on trainable residual connection and dual-scale convolutional Transformer. Background Art

[0002] With the development of artificial intelligence, language models have evolved from basic models to neural network models. Large language models such as ChatGPT have demonstrated outstanding capabilities with huge data sets, achieved success in many fields, and reshaped human-computer interaction. However, despite its success, the Transformer model still has challenges. It focuses on capturing global information and ignores capturing local information. It relies on residual connections, but as the number of network layers increases, the difficulty of information transmission increases, and the bottleneck problem is prominent when processing long sequence data. Summary of the invention

[0003] The object of the present invention is to provide a training method, device, computer equipment and readable storage medium for a large language model based on trainable residual connections and dual-scale convolutional Transformer.

[0004] In a first aspect, an embodiment of the present invention provides a training method for a large language model based on a trainable residual connection and a dual-scale convolutional Transformer, comprising:

[0005] Obtain a basic model based on a multi-layer Transformer architecture, wherein the basic model includes a plurality of cascaded Transformer decoder layers, each of the Transformer decoder layers includes a self-attention mechanism and a feedforward neural network, each of the Transformer decoder layers is embedded with a dual-scale convolution module, the dual-scale convolution modules of different scales are used to capture different local features, the outputs of the dual-scale convolution modules are fused as the outputs of the corresponding Transformer decoder layers, a trainable weight matrix is ​​configured between the input and output of each of the Transformer decoder layers, and the trainable weight matrix is ​​used to adjust the strength of the residual connection between each of the Transformer decoder layers;

[0006] Obtain preprocessed sample document data to build a training set;

[0007] The basic model based on the multi-layer Transformer architecture is trained based on the training set until a preset training termination condition is reached, thereby obtaining a trained large language model that integrates trainable residual connections and dual-scale convolutional Transformers.

[0008] In a possible implementation, the training of the basic model based on the multi-layer Transformer architecture based on the training set includes:

[0009] Perform unsupervised learning based on the training set to complete the pre-training of the basic model based on the multi-layer Transformer architecture to learn the statistical characteristics and contextual relationships of the language;

[0010] Perform supervised fine-tuning on the pre-trained base model based on the multi-layer Transformer architecture according to the preset downstream tasks;

[0011] The basic model based on the multi-layer Transformer architecture after supervised fine-tuning is combined with the reward model for reinforcement learning through a preset strategy optimization algorithm.

[0012] In a possible implementation, performing unsupervised learning according to the training set includes:

[0013] The text data included in the training set is converted into a plurality of digital sequences through a word segmenter; the numbers included in the digital sequences are index numbers of the text data in the dictionary;

[0014] Obtain multiple high-dimensional vectors of high dimension by embedding the multiple digital sequences;

[0015] Performing reasoning operations according to the multiple high-dimensional directions to obtain index sequences corresponding to the multiple digital sequences;

[0016] The word segmenter is used in combination with the index sequence to perform text restoration to obtain the target natural language.

[0017] In a possible implementation, performing inference operations according to the multiple high-dimensional vectors to obtain index sequences corresponding to the multiple digital sequences includes:

[0018] Normalizing the high-dimensional vector corresponding to each layer of the Transformer architecture, and processing the normalized original features through a linear layer;

[0019] The processed original features are divided into blocks on the time axis to obtain the first features; the block operation adopts an overlap factor of 50%;

[0020] A multi-scale Transformer using a trainable residual connection structure and adding a convolution module uses an intra-frame Transformer and an inter-frame Transformer to respectively learn the short-term and long-term dependencies of the first feature to obtain a second feature, and the intra-frame Transformer and the inter-frame Transformer have the same structure;

[0021] The second feature generated by the multi-scale Transformer is subjected to a PReLU activation function and a linear layer to obtain the third feature;

[0022] Performing an overlap-add operation on the third feature to obtain a fourth feature; the overlap rate of the overlap-add operation is 0.5;

[0023] Processing the fourth feature through a ReLU activation function and a feed-forward network layer, using a nonlinear transformation to extract and learn the feature representation of each token vocabulary;

[0024] When reaching the last layer of Transformer architecture, the SoftMax function is used to obtain the sampling probability of all digital sequences, and the next digital sequence is sampled according to the preset sampling algorithm;

[0025] Repeat the steps of normalizing the high-dimensional vector corresponding to each layer of the Transformer architecture, and processing the normalized original features through a linear layer, to the step of processing the fourth feature through a ReLU activation function and a feedforward network layer, and using a nonlinear transformation to extract and learn the feature representation of each token vocabulary, until a preset number of cycles is reached, and then generating the index sequence.

[0026] In a possible implementation, the multi-scale Transformer using a trainable residual connection structure and adding a convolution module uses an intra-frame Transformer and an inter-frame Transformer to respectively learn the short-term and long-term dependencies of the first feature to obtain the second feature, including:

[0027] Assigning a unique representation to each position of the first feature using absolute position coding to obtain an intermediate first feature;

[0028] Performing a linear transformation of weights and biases on the intermediate first feature through a feedforward neural network, and introducing a nonlinear transformation through a ReLU function to obtain an intermediate second feature;

[0029] Processing the intermediate second feature through a multi-head attention mechanism to obtain an intermediate third feature;

[0030] Processing the intermediate third feature through a dual-scale convolution module to obtain an intermediate fourth feature and an intermediate fifth feature;

[0031] Processing the intermediate fourth feature and the intermediate fifth feature through the Mish activation function to obtain an intermediate sixth feature and an intermediate seventh feature;

[0032] The intermediate sixth feature and the intermediate seventh feature are spliced ​​to obtain an intermediate eighth feature;

[0033] The intermediate eighth feature is subjected to adaptive average pooling to obtain an intermediate ninth feature, and the intermediate ninth feature is processed by the feedforward neural network to obtain the second feature.

[0034] In a possible implementation, the processing the intermediate ninth feature through the feedforward neural network to obtain the second feature includes:

[0035] The intermediate ninth feature in the intra-frame Transformer and the inter-frame Transformer is added to the input of each feedforward neural network layer through a trainable residual connection and then subjected to a layer normalization operation to obtain the second feature.

[0036] In a possible implementation, the trainable residual connection is represented by the formula: y=F(W*x);

[0037] Among them, x represents input, W represents the trainable weight matrix; F(x) represents the nonlinear mapping function from input to output, and y represents the output.

[0038] In a second aspect, an embodiment of the present invention provides a training device for a large language model based on a trainable residual connection and a dual-scale convolutional Transformer, comprising:

[0039] An acquisition module is used to acquire a basic model based on a multi-layer Transformer architecture, wherein the basic model includes multiple cascaded Transformer decoder layers, each of the Transformer decoder layers includes a self-attention mechanism and a feedforward neural network, each of the Transformer decoder layers is embedded with a dual-scale convolution module, and the dual-scale convolution modules of different scales are used to capture different local features. The outputs of the dual-scale convolution modules are fused as the outputs of the corresponding Transformer decoder layers, and a trainable weight matrix is ​​configured between the input and output of each Transformer decoder layer, and the trainable weight matrix is ​​used to adjust the strength of the residual connection between each Transformer decoder layer; obtain the preprocessed sample document data to build a training set;

[0040] The training module is used to train the basic model based on the multi-layer Transformer architecture based on the training set until a preset training termination condition is reached, thereby obtaining a trained large language model that integrates trainable residual connections and dual-scale convolutional Transformers.

[0041] In a third aspect, an embodiment of the present invention provides a computer device, comprising a processor and a non-volatile memory storing computer instructions, wherein when the computer instructions are executed by the processor, the computer device executes the method described in the first aspect.

[0042] In a fourth aspect, an embodiment of the present invention provides a readable storage medium, the readable storage medium includes a computer program, and when the computer program is running, the computer device where the readable storage medium is located is controlled to execute the method described in the first aspect. Compared with the prior art, the beneficial effects provided by the present invention include: using a training method, device, computer device and readable storage medium for a large language model based on a trainable residual connection and a two-scale convolutional Transformer disclosed in the present invention, including: first obtaining a basic model based on a multi-layer Transformer architecture, each layer of which contains a self-attention and a feedforward network and is embedded with a two-scale convolution module, and the output is fused as the output of the layer, and a trainable weight matrix is ​​configured between the input and output of each layer to adjust the residual connection strength. Obtain preprocessed sample documents to construct a training set, and based on this, train the basic model to preset conditions to obtain a large language model that integrates trainable residual connections and two-scale convolutions, which can improve model performance and generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only illustrate certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained based on these drawings without creative work.

[0044] Figure 1 A schematic flow chart of the steps of a method for training a large language model based on a trainable residual connection and a dual-scale convolutional Transformer provided in an embodiment of the present invention;

[0045] Figure 2 A schematic diagram of the architecture of a large language model based on trainable residual connections and dual-scale convolutional Transformer provided in an embodiment of the present invention;

[0046] Figure 3 A schematic diagram of the internal structure of the TransformerX and multi-scale Transformer modules with dual-scale convolution provided in the embodiments of the present invention;

[0047] Figure 4 A schematic diagram of the internal structure of a dual-scale convolutional Transformer with a trainable residual connection structure provided by an embodiment of the present invention;

[0048] Figure 5A schematic diagram of the internal structure of a dual-scale convolution module and a single convolution module provided in an embodiment of the present invention;

[0049] Figure 6 A schematic block diagram of the structure of a large language model training device based on a trainable residual connection and a dual-scale convolutional Transformer provided in an embodiment of the present invention;

[0050] Figure 7 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0052] The specific implementation modes of the present invention are described in detail below in conjunction with the accompanying drawings.

[0053] In order to solve the technical problems in the aforementioned background technology, Figure 1 A flow chart of a method for training a large language model based on trainable residual connections and a dual-scale convolutional Transformer is provided in an embodiment of the present disclosure. The method for training a large language model based on trainable residual connections and a dual-scale convolutional Transformer is introduced in detail below.

[0054] Step S201, obtaining a basic model based on a multi-layer Transformer architecture, wherein the basic model includes a plurality of cascaded Transformer decoder layers, each of the Transformer decoder layers includes a self-attention mechanism and a feedforward neural network, each of the Transformer decoder layers is embedded with a dual-scale convolution module, the dual-scale convolution modules of different scales are used to capture different local features, the outputs of the dual-scale convolution modules are fused as the outputs of the corresponding Transformer decoder layers, a trainable weight matrix is ​​configured between the input and output of each of the Transformer decoder layers, and the trainable weight matrix is ​​used to adjust the strength of the residual connection between each of the Transformer decoder layers;

[0055] Step S202, obtaining pre-processed sample document data to construct a training set;

[0056] Step S203, training the basic model based on the multi-layer Transformer architecture based on the training set until a preset training termination condition is reached, thereby obtaining a trained large language model that integrates trainable residual connections and dual-scale convolutional Transformers.

[0057] In the embodiment of the present invention, the following is an exemplary detailed scenario description:

[0058] Get the base model based on a multi-layer Transformer architecture:

[0059] The server receives a natural language processing task and needs to build a large language model. First, the server creates a base model based on a multi-layer Transformer architecture. For example, this base model consists of 12 Transformer decoder layers, each of which contains a self-attention mechanism and a feedforward neural network. In each Transformer decoder layer, a two-scale convolution module is embedded. The convolution kernel size of one of the two-scale convolution modules is 3 and 7, which are used to capture local features of different scales, respectively. The outputs of these two convolution kernels are fused by splicing as the output of the Transformer decoder layer. In addition, a trainable weight matrix is ​​configured between the input and output of each Transformer decoder layer to adjust the strength of the residual connection. As training progresses, this weight matrix will continue to learn and adjust to better adapt to different language tasks and data.

[0060] Get the preprocessed sample document data to build a training set:

[0061] The server collects a large amount of text data from the Internet, including news articles, academic papers, novels, etc. Then, these data are preprocessed to construct a training set.

[0062] First prepare the document:

[0063] Read data from WARC files to avoid the complexity of extracting content from HTML and reduce the interference of irrelevant information.

[0064] Perform preliminary filtering of URLs to exclude fraudulent and illegal websites, filter based on a domain block list and URL scoring based on a specific word list, and exclude URLs from high-quality text datasets such as Wikipedia and arXiv.

[0065] Use the Trafilatura library and regular expressions to extract the main content from an HTML page, ignoring menus, headers, footers, ads, etc., limiting new lines to two consecutive lines and removing all URL links.

[0066] Use language classifiers such as fastText for language identification, select texts in the corresponding language and delete documents with language scores below a threshold.

[0067] Then filter:

[0068] A heuristic approach is adopted to formulate rules to remove documents with too many lines, paragraphs or n-grams repeated to reduce costs and improve efficiency.

[0069] It mainly retains natural language documents written by humans, removes machine-generated spam, and uses quality filtering heuristics to remove outliers based on criteria such as document length and symbol-to-word ratio.

[0070] Use a linear correction filter to remove content irrelevant to the text (such as likes, navigation buttons, etc.) to ensure data quality.

[0071] Finally, remove the duplicates:

[0072] The MinHash algorithm is used to calculate the approximate similarity between documents and remove document pairs with high overlap.

[0073] Use the suffix array to find exact matches between strings, removing segments that repeat more than k consecutive tokens.

[0074] Removed duplicated URLs in cross-CC dumps to ensure uniqueness of data.

[0075] After these preprocessing steps, a training set containing a large amount of high-quality text data is finally constructed for subsequent model training.

[0076] Train the basic model based on the training set:

[0077] The server inputs the constructed training set into the basic model based on the multi-layer Transformer architecture for training.

[0078] In the pre-training phase, the model performs unsupervised learning with large-scale text data, with the goal of learning the statistical characteristics and contextual relationships of the language. For example, the model is trained by predicting the next word, filling in the blanks, or other language tasks to capture the grammatical, semantic, and contextual information of the language. The server continuously adjusts the parameters of the model so that the model can better understand the patterns and regularities of the language.

[0079] After pre-training is completed, the supervised fine-tuning phase begins. The model is fine-tuned using labeled data for specific downstream tasks, such as text classification, question answering, and dialogue generation. The server adjusts the model's parameters based on these labeled data to optimize performance on specific tasks.

[0080] Next, we enter the reward model phase. We manually sort multiple responses generated by the same prompt to construct training samples. The server creates a sample set with relative merits by comparing the quality of different responses. This sorting information is used to train the reward model, enabling it to predict the quality scores of different responses based on human feedback.

[0081] Finally, we enter the reinforcement learning phase. We use the reward model to score the results generated by the supervised fine-tuning model, and use these scores as feedback signals to train the model through a policy optimization algorithm (such as PPO). The goal is to maximize the reward for the model's generated responses, and by continuously adjusting the generation strategy, the model can output higher quality results in specific tasks.

[0082] During the training process, the server continuously iterates and updates the model parameters until the preset training termination conditions are met, such as reaching a certain accuracy rate, loss value convergence, etc. At this point, the server obtains a trained large language model that integrates trainable residual connections and dual-scale convolutional Transformers.

[0083] For example, after millions of training iterations, the model's accuracy on some common natural language processing tasks reaches more than 90%, and the loss value stabilizes at a low level. At this time, it can be considered that the model has been trained and can be used in actual natural language processing tasks.

[0084] In the embodiment of the present invention, the training of the basic model based on the multi-layer Transformer architecture based on the training set can be implemented through the following examples.

[0085] Perform unsupervised learning based on the training set to complete the pre-training of the basic model based on the multi-layer Transformer architecture to learn the statistical characteristics and contextual relationships of the language;

[0086] Perform supervised fine-tuning on the pre-trained base model based on the multi-layer Transformer architecture according to the preset downstream tasks;

[0087] The basic model based on the multi-layer Transformer architecture after supervised fine-tuning is combined with the reward model for reinforcement learning through a preset strategy optimization algorithm.

[0088] In the embodiment of the present invention, the following is an exemplary detailed scenario description:

[0089] Unsupervised learning stage (pre-training of the basic model to learn language statistics and contextual relationships):

[0090] The server receives a large amount of unlabeled text data, which covers a variety of topics and styles.

[0091] At this stage, the server begins unsupervised learning of the basic model based on the multi-layer Transformer architecture. For example, the server will let the model process a news report. The model uses the self-attention mechanism to analyze the relationship between each word and other words in the report, and gradually understands the meaning and role of the word in the context.

[0092] For example, for the sentence "White clouds float in the sky, and the breeze blows gently on the earth.", the self-attention mechanism of the basic model will focus on the relationship between words such as "sky", "clouds", "breeze", and "earth". It will find that "sky" and "clouds" often appear at the same time, and "breeze" and "earth" also have a certain connection. By continuously processing such a large amount of text, the model gradually learns the statistical characteristics of the co-occurrence pattern of words in the language, grammatical structure, etc.

[0093] As more text is processed, the model begins to be able to capture more complex contextual relationships. For example, the word "floating" may have different meanings in different contexts, and its usage and semantics may be different when describing clouds and other objects. The model can accurately understand this difference by observing and learning from a large amount of text.

[0094] During this process, the server continuously adjusts the parameters of the model so that the model can learn the statistical characteristics and contextual relationships of the language more accurately. For example, the model will give higher weights to some frequently occurring word combinations so that these combinations can be better processed in subsequent tasks. After a long period of unsupervised learning, the basic model has initially mastered the basic laws and patterns of the language, laying a solid foundation for subsequent tasks.

[0095] Supervised fine-tuning phase (fine-tuning the pre-trained basic model according to the preset downstream tasks):

[0096] Once the pre-training phase is complete, the server begins supervised fine-tuning of the base model for specific downstream tasks.

[0097] Assume that the preset downstream task is text classification, such as classifying news articles into different categories such as politics, technology, and entertainment. The server will select a portion of news text with clear classification labels as fine-tuning data.

[0098] For a specific political news text, the server inputs it into the pre-trained basic model. The model encodes and processes the text based on the language knowledge learned previously. Then, the server calculates the difference between the model output and the true label based on the actual classification label of the text, that is, the loss value.

[0099] Through the back-propagation algorithm, the server propagates this loss value back to the various parameters of the model, and then adjusts these parameters to make the model's output closer to the true classification label. For example, if the model initially misclassifies a political news item as a technology item, the server will adjust the model's parameters related to vocabulary understanding, semantic analysis, etc., so that the model pays more attention to the characteristics of political texts.

[0100] In this process, the server will perform such fine-tuning operations on a large number of texts with classification labels, and continuously optimize the performance of the model on specific downstream tasks. After multiple iterations and adjustments, the model gradually adapts to the specific text classification task, and its accuracy on the task will gradually improve.

[0101] Reinforcement learning stage (the basic model after supervision fine-tuning is combined with the reward model for reinforcement learning through the preset strategy optimization algorithm):

[0102] After supervised fine-tuning is completed, the server introduces the reward model and combines it with the preset strategy optimization algorithm for reinforcement learning.

[0103] First, multiple responses generated by the same prompt are manually sorted to construct training samples. For example, for a question "How to improve learning efficiency?", the model generates multiple different answers, such as "arrange time reasonably", "do more exercises", "maintain a good learning environment", etc. The human beings sort these answers according to their own judgment to determine which answers are better.

[0104] The server then inputs these sorted samples into the reinforcement learning process. For each generated answer, the reward model will give a reward score based on the results of the manual sorting. For example, the answer "reasonable time arrangement" is considered to be better, and the reward model will give a higher score, while the answer "eating more chocolate can improve learning efficiency" may get a lower score.

[0105] Next, the server uses a policy optimization algorithm (such as PPO) to adjust the model's generation strategy. The PPO algorithm will try to adjust the way the model generates answers based on the score given by the reward model, so that the model is more inclined to generate answers with high reward scores. For example, if the model finds that generating the answer "arrange time reasonably" can get a higher reward, then it will be more inclined to generate similar answers in the subsequent generation process.

[0106] During the reinforcement learning process, the server will continuously let the model generate answers and adjust it according to the feedback from the reward model until the model achieves a good effect in generating answers. For example, after multiple rounds of reinforcement learning, the quality of the answers generated by the model has been significantly improved, which can more accurately meet the needs of users and give more reasonable answers to various different questions.

[0107] Through these three stages of training, namely unsupervised learning, supervised fine-tuning and reinforcement learning, the server successfully trained the basic model based on the multi-layer Transformer architecture, enabling it to perform well in various natural language processing tasks and provide users with better language processing services.

[0108] In the embodiment of the present invention, the unsupervised learning based on the training set can be implemented through the following examples.

[0109] The text data included in the training set is converted into a plurality of digital sequences through a word segmenter; the numbers included in the digital sequences are index numbers of the text data in the dictionary;

[0110] Obtain multiple high-dimensional vectors of high dimension by embedding the multiple digital sequences;

[0111] Performing reasoning operations according to the multiple high-dimensional directions to obtain index sequences corresponding to the multiple digital sequences;

[0112] The word segmenter is used in combination with the index sequence to perform text restoration to obtain the target natural language.

[0113] In the embodiment of the present invention, the following is an exemplary detailed scenario description:

[0114] Convert the training set text data into multiple digital sequences through the tokenizer:

[0115] The server receives a training set containing a huge amount of text covering a wide range of fields and styles.

[0116] Take a news article about the development of science and technology as an example, such as "In recent years, artificial intelligence technology has made great breakthroughs, and various intelligent devices have emerged, bringing many conveniences to people's lives." The server starts the word segmenter, which is like a language "scissors". It will cut the article into words according to certain rules. For example, it will cut "in recent years" into two words, "in recent years" and "in the future", and "artificial intelligence technology" into three words, "artificial", "intelligent", and "technology", and so on.

[0117] When the word segmenter has finished processing the entire article, each word is segmented separately. Then, the server assigns each word an index number in the dictionary. This dictionary is like a "dictionary" of a language, which records all possible words and their corresponding numbers. For example, the word "artificial" is assigned the number 1001, "intelligence" is assigned the number 1002, "technology" is assigned the number 1003, and so on. In this way, the entire article is converted into a series of numerical sequences, each number representing the position of a word in the dictionary.

[0118] Multiple digital sequences are embedded to obtain multiple high-dimensional vectors of high dimensions:

[0119] After obtaining the text sequence composed of numbers, the server begins to use embedding technology to convert these number sequences into high-dimensional vectors.

[0120] Embedding is like giving each word a unique "fingerprint", which is a high-dimensional vector. For the number sequence [1001, 1002, 1003] in the previous example, the embedding process will generate a high-dimensional vector for the number 1001 corresponding to "manual", such as [0.2, 0.3, 0.1, 0.4, 0.7, ...] (this is just an example, the actual vector dimension and value will vary depending on the model settings), and generate another high-dimensional vector for the number 1002 corresponding to "intelligence", such as [0.1, 0.8, 0.2, 0.3, 0.6, ...], and generate [0.3, 0.4, 0.7, 0.1, 0.9, ...] for the number 1003 corresponding to "technology".

[0121] In this way, the original digital sequence is converted into a series of high-dimensional vectors, each of which contains rich semantic information about the corresponding word. These high-dimensional vectors not only retain the basic meaning of the word, but also capture the semantic relationship and contextual information between words.

[0122] Perform inference operations based on multiple high-dimensional vectors to obtain index sequences corresponding to multiple digital sequences:

[0123] The server uses an internal multi-layer Transformer architecture to perform complex reasoning operations on these high-dimensional vectors.

[0124] To illustrate with a simple example, suppose a layer in the Transformer architecture receives the three high-dimensional vectors [0.2, 0.3, 0.1, 0.4, 0.7, ...], [0.1, 0.8, 0.2, 0.3, 0.6, ...], and [0.3, 0.4, 0.7, 0.1, 0.9, ...] generated earlier. Through the self-attention mechanism, this layer will pay attention to the relationship between each vector and other vectors, trying to understand their role and relationship in the entire text.

[0125] These relationships are then further processed and calculated through a feedforward neural network, transforming and combining vectors according to certain rules and algorithms. After this series of inference operations, the layer outputs a new vector representation.

[0126] As the multi-layer Transformer architecture continues to process, a high-dimensional vector representing the entire text will eventually be obtained. Then, through some specific decoding operations, this high-dimensional vector is converted into a corresponding digital sequence, that is, an index sequence corresponding to multiple digital sequences is obtained. This index sequence is like an encoded representation of the original text, which contains the semantic information and structural information of the original text.

[0127] Use the word segmenter combined with the index sequence to restore the text and obtain the target natural language:

[0128] When the server obtains the index sequence corresponding to the digital sequence, it will use the tokenizer again to restore the text.

[0129] The word segmenter will convert each index number back to the corresponding word according to the previously established dictionary. For example, if the number of a position in the index sequence is 1001, the word segmenter will look up the dictionary and find the word "manual" corresponding to the number 1001, and then add the word "manual" to the restored text.

[0130] In this way, each number in the index sequence is converted into a word in turn, and finally the entire index sequence is restored to the target natural language. For example, for the previous example, after the restoration of the word segmenter, the index sequence may be restored to "In recent years, artificial intelligence technology has made great breakthroughs, and various smart devices have emerged, bringing many conveniences to people's lives." In this way, the server successfully restores the digital sequence to the original natural language text, completing the conversion process from training data to understandable natural language.

[0131] In the embodiment of the present invention, the inference operation is performed according to the multiple high-dimensional vectors to obtain the index sequences corresponding to the multiple digital sequences, which can be implemented through the following examples.

[0132] Normalizing the high-dimensional vector corresponding to each layer of the Transformer architecture, and processing the normalized original features through a linear layer;

[0133] The processed original features are divided into blocks on the time axis to obtain the first features; the block operation adopts an overlap factor of 50%;

[0134] A multi-scale Transformer using a trainable residual connection structure and adding a convolution module uses an intra-frame Transformer and an inter-frame Transformer to respectively learn the short-term and long-term dependencies of the first feature to obtain a second feature, and the intra-frame Transformer and the inter-frame Transformer have the same structure;

[0135] The second feature generated by the multi-scale Transformer is subjected to a PReLU activation function and a linear layer to obtain the third feature;

[0136] Performing an overlap-add operation on the third feature to obtain a fourth feature; the overlap rate of the overlap-add operation is 0.5;

[0137] Processing the fourth feature through a ReLU activation function and a feed-forward network layer, using a nonlinear transformation to extract and learn the feature representation of each token vocabulary;

[0138] When reaching the last layer of Transformer architecture, the SoftMax function is used to obtain the sampling probability of all digital sequences, and the next digital sequence is sampled according to the preset sampling algorithm;

[0139] Repeat the steps of normalizing the high-dimensional vector corresponding to each layer of the Transformer architecture, and processing the normalized original features through a linear layer, to the step of processing the fourth feature through a ReLU activation function and a feedforward network layer, and using a nonlinear transformation to extract and learn the feature representation of each token vocabulary, until a preset number of cycles is reached, and then generating the index sequence.

[0140] In the embodiment of the present invention, the following is an exemplary detailed scenario description:

[0141] Normalize the high-dimensional vectors corresponding to each layer of the Transformer architecture, and process the normalized original features through the linear layer:

[0142] The server receives multiple high-dimensional vectors after embedding processing, which represent the characteristics of each word or element in the text.

[0143] Taking an article about historical events as an example, one of the high-dimensional vectors may represent the characteristics of the event "Battle of Red Cliffs". For this high-dimensional vector, the server first performs a normalization operation.

[0144] The purpose of normalization is to make each high-dimensional vector have a similar distribution at a specific scale to avoid the adverse effects of excessively large or small values ​​on subsequent operations. By calculating the mean and standard deviation of the high-dimensional vector, then subtracting the mean from the value of each element and dividing it by the standard deviation, the mean of the vector becomes 0 and the variance becomes 1.

[0145] Next, the server uses the linear layer to process the normalized original features. The linear layer is like a filter that performs linear transformations on the features by performing matrix multiplication and addition operations with the original features. For example, for the normalized high-dimensional vector of "Battle of Red Cliffs", the weight matrix of the linear layer may be a matrix of size [original feature dimension, new feature dimension]. Through matrix multiplication and addition operations, the original features are mapped to a new feature space to obtain the processed features.

[0146] In this way, each high-dimensional vector is processed by normalization and linear layers, ready for subsequent feature learning and relationship extraction.

[0147] The processed original features are divided into blocks on the time axis to obtain the first feature; the block operation uses a 50% overlap factor:

[0148] The server divides the original features processed by the linear layer into blocks on the time axis. Taking the processed features of "Battle of Red Cliffs" as an example, the time axis can be regarded as the sequence position of the text.

[0149] The block operation uses a 50% overlap factor, which means that each block has half of the overlap. For example, if the original feature sequence of length 100 is divided into blocks of size 60, then there are 30 elements overlapping between two adjacent blocks. The advantage of this is that it can better capture the local features in the text and the continuity of the sequence.

[0150] Through the block operation, a series of first features are obtained, each of which represents the local information of the original feature at different time segments. These first features retain part of the information of the original features, while having a certain degree of independence, which is convenient for subsequent processing and analysis.

[0151] The multi-scale Transformer with a trainable residual connection structure and a convolution module uses intra-frame Transformer and inter-frame Transformer to learn the short-term and long-term dependencies of the first feature respectively to obtain the second feature. The intra-frame Transformer and inter-frame Transformer have the same structure:

[0152] For the first feature obtained in each block, the server uses intra-frame Transformer and inter-frame Transformer to learn its short-term and long-term dependencies respectively.

[0153] The intra-frame Transformer focuses on the local dependencies within the current time segment. It analyzes the relationships between elements within the current block through self-attention mechanisms and other methods to capture local semantic and grammatical information. For example, in the first feature of "The Battle of Red Cliffs", the intra-frame Transformer will focus on the relationship between elements such as "Red Cliffs" and "Battle", as well as their contextual information within the current block.

[0154] The inter-frame Transformer focuses on learning long-term dependencies between different time segments. It can capture the connections between different parts of the entire text sequence across time intervals. For example, in the entire text sequence about the "Battle of Red Cliffs", the inter-frame Transformer can learn the relationship between the "Battle of Red Cliffs" and previous historical events and subsequent historical influences.

[0155] Through the synergy of these two Transformers, the first feature is processed from different angles to obtain a richer and more comprehensive second feature.

[0156] The second feature generated by the multi-scale Transformer is obtained through the PReLU activation function and the linear layer to obtain the third feature:

[0157] After the second feature is processed by the intra-frame and inter-frame Transformer, the server uses the PReLU activation function to perform nonlinear transformation on it.

[0158] The PReLU activation function is linear in the positive part and has a certain slope in the negative part, which can introduce nonlinear characteristics and enable the model to learn more complex functional relationships. Through the processing of the PReLU activation function, the information in the second feature is further activated and enhanced, highlighting important features and suppressing some unimportant information.

[0159] Then, the second feature after PReLU activation is processed by a linear layer, similar to the previous linear layer operation, to perform linear transformation and mapping on the feature to obtain the third feature.

[0160] The fourth feature is obtained by performing an overlap-add operation on the third feature; the overlap ratio of the overlap-add operation is 0.5:

[0161] The server performs overlap-addition operations on the third feature after PReLU activation and linear layer processing.

[0162] Since the previous block operation uses an overlap factor of 50%, half of the elements between two adjacent third feature blocks overlap when performing overlap addition. By adding these overlapping parts, more information can be retained while avoiding excessive duplication of information.

[0163] In this way, a fourth feature with richer information is obtained, which integrates the features of multiple time segments and reflects the overall characteristics of the text more comprehensively.

[0164] The fourth feature is processed through the ReLU activation function and the feedforward network layer, and the feature representation of each token vocabulary is extracted and learned using nonlinear transformation:

[0165] The server uses the ReLU activation function to process the fourth feature. The ReLU activation function is linear in the part greater than 0, which can maintain the transmission of information. At the same time, it changes the value to 0 in the part less than 0, which plays a sparse role.

[0166] Through the processing of the ReLU activation function, some unimportant or irrelevant information in the fourth feature is suppressed, while the important information is retained and enhanced.

[0167] Then, the fourth feature after ReLU activation is further processed by the feedforward network layer. The feedforward network layer is a fully connected neural network layer that performs matrix multiplication and addition operations on the fourth feature, performs nonlinear transformation and learning on the feature, and extracts the feature representation of each token vocabulary.

[0168] In this process, the model continuously learns and adjusts parameters to better extract and represent the features of each token vocabulary, thereby providing a more accurate basis for subsequent generation and reasoning tasks.

[0169] When reaching the last layer of the Transformer architecture, the SoftMax function is used to obtain the sampling probability of all digital sequences, and the next digital sequence is sampled according to the preset sampling algorithm:

[0170] As the multi-layer Transformer architecture is processed, it gradually approaches the last layer. When it reaches the last layer, the server uses the SoftMax function to calculate the probability distribution of each possible sequence of numbers.

[0171] The SoftMax function converts the input numerical value into a probability value so that the sum of all probabilities is 1. Each digital sequence corresponds to a probability value, and the size of the probability value reflects the possibility of the occurrence of the digital sequence.

[0172] Then, according to the preset sampling algorithm, samples are taken from these probability distributions to obtain the next digital sequence. The sampling algorithm can be different methods such as random sampling and greedy sampling, which are selected according to specific tasks and requirements.

[0173] In this way, the server gradually generates a digital sequence, continuously performs feature learning and sequence generation until the preset number of cycles is reached, and finally generates a complete index sequence.

[0174] During the entire process, the server continuously performs various operations and processing on high-dimensional vectors. Through the synergy of the multi-layer Transformer architecture, it gradually extracts and learns the features of the text, and finally generates an index sequence with semantic and structural information, providing an important foundation for subsequent natural language processing tasks.

[0175] In an embodiment of the present invention, the multi-scale Transformer using a trainable residual connection structure and adding a convolution module uses an intra-frame Transformer and an inter-frame Transformer to respectively learn the short-term and long-term dependencies of the first feature to obtain a second feature, which can be implemented through the following examples.

[0176] Assigning a unique representation to each position of the first feature using absolute position coding to obtain an intermediate first feature;

[0177] Performing a linear transformation of weights and biases on the intermediate first feature through a feedforward neural network, and introducing a nonlinear transformation through a ReLU function to obtain an intermediate second feature;

[0178] Processing the intermediate second feature through a multi-head attention mechanism to obtain an intermediate third feature;

[0179] Processing the intermediate third feature through a dual-scale convolution module to obtain an intermediate fourth feature and an intermediate fifth feature;

[0180] Processing the intermediate fourth feature and the intermediate fifth feature through the Mish activation function to obtain an intermediate sixth feature and an intermediate seventh feature;

[0181] The intermediate sixth feature and the intermediate seventh feature are spliced ​​to obtain an intermediate eighth feature;

[0182] The intermediate eighth feature is subjected to adaptive average pooling to obtain an intermediate ninth feature, and the intermediate ninth feature is processed by the feedforward neural network to obtain the second feature.

[0183] In the embodiment of the present invention, the following is an exemplary detailed scenario description:

[0184] Absolute position encoding is used to assign a unique representation to each position of the first feature to obtain the intermediate first feature:

[0185] The server gets the first feature which contains various information, like a box filled with various objects, each position has its specific element.

[0186] In order to make the model better understand this position information, absolute position encoding is introduced. Take a text about natural landscape description as an example, assuming that the first feature represents the characteristics of each word in the sentence, such as the position of words such as "mountains", "lakes" and "forests" in the sequence.

[0187] Absolute position encoding is like attaching a unique label to each position, which is implemented through a specific mathematical formula. For example, for the i-th position in the sequence, the calculation formula for its absolute position encoding may be:

[0188] [Pos_i=[sin(2πi / 10000^{2j / d}), cos(2πi / 10000^{2j / d})]]

[0189] Among them, d is the dimension size of the model, and j controls the periodic changes of the sine and cosine functions. In this way, each position is assigned a two-dimensional vector containing sine and cosine values ​​as its unique representation, that is, the first intermediate feature. In this way, the model can use position information to better understand the order and structural relationship of words in the text, as if giving each word a specific "coordinate" at its position in the text.

[0190] The weight and bias of the first intermediate feature are linearly transformed through the feedforward neural network, and the nonlinear transformation is introduced through the ReLU function to obtain the second intermediate feature:

[0191] The server feeds the intermediate first feature with absolute position encoding into the feedforward neural network.

[0192] The feedforward neural network is like a complex processing machine that performs a series of operations on the intermediate first feature. First, the intermediate first feature is multiplied by the weight matrix through matrix multiplication, and then the bias vector is added. This is the linear transformation process of weights and biases. For example, the intermediate first feature is a matrix with a dimension of [100, 512] (100 represents the number of words and 512 represents the feature dimension), the weight matrix is ​​[512, 1024], and the bias vector is

[1024] . After multiplication and addition operations, a new intermediate feature is obtained, and its dimension becomes [100, 1024].

[0193] Then, a nonlinear transformation is introduced through the ReLU function. The role of the ReLU function is to change the negative values ​​in the input to 0, and the positive values ​​remain unchanged. This allows the model to learn more complex nonlinear relationships and avoid problems such as gradient disappearance. For example, for each element in the intermediate feature, if its value is less than 0, it is changed to 0, and the part greater than or equal to 0 remains unchanged. After being processed by the ReLU function, the intermediate second feature is obtained, which retains the important information in the intermediate first feature and introduces nonlinear factors, allowing the model to process text data more flexibly.

[0194] The second feature in the middle is processed by the multi-head attention mechanism to obtain the third feature in the middle:

[0195] The server inputs the intermediate second features into the multi-head attention mechanism.

[0196] The multi-head attention mechanism is like multiple "little eyes" observing the intermediate second feature at the same time, capturing information from different angles. First, the intermediate second feature is projected into three different representation spaces through linear transformation to obtain query, key, and value respectively. For example, the intermediate second feature is a matrix with a dimension of [100, 1024]. Through three different linear transformations, query, key, and value matrices with dimensions of [100, 32*64] are obtained respectively (here 64 is the number of heads and 32 is the dimension of each head).

[0197] Then, the query, key, and value are divided into 64 heads, each with a dimension of 32. For each head, its attention weight is calculated, which is the dot product of the query and the key divided by the square root of the sum of the squares of the keys, and then normalized by the softmax function. Next, the attention weight is multiplied by the value, and the outputs of all heads are concatenated to obtain the result of the multi-head attention mechanism, which is the intermediate third feature. Through the multi-head attention mechanism, the model can simultaneously pay attention to information at different positions, capture long-distance dependencies in the text, and thus better understand the semantics of the text.

[0198] The middle third feature is processed by a dual-scale convolution module to obtain the middle fourth feature and the middle fifth feature:

[0199] The server feeds the intermediate third feature into the dual-scale convolution module.

[0200] The dual-scale convolution module consists of two convolution kernels of different sizes, which are used to capture local features of different scales. One convolution kernel size is 3 and the other is 11. When inputting the intermediate third feature, its dimension is converted from [100, 32*64] to [100, 64, 32] for one-dimensional depth convolution operation.

[0201] For a convolution kernel of size 3, it is divided into 64 groups in the depth direction, and each group performs a 3-point convolution operation, and the output dimension is still [100, 64, 32]. For a convolution kernel of size 11, it is also divided into 64 groups in the depth direction, and an 11-point convolution operation is performed, and the output dimension is also [100, 64, 32]. In this way, two feature maps of different scales are obtained, namely the middle fourth feature and the middle fifth feature. Through the dual-scale convolution module, the model can capture local detail features and global context information at the same time, enriching the feature representation.

[0202] The intermediate fourth feature and the intermediate fifth feature are processed by the Mish activation function to obtain the intermediate sixth feature and the intermediate seventh feature:

[0203] The server applies the Mish activation function to the intermediate fourth feature and the intermediate fifth feature respectively.

[0204] The expression of the Mish activation function is (Mish(x)=x*tanh(softplus(x))), where (softplus(x)=ln(1+e^x)). The Mish activation function has a smaller slope in the negative part and a larger slope in the positive part, which can better balance the nonlinear characteristics of the model.

[0205] Each element in the intermediate fourth feature and the intermediate fifth feature is processed by the Mish activation function to obtain a richer and more nonlinear intermediate sixth feature and intermediate seventh feature. The introduction of the Mish activation function enables the model to better mine the potential information in the data and improve the model's expressiveness.

[0206] The sixth feature in the middle and the seventh feature in the middle are combined to obtain the eighth feature in the middle:

[0207] The server concatenates the intermediate sixth feature and the intermediate seventh feature processed by the Mish activation function in the feature dimension.

[0208] Since they have the same dimensions, the concatenation operation is to simply connect the two feature matrices in the channel dimension to obtain an intermediate eighth feature with a larger dimension. In this way, the model integrates the feature information obtained by convolution operations of different scales, enabling it to better process complex text data.

[0209] The eighth feature in the middle is obtained by adaptive average pooling to obtain the ninth feature in the middle, and the ninth feature in the middle is processed by a feedforward neural network to obtain the second feature:

[0210] The server performs an adaptive average pooling operation on the middle eighth feature.

[0211] Adaptive average pooling can automatically adjust the size of the pooling window according to the input features to adapt to features of different scales. It traverses all channels of the middle eighth feature, divides the features into regions of fixed size, and then takes the average of the feature values ​​in each region as the pooling result of the channel. In this way, the size of the output middle ninth feature remains fixed while retaining the main information of the input feature.

[0212] Finally, the ninth feature in the middle is sent to the feedforward neural network for processing. Just like the previous steps, the ninth feature in the middle is further transformed and learned through operations such as linear transformation and ReLU function, and finally the second feature is obtained. The second feature contains rich semantic information and feature representation, which is an important result obtained after the intra-frame Transformer processing, and provides a solid foundation for subsequent processing and analysis.

[0213] In the embodiment of the present invention, the processing of the intermediate ninth feature through the feedforward neural network to obtain the second feature can be implemented through the following example.

[0214] The intermediate ninth feature in the intra-frame Transformer and the inter-frame Transformer is added to the input of each feedforward neural network layer through a trainable residual connection and then subjected to a layer normalization operation to obtain the second feature.

[0215] In the embodiment of the present invention, the following is an exemplary detailed scenario description:

[0216] In the process of processing the natural language task, the server reaches the step of processing the intermediate ninth feature through a feedforward neural network to obtain the second feature.

[0217] Taking a long text about a historical story as an example, the ninth feature in the middle is like a set of key features extracted from the text after a series of complex operations. It contains various semantic, structural and other information about the story.

[0218] In the intra-frame Transformer, the feedforward neural network of each layer starts to process the middle ninth feature. First, the middle ninth feature of the current layer is fed into the feedforward neural network as input. The feedforward neural network here is like a complex network structure composed of multiple neurons, with multiple hidden layers and output layers.

[0219] In the process of inputting into the feedforward neural network, the output of the previous layer is added through a trainable residual connection. The role of the trainable residual connection is to enable information to flow more smoothly in the network, avoiding information loss or transmission difficulties as the number of network layers increases. For example, the ninth feature in the middle of the input of the current layer is a tensor of shape [B, C, H] (B represents the batch size, C represents the number of channels, and H represents the height or width of the feature, etc.), and the output of the previous layer is also a tensor of the same shape after some processing. These two tensors are added through a trainable residual connection, that is, the elements in the corresponding positions are added, so that the input of the current layer contains both new information and retains the information of the previous layer, which facilitates the transmission and fusion of information.

[0220] The result after adding the trainable residual connection then enters the layer normalization operation. The purpose of layer normalization is to make the input of each layer close to 0 and the variance close to 1 after various operations, thereby accelerating the training process and improving the stability of the model. Specifically, for the tensor added by the trainable residual connection, calculate its mean and standard deviation in each dimension. Then, subtract the mean from the value of each element and divide it by the standard deviation to make the data distribution more stable. After the layer normalization operation, an adjusted intermediate feature is obtained.

[0221] The same process is repeated in the inter-frame Transformer. Each layer of the feedforward neural network takes the intermediate features after trainable residual connection and layer normalization as input and continues a series of operations and processing. With the sequential processing of multiple layers of feedforward neural networks in the intra-frame Transformer and inter-frame Transformer, the intermediate ninth feature is continuously refined, adjusted and optimized, gradually integrating more semantic and structural information.

[0222] After the entire feedforward neural network processing, the second feature is finally obtained. This second feature is like the essence of the text after deep processing and refinement. It contains the most core semantic and structural information about the entire text, and can provide strong support for subsequent language generation, semantic understanding and other tasks. For example, when performing text generation tasks, the server can generate new text content related to the original text based on this second feature; when performing semantic understanding tasks, this second feature can help the server understand the meaning and intention of the text more accurately.

[0223] In an embodiment of the present invention, the trainable residual connection is represented by the formula: y=F(W*x);

[0224] Among them, x represents the input, W represents the trainable weight matrix; in the weight matrix, α close to the current layer is initialized to 1, while the previous farther layer is initialized to 0; F(x) represents the nonlinear mapping function from input to output, and y represents the output.

[0225] In the embodiment of the present invention, the following is an exemplary detailed scenario description:

[0226] Trainable residual connections play a crucial role in the server's process of building and training large language models based on trainable residual connections.

[0227] Taking a specific language translation task as an example, the server receives a text in the source language, such as an English text about an introduction to a technology product.

[0228] When entering the network structure with trainable residual connections, we first look at the input part. The input is the vector representation of the English text to be translated after various preprocessing and feature extraction, which is recorded as ({X}).

[0229] The trainable weight matrix is ​​like a regulating valve in this process. It dynamically adjusts the flow of information according to the process of network training and the relationship between different layers. Taking the first layer of the network as an example, the weight matrix here is recorded as (W_1), where (α) close to the current layer is initialized to 1, which means that at the beginning of the first layer, the input ({X}) is almost intactly passed to the output through the residual connection. For the previous more distant layers, their (α) is initialized to 0, which means that in the early stages, these layers have less influence on the current layer.

[0230] As the network goes deeper layer by layer, the value of (α) will be adjusted according to the training situation. For example, in the second layer, (α) may be adjusted to a value between 0 and 1 based on the information interaction between the second layer and the first layer and the contribution to the entire translation task. If the second layer finds through interaction with the first layer that keeping more input information helps improve translation accuracy, then (α) may be closer to 1; conversely, if it is found that more inter-layer interaction and information fusion are needed, (α) may be adjusted to a value closer to 0, but a certain amount of input information will still be retained.

[0231] By multiplying the input ({X}) with the trainable weight matrix (W_1), that is, (W_1{X}), and then adding it to the input ({X}), that is, ({X}+W_1{X}), an intermediate result is obtained. This intermediate result is actually the fusion of the input ({X}) after some transformation of the weight matrix (W_1) and the input itself.

[0232] Next, this intermediate result will pass through a nonlinear mapping function (f) from input to output. The nonlinear mapping function (f) can be an activation function such as ReLU (rectified linear unit) and Leaky ReLU, which can introduce nonlinear factors and enable the model to learn more complex relationships. For example, for the ReLU function, when an element in the intermediate result is greater than 0, it remains unchanged; when an element is less than 0, it is changed to 0. This allows the model to filter out important information and suppress some unimportant information, so as to better adapt to various situations in language translation tasks.

[0233] After being processed by the nonlinear mapping function (f), the final output is obtained, denoted as ({Y}). ({Y}) is the output after being processed by the trainable residual connection, which contains the information of the input ({X}) as well as the new information and features introduced by the weight matrix and the nonlinear mapping function.

[0234] In subsequent network layers, the same trainable residual connection mechanism will continue to play a role. The input of each layer is the output of the previous layer after being processed by the trainable residual connection. By continuously adjusting the weight matrix and nonlinear mapping function, information is gradually transmitted, integrated and optimized in the network, and ultimately the entire model can accurately perform tasks such as language translation.

[0235] For example, in the translation process, if some complex grammatical structures or semantic expressions are encountered, the model can better retain and utilize the knowledge about grammar and semantics learned by the previous layer through the trainable residual connection mechanism, while being able to adjust and innovate according to the characteristics of the current layer, so as to more accurately translate the source language text into the target language text. When processing different types of scientific texts, literary texts or news texts, the trainable residual connection mechanism can flexibly adjust the flow and processing of information according to the characteristics and needs of the text, making the model more generalizable and adaptable.

[0236] In order to more clearly describe the solution provided by the embodiment of the present invention, a relatively complete implementation method is provided below. Figure 2 , Figure 2 A schematic diagram of the architecture of a large language model based on trainable residual connections and dual-scale convolutional Transformer provided in an embodiment of the present invention.

[0237] In order to address the shortcomings of current language models in processing complex language structures, long-distance dependencies, and capturing multi-scale semantic features, this paper proposes an innovative large language model architecture - "a large language model architecture based on trainable residual connections and dual-scale convolutional Transformers". This solution combines two core technological innovations, thereby achieving a significant improvement in model performance and generalization capabilities.

[0238] First, in response to the fixedness and limitations of the residual connection method in the traditional Transformer, this paper pioneered the introduction of a new trainable residual connection structure. This innovative design gives the model the ability to dynamically adjust the information flow path, enabling it to intelligently optimize the hierarchical transmission process of information according to the characteristics of the input data and task requirements. This not only significantly enhances the model's ability to capture deep-level language features, but also promotes the effective fusion of information between different levels, thereby greatly improving the model's expressiveness and depth of understanding. Please refer to Figure 4 , Figure 4 A schematic diagram of the internal structure of a dual-scale convolutional Transformer with a trainable residual connection structure provided by an embodiment of the present invention;

[0239] Secondly, in order to make up for the shortcomings of traditional models in capturing multi-scale semantic features, the present invention cleverly embeds a dual-scale convolution module in the Transformer architecture. This design can simultaneously capture local detail features and global context information in language data by processing convolution kernels of different sizes in parallel, thus achieving a comprehensive and in-depth analysis of language information. This innovation greatly enriches the feature dimensions that can be extracted by the model and significantly improves the model's ability to process complex language structures and understand implicit semantic relationships. Please refer to Figure 3 , Figure 3 Schematic diagram of the internal structure of the TransformerX and multi-scale Transformer modules with dual-scale convolution provided in the embodiments of the present invention.

[0240] Based on the above two core technology innovations, the present invention successfully built a new component - TransformerX. This component not only inherits the advantages of the Transformer architecture, but also significantly improves the performance and generalization ability of the model in processing complex language tasks through the introduction of trainable residual connections and dual-scale convolution modules. Through the introduction of TransformerX, the large language model architecture of the present invention has performed well in multiple natural language processing tasks, demonstrating its strong application potential.

[0241] In order to achieve the above-mentioned purpose, the technical solution adopted by the present invention is: a large language model based on trainable residual connection and dual-scale convolutional Transformer, which specifically includes the following steps:

[0242] To build a large language model, first of all, a basic model based on the TransformerX architecture is designed. The model consists of multiple TransformerX decoder layers, each of which contains a self-attention mechanism and a feedforward neural network. In order to enhance the local feature capture capability of the model, a dual-scale convolution module is embedded in each TransformerX layer. The dual-scale convolution module consists of two convolution kernels of different sizes, which are used to capture local features of different scales. Depth-separable convolution is used to replace the traditional standard convolution to reduce the number of parameters and computational complexity of the model. The outputs of the two convolution kernels are combined by splicing or weighted summation as the output of the module and passed to the subsequent TransformerX layer. In traditional Transformers, residual connections are fixed to keep the flow of information. In the present invention, a trainable residual connection structure is proposed. This structure allows the model to learn how to better adjust the flow path of information during training, thereby optimizing the efficiency of information transmission. In specific implementation, a trainable weight matrix is ​​added between the input and output of each TransformerX layer to adjust the strength of the residual connection.

[0243] Data collection and processing: collect text data from the Internet and other channels to obtain training sets. Specifically,

[0244] S1, document preparation, document preparation includes data reading, URL filtering, text extraction and language identification. The specific process of S1 is:

[0245] S1.1, Data reading: Text data can be extracted from WET files or WARC files. In order to avoid the complexity of extracting content from HTML, it is preferred to read data from WARC files to reduce the interference of irrelevant information;

[0246] S1.2, Filtering URLs: Before processing text data, perform preliminary filtering on URLs to exclude fraudulent and illegal websites. Filtering is based on a domain block list and URL scoring based on a specific word list. Also exclude URLs from high-quality text datasets (such as Wikipedia and arXiv);

[0247] S1.3, Extract Text: Extract the main content from the HTML page, ignoring menus, headers, footers, and advertisements, etc. Use the Trafilatura library and regular expressions for text extraction, and finally limit new lines to two consecutive lines and remove all URL links;

[0248] S1.4, language identification: language identification is performed using language classifiers such as fastText, selecting texts in the corresponding language and deleting documents with language scores below a threshold;

[0249] S2, filtering. The purpose of filtering is to improve the quality of text, remove repeated paragraphs, irrelevant content and non-natural language, including document level and line level filtering. The specific process of S2 is as follows:

[0250] S2.1, Duplicate document removal: adopt heuristic methods and formulate rules to delete documents with too many repeated lines, paragraphs or n-grams to reduce costs and improve efficiency;

[0251] S2.2, document filtering: mainly retain natural language documents written by humans, remove machine-generated spam, use quality filtering heuristics, and remove outliers based on criteria such as document length and symbol-to-word ratio;

[0252] S2.3, row-level filtering: Use linear correction filters to remove content that is not related to the text (such as likes, navigation buttons, etc.) to ensure data quality;

[0253] S3, deduplication. After filtering, although the data quality is improved, there are still duplicate documents. Combining fuzzy matching and exact matching methods to deduplication, the specific process of S3 is as follows:

[0254] S3.1, fuzzy deduplication: use the MinHash algorithm to calculate the approximate similarity between documents and delete document pairs with high overlap;

[0255] S3.2, accurate deduplication: Use suffix arrays to find exact matches between strings and delete paragraphs with more than k consecutive tokens;

[0256] S3.3, URL deduplication: remove the URLs that are repeatedly accessed in the cross-CC dump to ensure the uniqueness of the data;

[0257] The training process of a large language model. Specifically,

[0258] S1, Pre-training: In this stage, the model performs unsupervised learning through large-scale text data, with the goal of learning the statistical characteristics and contextual relationships of the language. The model is trained by predicting the next word, filling in the blanks, or other language tasks to capture the grammatical, semantic, and contextual information of the language.

[0259] S2, supervised fine-tuning: After pre-training, the model will be supervised fine-tuned on specific downstream tasks. This process usually uses labeled data, and the model learns based on this data to optimize performance on specific tasks, such as text classification, question answering, dialogue generation, etc.

[0260] S3, Reward Modeling: In this stage, multiple responses generated by the same prompt are manually ranked to construct training samples. By comparing the quality of different responses, a sample set with relative advantages and disadvantages is created. This ranking information is used to train the reward model so that it can predict the quality scores of different responses based on human feedback. Specifically, the goal of the reward model is to learn a function that can assign a score to each generated response to reflect its quality in a specific task.

[0261] S4, Reinforcement Learning: In the reinforcement learning phase, a reward model is used to score the results generated by the supervised fine-tuning (SFT) model. Using these scores as feedback signals, the model is trained through a policy optimization algorithm (such as PPO, Proximal Policy Optimization). The goal is to maximize the rewards for the model's generated responses, and by continuously adjusting the generation strategy, the model can output higher quality results in specific tasks. Finally, after multiple rounds of optimization and training, the final model is obtained, and its generation ability has been significantly improved, which can better meet the needs and expectations of users.

[0262] The significant advantage of the present invention is that it innovatively introduces the TransformerX component within and between frames of the traditional large language model. This component embeds a dual-scale convolution module and cleverly uses deep separable convolution to replace traditional convolution, which greatly enhances the model's ability to capture local features. At the same time, this design also significantly reduces the number of parameters and computational overhead of the model, achieving efficient and lightweight model. Not only that, TransformerX also incorporates additional feedforward neural network structures, advanced activation functions, and layer normalization techniques, which further enhance the model's representation ability and robustness, making it more handy when dealing with complex and changeable natural language tasks. It is particularly worth mentioning that the present invention pioneered the use of a trainable residual connection structure to replace the traditional fixed residual connection method. This revolutionary design enables the model to autonomously transmit hierarchical information, significantly optimizes the training process, and improves the optimization effect of the model. Through the integration of this series of innovative designs and technical applications, the present invention has demonstrated excellent performance in multiple natural language processing tasks without significantly increasing model parameters. This has undoubtedly injected new vitality into the development of natural language processing and promoted the field to move towards a more efficient and intelligent direction. Figure 5 , Figure 5 A schematic diagram of the internal structure of a dual-scale convolution module and a single convolution module provided in an embodiment of the present invention;

[0263] The core of the large language model architecture of this embodiment lies in two key improvements to the Transformer structure. First, an innovative trainable residual connection structure is proposed to replace the traditional fixed connection method in Transformer. This new structure can autonomously transmit hierarchical information, significantly improving the expressiveness and generalization capabilities of the model. Secondly, a dual-scale convolution module is incorporated to more effectively capture semantic features of different scales and enhance the model's ability to process complex language information. Finally, by integrating these two technologies, a new TransformerX component was launched, and an excellent large language model architecture was built on this basis.

[0264] The above-mentioned solution of the present invention is specifically applied to the following embodiments:

[0265] S1 and S2 in step 3) of this embodiment are specifically:

[0266] S1, the input text is first converted into a sequence of numbers by the tokenizer. These numbers are the index numbers of the words in the dictionary (vocab);

[0267] S2, obtains a high-dimensional vector through embedding of the digital sequence;

[0268] S3, the above vectors are subjected to complex reasoning operations through the decoder to generate the digital index of the next word. The entire index sequence can be generated by looping the operations multiple times. Further, the specific process of S3 is as follows:

[0269] S3.1, the input code represents h0 after layer normalization, the specific formula is:

[0270]

[0271] Where x is the input data, μ is the mean of the input data in a certain dimension, and the calculation formula is:

[0272]

[0273] D is the total number of dimensions of the input data, i is a certain dimension of the input data. σ is the standard deviation of the input data in a certain dimension, and the calculation formula is:

[0274]

[0275] ε is a very small constant used to avoid the denominator being zero. g is a learnable scaling parameter, b is a learnable bias parameter, and ⊙ represents element-wise multiplication.

[0276] Layer normalization normalizes the input data so that the input of each layer is within the distribution range of mean 0 and variance 1, thereby alleviating the internal covariate shift problem in the neural network;

[0277] S3.2, use the linear layer to process the layer-normalized feature h0. The linear layer is a commonly used network layer in deep learning, also known as a fully connected layer or an affine layer. The specific formula is:

[0278] y=Wx+b,

[0279] Where x is the input vector, W is the weight matrix, b is the bias vector, and y is the output vector. This formula means that the input vector x is obtained by matrix multiplication and addition of the bias term to obtain the output vector y.

[0280] Specifically, if x is a vector of dimension n, then W is a matrix of dimension m×n, where m represents the weight of each neuron for each input feature; b is a vector of dimension m, representing the bias of each neuron; y is an output vector of dimension m, representing the output result of each neuron;

[0281] S3.3, the layer-normalized feature h0 is divided into blocks on the time axis to obtain h1, and the block operation adopts an overlap factor of 50%;

[0282] S3.4, a multi-scale Transformer with a trainable residual connection structure and a convolution module is used to learn the short-term and long-term dependencies of feature h1 using intra-frame Transformer and inter-frame Transformer to obtain feature h2. The intra-frame Transformer and inter-frame Transformer have the same structure, specifically:

[0283] S3.4.1, use absolute position encoding to assign a unique representation to each position of the input feature, and obtain feature X1. The specific formula of absolute position encoding is:

[0284]

[0285] Among them, pos represents the position index in the sequence, i represents the dimension index in the position encoding vector, and d represents the dimension size of the model. The sine function formula maps the position index to a small numerical range, and uses the periodic characteristics of the sine function to ensure that the encoding values ​​of adjacent positions are different. Similar to the sine function, the cosine function also has periodic characteristics. The cosine function is used to encode odd dimensions to further increase the diversity of encoding. The calculated results of the sine and cosine functions are combined into a position encoding vector. The position encoding vector is added to the audio feature vector in the time window to obtain the final position encoding representation. By introducing absolute position encoding, the model can use information at different positions in the sequence to better understand the structure and sequential relationship of text information;

[0286] 3.4.2, feature X2 is obtained through a feedforward neural network. In the feedforward neural network, the input is first transformed linearly by the linear layer with weights and biases, and then a nonlinear transformation is introduced by the ReLU function. The ReLU function changes the part of the input value less than 0 to 0, and keeps the part of the input value greater than or equal to 0 unchanged. The mathematical expression is:

[0287] ReLU(x)=max(0,x)

[0288] Using the ReLU activation function in a feedforward neural network can introduce nonlinearity, alleviate the gradient vanishing problem, extract useful features, and improve computational efficiency, which helps improve the performance of speech separation tasks;

[0289] 3.4.3, feature X2 is used to obtain feature X3 through a multi-head self-attention mechanism. The specific steps are as follows:

[0290] First, the input sequence X is projected into three different representation spaces: query, key, and value through linear transformation. This is achieved through the following formula:

[0291] Q=XW q

[0292] K=XW k

[0293] V=XW v

[0294] Where W q , W k , W v is a learnable weight matrix, which is used for linear transformation of query, key and value respectively. Next, the obtained Q, K, V are divided into 8 heads (or subspaces), each with a dimension of 32:

[0295]

[0296] Among them, Q i , K i and V i is the query, key, and value of the i-th head. Then, the attention weight of each head is calculated:

[0297]

[0298] Among them, softmax represents the normalization of the attention score, d h is the dimension of each head, which is 32. Divided by The purpose is to scale the attention scores so as to better handle attention in different dimensions. Finally, the outputs of all heads are concatenated to obtain the result of the multi-head attention mechanism;

[0299] S3.4.4, feature X3 is converted into features X41 and X42 through a dual-scale convolution module. The internal structures of the two convolution modules are the same. The specific steps are as follows:

[0300] First, the input tensor x is dimensionalized and the shape is changed from [B, N, L] to [B, L, N], where B represents the batch size, N represents the number of channels, and L represents the sequence length.

[0301] Next, we use one-dimensional depthwise convolution to perform convolution operations on the features simultaneously. The number of input and output channels of the convolution layer is 256, one convolution kernel size is 3, the other is 11, the stride is 1, and the padding size is 1. The input and output channels are divided into 256 groups for convolution calculation. The channels within each group are independent of each other, and no convolution operation is performed between them.

[0302] Next, a one-dimensional point-by-point convolution is used to convolve the output of the depthwise convolution, that is, the position information on each channel is exchanged. The number of input and output channels of this convolution layer is still 256, the size of the convolution kernel is 1, the stride is 1, and there is no padding;

[0303] S3.4.5, features X41 and X42 are obtained by Mish activation function to obtain features X51 and X52. The expression of Mish activation function is as follows:

[0304] Mish(x) = x·tanh(ln(1+e x ))

[0305] S3.4.6, feature X51 and X52 are concatenated to obtain feature X6. The specific steps are as follows:

[0306] The shape of feature X51 is [B, L1, N]. The shape of feature X52 is [B, L2, N], where B represents the batch size, N represents the number of channels, and L represents the sequence length. The concatenation operation connects the two feature maps in the depth dimension to generate a deeper feature X6 with a shape of [B, L1+L2, N].

[0307] S3.4.7, feature X6 is obtained by adaptive average pooling to obtain feature X7, the steps are: traverse all channels of feature X6, divide feature X6 into regions of fixed size, and take the average of the feature values ​​in each region as the pooling result of the channel, so that the output size remains fixed;

[0308] S3.4.8, feature X7 is processed again by the feedforward neural network;

[0309] In the intra-frame Transformer and inter-frame Transformer, the input of each feed-forward neural network layer is added to the output through a trainable residual connection and then undergoes layer normalization. The convolution module and the multi-head attention module also use this method, and the input is added to the output through a trainable residual connection and then undergoes layer normalization.

[0310] The expression of the trainable residual connection is:

[0311] y=F(W*x)

[0312] Among them, x represents input, and W represents a trainable weight matrix. In the weight matrix, the alpha close to the current layer is initialized to 1, while the previous and farther layers are initialized to 0. After training, the model can autonomously select the information of the required layer. F(x) represents the nonlinear mapping function from input to output, and y represents the output. Through training, the value of W is learned and multiplied by F(x), and then added to the input x, which realizes the trainable residual connection;

[0313] S3.5, feature h2 generated by multi-scale Transformer is passed through PReLU activation function and linear layer to obtain feature h3. The expression of PReLU activation function is:

[0314] PReLU(x)=max(0,x)+a·min(0,x)

[0315] S3.6, perform overlap-addition operation on feature h3 to obtain feature h4, with an overlap ratio of 0.5. This step is used to restore the divided features to their original length;

[0316] S3.7, feature h4 is processed through the ReLU activation function and the feedforward network layer, and the feature representation of each token vocabulary is extracted and learned using nonlinear transformation. After multiple layers, the above process is repeated until the last layer, and finally the sampling probability of all digital sequences is obtained through the SoftMax function, and the next digital sequence is sampled according to a specific sampling algorithm;

[0317] S3.8, after multiple cycles of calculation, the entire index sequence can be generated;

[0318] S4 uses a tokenizer to restore these digital sequences to natural language that can be understood by humans.

[0319] Through the large language model architecture based on trainable residual connections and dual-scale convolutional Transformer proposed in the present invention, significant experimental effect improvement has been successfully achieved in the field of natural language processing.

[0320] Please refer to Figure 6 , Figure 6 A training device 110 for a large language model based on a trainable residual connection and a dual-scale convolutional Transformer provided in an embodiment of the present invention includes:

[0321] An acquisition module 1101 is used to acquire a basic model based on a multi-layer Transformer architecture, wherein the basic model includes a plurality of cascaded Transformer decoder layers, each of the Transformer decoder layers includes a self-attention mechanism and a feedforward neural network, each of the Transformer decoder layers is embedded with a dual-scale convolution module, the dual-scale convolution modules of different scales are used to capture different local features, the outputs of the dual-scale convolution modules are fused as the outputs of the corresponding Transformer decoder layers, a trainable weight matrix is ​​configured between the input and output of each Transformer decoder layer, and the trainable weight matrix is ​​used to adjust the strength of the residual connection between each Transformer decoder layer; obtaining preprocessed sample document data to construct a training set;

[0322] The training module 1102 is used to train the basic model based on the multi-layer Transformer architecture based on the training set until a preset training termination condition is reached, thereby obtaining a trained large language model that integrates trainable residual connections and dual-scale convolutional Transformers.

[0323] It should be noted that the implementation principle of the aforementioned large language model training device 110 based on trainable residual connection and dual-scale convolutional transformer can refer to the implementation principle of the aforementioned large language model training method based on trainable residual connection and dual-scale convolutional transformer, which will not be repeated here. It should be understood that the division of the various modules of the above device is only a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also be implemented in the form of hardware; some modules can also be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example, the large language model training device 110 based on trainable residual connection and dual-scale convolutional transformer can be a separately established processing element, or it can be integrated in a chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and a processing element of the above device calls and executes the functions of the large language model training device 110 based on trainable residual connection and dual-scale convolutional transformer. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by an integrated logic circuit of hardware in the processor element or an instruction in the form of software.

[0324] For example, the above modules may be one or more integrated circuits configured to implement the above methods, such as one or more application specific integrated circuits (ASIC), or one or more microprocessors (digital signal processors, DSP), or one or more field programmable gate arrays (FPGA), etc. For another example, when a module above is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0325] The embodiment of the present invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned large language model training device 110 based on trainable residual connection and dual-scale convolutional Transformer. Figure 7 As shown, Figure 7 The computer device 100 is a block diagram of a structure of a computer device 100 provided in an embodiment of the present invention. The computer device 100 includes a training device 110 for a large language model based on a trainable residual connection and a dual-scale convolutional Transformer, a memory 111, a processor 112, and a communication unit 113.

[0326] To achieve data transmission or interaction, the memory 111, the processor 112 and the communication unit 113 are electrically connected to each other directly or indirectly. For example, the electrical connection between these elements can be achieved through one or more communication buses or signal lines. The training device 110 for a large language model based on trainable residual connections and dual-scale convolutional transformers includes at least one software function module that can be stored in the memory 111 in the form of software or firmware or solidified in the operating system (OS) of the computer device 100. The processor 112 is used to execute the training device 110 for a large language model based on trainable residual connections and dual-scale convolutional transformers stored in the memory 111, such as the software function modules and computer programs included in the training device 110 for a large language model based on trainable residual connections and dual-scale convolutional transformers.

[0327] An embodiment of the present invention provides a readable storage medium, which includes a computer program. When the computer program is running, it controls the computer device where the readable storage medium is located to execute the aforementioned training device 110 for the large language model based on trainable residual connection and dual-scale convolutional Transformer.

[0328] For illustrative purposes, the foregoing description is made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the present disclosure to the precise form disclosed. Numerous modifications and variations are possible in accordance with the above teachings. These embodiments are selected and described in order to best illustrate the principles of the present disclosure and its practical application, so that those skilled in the art can best utilize the present disclosure and utilize various embodiments with different modifications to suit the intended specific application.

Claims

1. A training method for a large language model based on a trainable residual connection and a dual-scale convolutional Transformer, characterized in that: include: Obtain a basic model based on a multi-layer Transformer architecture, wherein the basic model includes a plurality of cascaded Transformer decoder layers, each of the Transformer decoder layers includes a self-attention mechanism and a feedforward neural network, each of the Transformer decoder layers is embedded with a dual-scale convolution module, the dual-scale convolution modules of different scales are used to capture different local features, the outputs of the dual-scale convolution modules are fused as the outputs of the corresponding Transformer decoder layers, a trainable weight matrix is ​​configured between the input and output of each of the Transformer decoder layers, and the trainable weight matrix is ​​used to adjust the strength of the residual connection between each of the Transformer decoder layers; Obtain preprocessed sample document data to build a training set; The basic model based on the multi-layer Transformer architecture is trained based on the training set until a preset training termination condition is reached, thereby obtaining a trained large language model that integrates trainable residual connections and dual-scale convolutional Transformers.

2. The method according to claim 1, characterized in that The training of the basic model based on the multi-layer Transformer architecture based on the training set includes: Perform unsupervised learning based on the training set to complete the pre-training of the basic model based on the multi-layer Transformer architecture to learn the statistical characteristics and contextual relationships of the language; Perform supervised fine-tuning on the pre-trained base model based on the multi-layer Transformer architecture according to the preset downstream tasks; The basic model based on the multi-layer Transformer architecture after supervised fine-tuning is combined with the reward model for reinforcement learning through a preset strategy optimization algorithm.

3. The method according to claim 2, characterized in that The performing unsupervised learning according to the training set comprises: The text data included in the training set is converted into a plurality of digital sequences through a word segmenter; the numbers included in the digital sequences are index numbers of the text data in the dictionary; Obtain multiple high-dimensional vectors of high dimension by embedding the multiple digital sequences; Performing reasoning operations according to the multiple high-dimensional directions to obtain index sequences corresponding to the multiple digital sequences; The word segmenter is used in combination with the index sequence to perform text restoration to obtain the target natural language.

4. The method according to claim 3, characterized in that The performing inference operations according to the multiple high-dimensional vectors to obtain index sequences corresponding to the multiple digital sequences includes: Normalizing the high-dimensional vector corresponding to each layer of the Transformer architecture, and processing the normalized original features through a linear layer; The processed original features are divided into blocks on the time axis to obtain the first features; the block operation adopts an overlap factor of 50%; A multi-scale Transformer using a trainable residual connection structure and adding a convolution module uses an intra-frame Transformer and an inter-frame Transformer to respectively learn the short-term and long-term dependencies of the first feature to obtain a second feature, and the intra-frame Transformer and the inter-frame Transformer have the same structure; The second feature generated by the multi-scale Transformer is subjected to a PReLU activation function and a linear layer to obtain the third feature; Performing an overlap-add operation on the third feature to obtain a fourth feature; the overlap rate of the overlap-add operation is 0.5; Processing the fourth feature through a ReLU activation function and a feed-forward network layer, using a nonlinear transformation to extract and learn the feature representation of each token vocabulary; When reaching the last layer of Transformer architecture, the SoftMax function is used to obtain the sampling probability of all digital sequences, and the next digital sequence is sampled according to the preset sampling algorithm; Repeat the steps of normalizing the high-dimensional vector corresponding to each layer of the Transformer architecture, and processing the normalized original features through a linear layer, to the step of processing the fourth feature through a ReLU activation function and a feedforward network layer, and using a nonlinear transformation to extract and learn the feature representation of each token vocabulary, until a preset number of cycles is reached, and then generating the index sequence.

5. The method according to claim 4, characterized in that The multi-scale Transformer using a trainable residual connection structure and adding a convolution module uses an intra-frame Transformer and an inter-frame Transformer to respectively learn the short-term and long-term dependencies of the first feature to obtain a second feature, including: Assigning a unique representation to each position of the first feature using absolute position coding to obtain an intermediate first feature; Performing a linear transformation of weights and biases on the intermediate first feature through a feedforward neural network, and introducing a nonlinear transformation through a ReLU function to obtain an intermediate second feature; Processing the intermediate second feature through a multi-head attention mechanism to obtain an intermediate third feature; Processing the intermediate third feature through a dual-scale convolution module to obtain an intermediate fourth feature and an intermediate fifth feature; Processing the intermediate fourth feature and the intermediate fifth feature through the Mish activation function to obtain an intermediate sixth feature and an intermediate seventh feature; The intermediate sixth feature and the intermediate seventh feature are spliced ​​to obtain an intermediate eighth feature; The intermediate eighth feature is subjected to adaptive average pooling to obtain an intermediate ninth feature, and the intermediate ninth feature is processed by the feedforward neural network to obtain the second feature.

6. The method according to claim 5, characterized in that The step of processing the intermediate ninth feature through the feedforward neural network to obtain the second feature includes: The intermediate ninth feature in the intra-frame Transformer and the inter-frame Transformer is added to the input of each feedforward neural network layer through a trainable residual connection and then subjected to a layer normalization operation to obtain the second feature.

7. The method according to claim 6, characterized in that The trainable residual connection is represented by the formula: y=F(W*x); Among them, x represents input, W represents the trainable weight matrix; F(x) represents the nonlinear mapping function from input to output, and y represents the output.

8. A training device for a large language model based on a trainable residual connection and a dual-scale convolutional Transformer, characterized in that: include: An acquisition module is used to acquire a basic model based on a multi-layer Transformer architecture, wherein the basic model includes multiple cascaded Transformer decoder layers, each of the Transformer decoder layers includes a self-attention mechanism and a feedforward neural network, each of the Transformer decoder layers is embedded with a dual-scale convolution module, and the dual-scale convolution modules of different scales are used to capture different local features. The outputs of the dual-scale convolution modules are fused as the outputs of the corresponding Transformer decoder layers, and a trainable weight matrix is ​​configured between the input and output of each Transformer decoder layer, and the trainable weight matrix is ​​used to adjust the strength of the residual connection between each Transformer decoder layer; obtain the preprocessed sample document data to build a training set; The training module is used to train the basic model based on the multi-layer Transformer architecture based on the training set until a preset training termination condition is reached, thereby obtaining a trained large language model that integrates trainable residual connections and dual-scale convolutional Transformers.

9. A computer device, characterized in that: The computer device comprises a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device executes the method according to any one of claims 1 to 7.

10. A readable storage medium, characterized in that: The readable storage medium includes a computer program, and when the computer program is executed, the computer device where the readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Text matching method based on enhanced pre-training text matching model

    CN117540009A

  • Extraction type text abstract generation method based on clause coding

    CN117875268A

  • Large robot model and training method and device thereof

    CN118568504A

  • Agricultural field large language model training method and device and medium

    CN119128070A

  • Training method of natural language processing model based on Transform architecture

    CN119129654A

Cited By

  • Regional meteorological model error control method based on neural network

    CN121327748A