Generation method and device of large language model, equipment and storage medium

Through phased training and parameter fusion methods, high-quality short text and medium- and low-quality long text data are used to train large language models, which solves the problems of complex, high cost and poor balance of capabilities in the existing technology medium- and long text training process, and achieves efficient and flexible long- and short text processing capabilities.

CN120336849APending Publication Date: 2025-07-18BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510402633.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The long text training method of existing large language models is complex and expensive, it is difficult to obtain high-quality long text data, poor balance of long and short text capabilities, insufficient universality of training strategies, and difficult to meet the needs of different task scenarios.

Method used

Through phased training and parameter fusion methods, the initial model is trained using high-quality short text and medium- and low-quality long text data, and the short text model and long text model are obtained respectively, and then parameter fusion is performed to generate a large language model.

Benefits of technology

It improves the processing ability of large language models for texts of different lengths, reduces training difficulty and cost, enhances the flexibility and applicability of the model, and can take into account both efficiency and applicability in multiple scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336849A_ABST
    Figure CN120336849A_ABST
Patent Text Reader

Abstract

The invention provides a large language model generation method and device, equipment and a storage medium, and relates to the technical field of artificial intelligence, in particular to the technical fields of large language models, model training, text processing and the like. According to the specific implementation scheme, an initial model is trained according to first training data, and a first model is obtained; wherein the first training data comprises a first type of text; training the initial model according to second training data to obtain a second model; wherein the second training data comprises the first type of text and the second type of text; and performing parameter fusion according to the first model and the second model to obtain the trained large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and particularly to the fields of large language models, model training, text processing, etc. Background Art

[0002] The long text ability of a model includes the abilities to process, analyze, generate, etc. long texts. A large language model (which can be abbreviated as large language model) can capture the context information in long texts and understand the semantic and logical relationships of the texts. A long text training method for a large language model can include pre-training on a large amount of long texts based on a short text model, and then performing multiple fine-tuning using short texts, long and short texts, etc. This training method has a complex process and requires huge computing resources. Summary of the Invention

[0003] This disclosure provides a method, apparatus, device, and storage medium for generating a large language model.

[0004] According to one aspect of this disclosure, there is provided a training method for a large language model, including:

[0005] Training an initial model according to first training data to obtain a first model; wherein, the first training data includes a first type of text;

[0006] Training the initial model according to second training data to obtain a second model; wherein, the second training data includes the first type of text and a second type of text;

[0007] Performing parameter fusion according to the first model and the second model to obtain the trained large language model.

[0008] According to another aspect of this disclosure, there is provided a method for generating a large language model, including:

[0009] Inputting a text to be processed into the large language model and outputting a generation result; wherein, the large language model is trained according to the above-mentioned training method for the large language model.

[0010] According to another aspect of this disclosure, there is provided a training apparatus for a large language model, including:

[0011] A first training module, configured to train an initial model according to first training data to obtain a first model; wherein, the first training data includes a first type of text;

[0012] A second training module, configured to train the initial model according to second training data to obtain a second model; wherein, the second training data includes the first type of text and a second type of text;

[0013] A fusion module, configured to perform parameter fusion based on the first model and the second model to obtain the trained large language model.

[0014] According to another aspect of the present disclosure, there is provided a generating device for a large language model, including:

[0015] A generating module, configured to input the text to be processed into the large language model and output a generation result; wherein, the large language model is trained by the training method of the large language model according to any embodiment of the present disclosure.

[0016] According to another aspect of the present disclosure, there is provided an electronic device, including:

[0017] At least one processor; and

[0018] A memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute any method in the embodiments of the present disclosure.

[0020] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any method in the embodiments of the present disclosure.

[0021] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, which when executed by a processor implements any method in the embodiments of the present disclosure.

[0022] According to another aspect of the present disclosure, there is provided a large language model, which is trained according to the training method of the large language model described above and is used to implement the generation method based on the large language model described above.

[0023] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0025] Figure 1 is a flowchart of a training method for a large language model according to an embodiment of the present disclosure;

[0026] Figure 2 is a flowchart of a generation method for a large language model according to an embodiment of the present disclosure;

[0027] Figure 3 It is a schematic flowchart of a method for fusing a short text model and a long text model according to an embodiment of the present disclosure;

[0028] Figure 4 It is a schematic structural diagram of a training device for a large language model according to an embodiment of the present disclosure;

[0029] Figure 5 It is a schematic structural diagram of a training device for a large language model according to another embodiment of the present disclosure;

[0030] Figure 6 It is a schematic structural diagram of a generation device for a large language model according to an embodiment of the present disclosure;

[0031] Figure 7 It is a block diagram of an electronic device for implementing the embodiments of the present disclosure. Specific Embodiments

[0032] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.

[0033] An example of a long text training method for a large language model may include the following steps:

[0034] 1. Obtain a base short text model (Base Model): This is a short window model. For example, the window of the model has a context length of 4k or 8k.

[0035] 2. Long context continue pretraining: Further pretrain the Base Model using a large amount of data, such as 500 million to 10 billion tokens, and train it on a larger context length such as 32k or 128k to obtain basic long text capabilities.

[0036] 3. General short instruction fine-tuning (Short SFT): Further fine-tune the model using a general short text instruction dataset so that the model has the ability to follow general instructions.

[0037] 4. Short-long mix instruction data fine-tuning (Short-Long Mix SFT): Combine short text and long text task instruction datasets for mixed fine-tuning to improve the model's ability to follow instructions in long text tasks.

[0038] However, there are some problems with the methods for expanding the long text capabilities of large language models:

[0039] 1. The training process is complex and costly: It requires going through a complex training process in multiple stages. Each stage requires a large amount of computing resources and multiple rounds of training, resulting in a high overall training cost and making it difficult to meet the needs of scenarios with limited computing power.

[0040] 2. It is difficult to obtain high-quality long text data: It is difficult to obtain high-quality long text data (Low-SFT) generally exceeding 32k. It is hard to ensure consistency and high quality under a large amount of data, and often relies on medium-quality / low-quality data for model training.

[0041] 3. The balance between long and short text capabilities is poor: After mixed training with medium-quality / low-quality long text data and high-quality short text data, although the long text task capabilities can be improved, it often leads to a significant decline in short text task performance, making it difficult to meet the needs of different task scenarios simultaneously.

[0042] 4. The lack of strategy generality: The general scalability of long text training strategies is poor. The lack of universality and scalability further increases the trial-and-error costs of development and deployment.

[0043] The embodiments of the present disclosure can improve the long text training of large models in one or more aspects such as data quality, computing power cost, ability balance, and generality.

[0044] Figure 1 FIG. 100 is a schematic flowchart of a training method for a large language model according to an embodiment of the present disclosure. In one implementation, the method may include:

[0045] S101. Train an initial model according to first training data to obtain a first model; wherein, the first training data includes a first type of text;

[0046] S102. Train the initial model according to second training data to obtain a second model; wherein, the second training data includes the first type of text and a second type of text;

[0047] S103. Perform parameter fusion according to the first model and the second model to obtain the trained large language model.

[0048] In the embodiments of the present disclosure, the initial model may include a base short text model (Base Model), such as a short window model with a context length of 4k or 8k. The initial model may have some general instruction processing capabilities.

[0049] In the embodiments of the present disclosure, the first training data may include multiple first-type texts. The first-type texts may include text instructions. Taking the first-type texts as short text instructions as an example, the first-type texts may include question-and-answer pairs, that is, a question part and an answer part. The question part of the first-type texts in the first training data may be input into the initial model to obtain an output result. Based on the answer part of the first-type texts, the output result, etc., the initial model may be fine-tuned to obtain a first model that can process short text instructions such as the first-type texts. A fine-tuning method may include: calculating a loss function based on the answer part and the output result. If the loss function does not converge, the parameters of the initial model are fine-tuned and then the first training data is continued to be used for training. The first-type texts included in the first training data used each time may be the same or different. If the loss function converges, the training may be stopped to obtain a trained first model. If the initial model is trained using short text instructions, it can adapt to the semantic compression and fast response requirements of short texts, and optimize the model's understanding and generation capabilities for short texts, such as sentiment analysis, keyword extraction, etc.

[0050] In the embodiments of the present disclosure, the second training data may include multiple first-type texts and multiple second-type texts. That is to say, the second training data is mixed data of different types of texts. The first-type texts in the second training data may be the same as those in the above-mentioned first training data. The second-type texts may include text instructions. Taking the second-type texts as long text instructions as an example, the second-type texts may include text contents such as novels, news, papers, etc. and their corresponding annotation contents. The question part of the first-type texts and the text contents of the second-type texts in the second training data may be input into the initial model to obtain an output result. Based on the answer part of the first-type texts, the annotation contents of the second-type texts, the output result, etc., the initial model may be fine-tuned to obtain a second model that can process second-type texts such as long text instructions. The fine-tuning method may refer to the relevant description of the loss function of the first model above. If the initial model is trained using a mixture of short and long text instructions, the large language model (which may be abbreviated as the large model) can learn long text capabilities, such as: the ability to process long texts, the ability to generate long texts, the ability to semantically condense long texts, the ability to jointly process complex contexts and multi-scale texts, etc.

[0051] In the embodiments of the present disclosure, the initial model trained using the first training data and the initial model trained using the second training data can be the same model. Before training, the structures and parameter values of the initial models can be the same. After training, the structures of the first model and the second model are the same, but the parameter values are different. The parameters of each layer of the first model and the second model can be fused layer by layer, or specific parameters of the first model and the second model can be fused to obtain new parameters. Parameter fusion can include methods of combining the parameters of the two models according to specific rules, such as weighted average, layer-by-layer interpolation, etc., which are not limited in the present disclosure. By parameter fusion, the weights of the new model can be generated to inherit the advantageous features of different models. For example, the corresponding layer parameters of the first model and the second model can be weighted and fused according to weights such as 0.3:0.7 to obtain a large language model.

[0052] According to the embodiments of the present disclosure, through staged training and parameter fusion, the trained large language model can balance efficiency and applicability in multiple scenarios, improving the processing ability of the large language model for more types of texts. For example, through staged training, gradient conflicts during mixed data training can be avoided; through the first training data, the response efficiency of the large language model can be optimized, and through the second training data, the text processing length of the large language model can be extended. Through parameter fusion, the advantages of the two models can be combined, reducing performance degradation caused by continued mixed training.

[0053] In one implementation, the length of the first type of text is less than the length of the second type of text; the first model is a short text model; the second model is a long text model.

[0054] In the embodiments of the present disclosure, the first type of text can include short text instructions. Short text instructions can include text segments with short lengths and high information density, such as text segments with a single theme or simple semantics. For example: social media comments, news headlines, question-and-answer pairs, etc. The second type of text can include long text instructions. Long text instructions can include paragraphs, chapters, books, etc., with complex logic and context relevance. The length of the first type of text can be much less than the length of the second type of text. The quantity of the first type of text can be much greater than that of the second type of text. The quality of the first type of text can be higher than that of the second type of text. For example, the first type of text includes 200,000 high-quality general short text supervised fine-tuning (SFT) data. The second type of text includes 20,000 medium-quality and / or low-quality long and short mixed SFT data.

[0055] In the embodiments of the present disclosure, by training an initial model with high-quality short text instructions, a short text model with the ability to follow short text instructions can be obtained. By training the initial model with a mixed text of high-quality short text instructions and medium- and low-quality long text instructions, a long text model with the ability to follow long context instructions can be obtained. By performing parameter fusion based on the short text model and the long text model, a trained large language model can be obtained.

[0056] According to the embodiments of the present disclosure, a short text model and a long text model can be trained based on the same model with different training data, and thus the fused large language model has strong short and long text capabilities.

[0057] In one implementation, the length of the second type of text is longer than the window length of the initial model; the window length of the initial model is the maximum text length that the initial model can process at one time.

[0058] In the embodiments of the present disclosure, in order to expand the processing ability of the large language model for texts of different lengths and enhance the processing ability of the large language model for texts longer than its own window length, the second type of text with a length longer than the window of the large language model can be added to the training data. For example, if the window length of the initial model is L, the window length range of the finally enhanced long text model can reach 4L - 32L (4 to 32 times the initial window) according to different computing powers, such as expanding from a 4096 window to a length of 32768 or 131072.

[0059] According to the embodiments of the present disclosure, by fine-tuning the initial large language model with the mixed data of the first type of text and the second type of text, the maximum text length that the obtained second model can process at one time will be increased compared with the initial model, which is beneficial to enhancing the processing ability of the large language model for texts of different lengths, especially for long texts.

[0060] In one implementation, the first model includes first parameters, and the second model includes second parameters; in S103, performing parameter fusion based on the first model and the second model to obtain the trained large language model further includes: obtaining the large language model according to the first parameters, the second parameters, and a fusion coefficient.

[0061] In the embodiments of the present disclosure, the parameters in the first model can be referred to as the first parameters, and the parameters in the second model can be referred to as the second parameters. The fusion coefficient of the first parameters and the second parameters can be set according to requirements. The fusion coefficient can also be referred to as the fusion weight, fusion ratio, etc. The fusion coefficient can represent the importance of the parameters of different models in the parameters of the finally fused large language model. The fusion coefficient can be a numerical value, a vector or a matrix, and can be specifically determined according to the structure of the initial model, the type of parameters, etc. For example, the fusion coefficient corresponding to all the first parameters in the first model is 0.4, and the fusion coefficient corresponding to all the second parameters in the second model is 0.6. For another example, in the first model, the fusion coefficient corresponding to the first parameters in the first layer is a, the fusion coefficient corresponding to the first parameters in the second layer is b, and the fusion coefficient corresponding to the first parameters in other layers is c; in the second model, the fusion coefficient corresponding to the first parameters in the first layer is 1 - a, the fusion coefficient corresponding to the first parameters in the second layer is 1 - b, and the fusion coefficient corresponding to the first parameters in other layers is 1 - c. In this example, the fusion coefficients of the first model and the second model can be represented by a vector. For another example, if different parameters in different layers of the model may be fused with different values, the fusion system can also be represented by a matrix related to the model parameters.

[0062] In the embodiments of the present disclosure, by adjusting the fusion system, the processing ability of the finally obtained large language model for different types of texts can be adjusted, so that the large language model focuses on different functions. For example, by adjusting the fusion coefficients of the short text model and the long text model, a large language model with both short text instruction following ability and long text instruction following ability can be obtained.

[0063] According to the embodiments of the present disclosure, by fusing the parameters of different models through the fusion coefficient, the processing ability of the large language model for various types of texts such as short and long texts can be optimized, the flexibility and generality of the large language model can be improved, and the training difficulty of the large language model can be reduced.

[0064] In one implementation manner, obtaining the large language model according to the first parameters, the second parameters and the fusion coefficient includes:

[0065] Calculating a weighted sum according to the first parameters, the first fusion coefficient, the second parameters and the second fusion coefficient to obtain the large language model; the first fusion coefficient represents the proportion of the first model in the large language model; the second fusion coefficient represents the proportion of the second model in the large language model.

[0066] In the embodiments of the present disclosure, the first fusion coefficient can be used as the weight of the first parameters, and the second fusion coefficient can be used as the weight of the second parameters. An example of a calculation method for parameter fusion is as follows:

[0067] Θ merged =λ shortΘ short + λ long Θ long

[0068] Among them, Θ short and Θ long represent the parameters of the first model and the second model respectively, and λ short and λ long are the fusion coefficients of the two models, indicating the proportions of the two models. The fused parameter Θ merged serves as the parameter of the final large language model.

[0069] According to the embodiments of the present disclosure, through weighted fusion calculation, a large language model with good processing capabilities for various texts such as long and short texts can be obtained, which can not only improve the applicability of the large language model but also enhance the flexibility of the large language model.

[0070] Figure 2 is a schematic flowchart of the generation method 200 of the large language model according to an embodiment of the present disclosure. In one implementation, the method may include:

[0071] S201. Input the text to be processed into the large language model and output the generation result; among them, the large language model is trained according to the training method of the large language model in any of the above embodiments.

[0072] In the embodiments of the present disclosure, the text to be processed can be directly input into the large language model without preprocessing. Multiple texts can be combined and input into the large language model as the text to be processed.

[0073] In the embodiments of the present disclosure, using the trained large language model, functions such as long text classification, long information retrieval, sentiment analysis, text analysis, summary generation, image generation, video generation, audio generation, and dialogue can be realized according to the input text to be processed.

[0074] According to the embodiments of the present disclosure, corresponding processing can be performed according to the input text to be processed to obtain a result that meets expectations.

[0075] In one implementation, the text to be processed includes the first type of text and / or the second type of text.

[0076] In the embodiments of the present disclosure, for the explanations and examples of the first type of text and / or the second type of text, reference can be made to the relevant descriptions in the above training method, which will not be elaborated here.

[0077] According to the embodiments of the present disclosure, the model can process different types of texts to be processed and has strong flexibility and applicability.

[0078] Figure 3A method for integrating a short text model and a long text model according to an embodiment of the present disclosure is as follows Figure 3 As shown, this method can efficiently expand the long text ability of the model based on model integration. The method may include the following steps:

[0079] S301. Short text model training: Train using a first quantity, such as 300,000 high-quality general short text SFT data (labeled as Short-SFT data), to obtain a first model. The first model has the ability to follow general instructions.

[0080] S302. Long text model training: Use a second quantity, such as 40,000 to 50,000 medium / low-quality long text SFT data (labeled as Long-SFT data) and 300,000 high-quality general short text SFT data for mixed training to obtain a second model. The second model can learn to handle long context instruction following tasks.

[0081] S303. Weight fusion: Fusion the weights of the two models obtained in steps S301 and S302. An example of the fusion formula is as follows:

[0082] Θ merged =λ short Θ short +λ long Θ long

[0083] Where Θ short and Θ long represent the weights of the short text model and the long text model respectively, and λ short and λ long are fusion coefficients representing the proportions of the two models. The fused weight Θ merged is used as the final long text model.

[0084] The method of the present disclosure can be applied to LLM long text task processing. The following are specific application examples:

[0085] 1. Long text classification: The enhanced long text model of the present disclosure can be used to classify longer content input by users, such as novels, news, etc.

[0086] 2. Long information retrieval. In a search engine, the query terms input by users are usually short texts, while the returned results may be long articles. The long text model enhanced by the algorithm of the present disclosure can optimize the relevance of search results and provide answers that better meet the user's needs.

[0087] 3. Sentiment analysis: In market research, consumers may leave short reviews or detailed feedback. The long text model enhanced by the algorithm of the present disclosure can capture the potential sentiment tendencies in long reviews while analyzing short reviews, thereby providing more comprehensive insights for enterprises.

[0088] Using the efficient long text expansion method proposed in this disclosure, a model with relatively excellent capabilities for both long and short texts can be obtained after fusion. For example, in one training result, the general short text ability of the short text model is 8.3, the general short text ability of the long text model is 6.94, and the general short text ability of the fusion model is 7.81. The general long text ability of the short text model is poor, the general long text ability of the long text model is good, and the general long text ability of the fusion model is close to that of the long text model and can, for example, process content of 128k.

[0089] Without the need for large-scale continue pre-training of the original short window context model, this disclosure can use low-quality Long-SFT corpus to perform supervised fine-tuning on the short text model, and then fuse it with a relatively excellent short text model to finally obtain a large language model with good capabilities for processing both long and short text tasks. This method is suitable for application scenarios that require efficient expansion of the long text ability of the model. Especially in scenarios with limited computing power and limited high-quality data annotation, it can better expand the long text task processing ability with relatively low computing resources and relatively low data quality.

[0090] The application scenarios of this disclosure include but are not limited to: expansion of the long text ability of large language models: tasks such as long conversations, long text retrieval and summarization.

[0091] According to this disclosure, it is possible to efficiently expand the long text task processing ability of the model with relatively low computing resources and relatively low data quality, while maximizing the performance of the existing short text model, reducing the overall system cost, and improving the practicability and popularity.

[0092] Figure 4 FIG. 15 is a schematic structural diagram of a training device 400 of a large language model according to an embodiment of this disclosure. In one implementation, the device may include:

[0093] A first training module 401, configured to train an initial model according to first training data to obtain a first model; wherein, the first training data includes a first type of text;

[0094] A second training module 402, configured to train the initial model according to second training data to obtain a second model; wherein, the second training data includes the first type of text and a second type of text;

[0095] A fusion module 403, configured to perform parameter fusion according to the first model and the second model to obtain the trained large language model.

[0096] In one embodiment, the length of the first type of text is less than the length of the second type of text; the first model is a short text model; the second model is a long text model.

[0097] In one embodiment, the length of the second type of text is longer than the window length of the initial model; the window length of the initial model is the maximum text length that the initial model can process at one time.

[0098] Figure 5 It is a schematic structural diagram of a training device 500 for a large language model according to another embodiment of the present disclosure. The device 500 may include: a first training module 501, a second training module 502, and a fusion module 503. The functions of the above modules can refer to the functions of the respective modules of the training device 400 for the large language model in the above embodiment. In one embodiment, the first model includes first parameters, and the second model includes second parameters; the fusion module 503 includes:

[0099] A calculation sub-module 5031, configured to obtain the large language model according to the first parameter, the second parameter, and the fusion coefficient.

[0100] In one embodiment, the calculation sub-module 5031 is further configured to calculate a weighted sum according to the first parameter, the first fusion coefficient, the second parameter, and the second fusion coefficient to obtain the large language model; the first fusion coefficient represents the proportion of the first model in the large language model; the second fusion coefficient represents the proportion of the second model in the large language model.

[0101] Figure 6 It is a schematic structural diagram of a generating device 600 for a large language model according to an embodiment of the present disclosure. In one embodiment, the device may include:

[0102] A generating module 601, configured to input the text to be processed into the large language model and output a generation result; wherein, the large language model is trained by the training device for any large language model in the above embodiment.

[0103] In one embodiment, the text to be processed includes the first type of text and / or the second type of text.

[0104] For the specific functions and examples of the respective modules and sub-modules of the device in the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be elaborated herein.

[0105] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0106] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0107] Figure 7 FIG. shows a schematic block diagram of an exemplary electronic device 700 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital assistant, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0108] As Figure 7 shown, the device 700 includes a computing unit 701 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0109] A plurality of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0110] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the training method of a large language model and / or the generation method of a large language model. For example, in some embodiments, the training method of a large language model and / or the generation method of a large language model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the training method of the large language model and / or the generation method of the large language model described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute the training method of the large language model and / or the generation method of the large language model in any other suitable manner (e.g., by means of firmware).

[0111] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0112] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program codes may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0113] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0114] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0115] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0116] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0117] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0118] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A training method for a large language model, comprising: Training an initial model according to first training data to obtain a first model; wherein, the first training data includes a first type of text; Training the initial model according to second training data to obtain a second model; wherein, the second training data includes the first type of text and a second type of text; Performing parameter fusion according to the first model and the second model to obtain the trained large language model.

2. The method according to claim 1, wherein, The length of the first type of text is less than the length of the second type of text; the first model is a short text model; the second model is a long text model.

3. The method according to claim 1 or 2, wherein The length of the second type of text is longer than the window length of the initial model; the window length of the initial model is the maximum text length that the initial model can process at one time.

4. The method according to any one of claims 1 to 3, wherein, The first model includes first parameters, and the second model includes second parameters; the performing parameter fusion according to the first model and the second model to obtain the trained large language model includes: Obtaining the large language model according to the first parameters, the second parameters, and a fusion coefficient.

5. The method according to claim 4, wherein, Obtaining the large language model according to the first parameters, the second parameters, and a fusion coefficient includes: Calculating a weighted sum according to the first parameters, a first fusion coefficient, the second parameters, and a second fusion coefficient to obtain the large language model; the first fusion coefficient represents the proportion of the first model in the large language model; the second fusion coefficient represents the proportion of the second model in the large language model.

6. A generation method for a large language model, comprising: Inputting a text to be processed into the large language model and outputting a generation result; wherein, the large language model is trained according to the training method of the large language model according to any one of claims 1 to 5.

7. The method according to claim 6, wherein, The text to be processed includes the first type of text and / or the second type of text.

8. A training device for a large language model, comprising: A first training module for training an initial model according to first training data to obtain a first model; wherein, the first training data includes a first type of text; A second training module for training the initial model according to second training data to obtain a second model; wherein, the second training data includes the first type of text and a second type of text; A fusion module for performing parameter fusion according to the first model and the second model to obtain the trained large language model.

9. A generation device for a large language model, comprising: A generation module for inputting a text to be processed into the large language model and outputting a generation result; wherein, the large language model is trained according to the training method of the large language model according to any one of claims 1 to 5.

10. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-5 or any one of claims 6-7.

11. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-5 or any one of claims 6-7.

12. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-5 or any one of claims 6-7.

13. A large language model trained according to the training method of the large language model according to any one of claims 1-5 for implementing the method according to any one of claims 6-7.