Pre-training method and system of large language model, interaction method and system and storage medium

Through a multi-stage pre-training method, a large language model is trained using a code and mathematics-focused dataset, which solves the problem of insufficient general model capabilities in existing technologies, improves the model's capabilities in code and mathematics, and achieves high efficiency and accuracy in application scenarios.

CN120706555APending Publication Date: 2025-09-26ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510805244.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The large language models in existing technologies only have general capabilities during the pre-training stage, lacking professionalism and reasoning capabilities, especially in terms of application capabilities in code and mathematics.

Method used

By obtaining a pre-training data set focusing on coding and mathematical abilities, the basic large language model is pre-trained in multiple stages to improve its coding and mathematical abilities respectively, and combined with general ability training to form a target large language model.

Benefits of technology

This achieves a simultaneous improvement in the reasoning and general capabilities of the target large language model, improves its effectiveness and accuracy in specific application scenarios, and enhances its professionalism in code and mathematics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706555A_ABST
    Figure CN120706555A_ABST
Patent Text Reader

Abstract

The invention provides a pre-training method, an interaction method, a system and a storage medium, and the method comprises the steps: obtaining a pre-training data set which is used for learning reasoning capability and universal capability, the reasoning capability comprises code capability and mathematical capability, and the pre-training data set comprises a first data set and a second data set, the first data set focuses on code learning ability, the second data set focuses on mathematics learning ability, multi-stage pre-training is performed on the basic large language model according to the pre-training data set to obtain a target large language model, and the fine-tuned target large language model is used for determining an output answer corresponding to the input question. The method achieves the synchronous improvement of the reasoning capability and the universal capability of the target large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of artificial intelligence technology, and in particular to a pre-training method, interaction method, system, and storage medium for a large language model. Background Art

[0002] With the development of artificial intelligence technology, large language models are applied to natural language processing.

[0003] The training of a large language model consists of two phases: pre-training and fine-tuning. In related technologies, the pre-training phase primarily helps the large language model learn general capabilities, while the fine-tuning phase primarily helps the large language model learn specific application capabilities corresponding to specific tasks in specific scenarios.

[0004] However, through the methods in the above-mentioned related arts, the large language model after the pre-training stage usually has only general capabilities.

[0005] It is worth noting that the content of the above-mentioned related technologies is only information known to the inventor personally, and does not mean that the above-mentioned information has entered the public domain before the filing date of this specification, nor does it mean that it can become the prior art of this specification. Summary of the Invention

[0006] This specification provides a pre-training method, an interactive method, a system, and a storage medium to avoid the above-mentioned technical problems.

[0007] In a first aspect, this specification provides a pre-training method for a large language model, comprising:

[0008] Obtaining a pre-training dataset, wherein the pre-training dataset is used to learn reasoning ability and general ability, the reasoning ability including coding ability and mathematical ability, the pre-training dataset including a first dataset and a second dataset, the first dataset focusing on learning coding ability, and the second dataset focusing on learning mathematical ability; and

[0009] Performing multi-stage pre-training on the basic large language model according to the pre-training dataset to obtain a target large language model;

[0010] Among them, the fine-tuned target large language model is used to determine the output answer corresponding to the input question.

[0011] In a second aspect, this specification provides an interaction method, including:

[0012] Get input questions;

[0013] An output answer corresponding to the input question is determined based on the fine-tuned target large language model, wherein the target large language model is obtained based on the pre-training method as described in the first aspect.

[0014] In a third aspect, this specification provides a pre-training system for a large language model, including:

[0015] at least one storage medium storing at least one instruction set for pre-training a large language model;

[0016] At least one processor is communicatively connected to the at least one storage medium, wherein when the at least one processor is running, it reads the at least one instruction set and executes the pre-training method as described in the first aspect according to the instructions of the at least one instruction set.

[0017] In a fourth aspect, this specification provides an interactive system, including:

[0018] at least one storage medium storing at least one instruction set for pre-training a large language model;

[0019] At least one processor is communicatively connected to the at least one storage medium, wherein when the at least one processor is running, it reads the at least one instruction set and executes the interaction method as described in the second aspect according to the instructions of the at least one instruction set.

[0020] In a fifth aspect, this specification provides a computer-readable non-transitory storage medium, wherein the computer-readable non-transitory storage medium stores at least one instruction set, and the at least one instruction set is executed by at least one processor to implement the pre-training method described in the first aspect or the interaction method described in the second aspect.

[0021] As can be seen from the above technical solutions, the pre-training method, interactive method, system and storage medium provided in this specification, in the pre-training method, by using the corresponding data sets of code and mathematics in a staged and focused manner for pre-training, effectively improves the reasoning ability of the target large language model (especially the ability of the code and mathematics dimensions). In addition, the general language understanding ability of the target large language model is guaranteed through general capabilities, thereby ultimately obtaining a target large language model that is both professional and general, that is, achieving the simultaneous improvement of the reasoning ability and general capabilities of the target large language model.

[0022] Other features of the pre-training method, interactive method, system, and storage medium provided in this specification are partially listed in the following description. The creative aspects of the pre-training method, interactive method, system, and storage medium provided in this specification can be fully explained by practicing or using the methods, devices, and combinations described in the following detailed examples. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of this specification, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0024] Figure 1 A schematic diagram of an application scenario of the large language model pre-training method provided in the embodiments of this specification;

[0025] Figure 2 A schematic diagram of the structure of a large language model pre-training system provided in an embodiment of this specification;

[0026] Figure 3 A flowchart of a method for pre-training a large language model according to one embodiment of this specification is provided;

[0027] Figure 4 A schematic diagram illustrating the principle of a large language model pre-training method provided in one embodiment of this specification;

[0028] Figure 5 A flowchart of a method for pre-training a large language model according to another embodiment of this specification;

[0029] Figure 6 A schematic diagram illustrating the principle of a large language model pre-training method provided in another embodiment of this specification;

[0030] Figure 7 A flowchart of the interaction method provided in the embodiments of this specification. DETAILED DESCRIPTION

[0031] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with this specification. Rather, they are merely examples of apparatus and methods consistent with certain aspects of this specification, as detailed in the appended claims.

[0032] It should be understood that the terms "including" and "having" and any variations thereof in the embodiments of this specification are intended to cover but not exclude inclusion. For example, a product or device that includes a series of components is not necessarily limited to those components explicitly listed, but may include other components that are not explicitly listed or are inherent to these products or devices.

[0033] In the embodiments of this specification, the term "and / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.

[0034] The term "plurality" in the embodiments of this specification refers to two or more than two, and other quantifiers are similar to it.

[0035] The terms "first," "second," "third," "base," "target," and the like in this specification are used to distinguish similar or similar objects or entities and are not necessarily intended to limit a particular order or precedence, unless otherwise indicated. It should be understood that the terms used in this manner are interchangeable where appropriate, for example, enabling implementation in an order other than that given in the illustrations or descriptions of the embodiments of this specification.

[0036] The term "unit / module" as used in this specification refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic or combination of hardware and / or software code that is capable of performing the functions associated with that element.

[0037] In order to avoid at least one of the technical problems mentioned in the above background technology, this specification proposes a technical conception that has been creatively worked on: in the pre-training stage of the large language model, a pre-training data set for learning reasoning ability and general ability can be obtained, so that a large language model (also called a base model) with both reasoning ability and general ability can be trained based on the pre-training data set. And the reasoning ability can include coding ability and mathematical ability. That is to say, in the pre-training stage, a multi-stage approach can be adopted to enable the large language model to learn coding ability, mathematical ability, and general ability.

[0038] For example, first, a pre-training dataset can be obtained. The pre-training dataset is used to learn reasoning and general abilities, including coding and mathematical abilities. For example, the pre-training dataset includes a first dataset and a second dataset, with the first dataset focusing on learning coding and the second dataset focusing on learning mathematical abilities.

[0039] Then, the basic large language model is pre-trained in multiple stages based on the first and second datasets. For example, the large language model is pre-trained in the first stage based on the first dataset, and then pre-trained in the second stage based on the results of the first stage pre-training based on the second dataset. This allows the basic network model to first learn coding capabilities based on the first dataset, and then learn mathematical capabilities based on the second dataset, thereby obtaining a large language model with both coding and mathematical capabilities.

[0040] It is understandable that the order of multi-stage pre-training can also be adjusted, such as using mathematical ability learning as the first stage of pre-training and coding ability learning as the second stage of pre-training.

[0041] Accordingly, after pre-training the large language model, the pre-trained large language model can be fine-tuned, and the fine-tuned large language model can be deployed in a specific application scenario so as to perform specific applications of the fine-tuned large language model.

[0042] The technical solution provided in this specification is based on the above-mentioned technical concept. Based on the description of the above-mentioned technical concept, it can be seen that in the technical solution provided in this specification, by pre-training the large language model from the two dimensions of reasoning ability and general ability, the reasoning ability and general ability of the pre-trained large language model can be balanced and improved. This can improve the effectiveness, accuracy and reliability of the fine-tuned large language model in specific application scenarios. It can also improve the generalization ability of the pre-trained large language model.

[0043] To facilitate readers' understanding of this manual, the application scenarios of this manual are now introduced.

[0044] The technical solutions provided in this specification are suitable for scenarios that require pre-training of large language models. For example, the technical solutions provided in this specification can be applied to scenarios such as intelligent programming assistants, online education platforms, internal enterprise knowledge management systems, and financial risk assessment systems.

[0045] Take the above scenario of the intelligent programming assistant as an example:

[0046] The technical solution provided in this specification can be used to pre-train a large language model. After fine-tuning the large language model, a large language model can be deployed in an intelligent programming assistant.

[0047] When a software developer encounters a challenge developing a new feature and needs to quickly understand a complex algorithm, they can ask the intelligent programming assistant questions such as "Explain the function of this code and point out possible performance bottlenecks."

[0048] Accordingly, the intelligent programming assistant can determine and output a response to the question based on the large language model deployed within it, such as "This code implements a variant of the quick sort algorithm, which mainly sorts an array recursively. However, in the worst case (when the array is already sorted), its time complexity degenerates to X. It is recommended to randomly select a benchmark value before each recursion to optimize performance."

[0049] Take the above scenario of online education platform as an example:

[0050] The technical solution provided in this specification can be used to pre-train a large language model. After fine-tuning the large language model, a large language model can be deployed on the online education platform.

[0051] When students encounter difficult-to-understand advanced math problems, they can submit questions to online education platforms, such as "Please provide a detailed answer to this calculus word problem: XXX."

[0052] Accordingly, the online education platform can determine and output a response corresponding to the question based on the large language model deployed within it, such as "First, take the derivative to get XXXX."

[0053] It is worth noting that the above examples are only used to illustrate the application scenarios to which the technical solutions of this specification can be applied, and should not be understood as limiting the application scenarios.

[0054] Figure 1 This is a schematic diagram of an application scenario of a large language model pre-training method (hereinafter referred to as a pre-training method) according to an embodiment of this specification. The pre-training method of this specification can be applied to Figure 1 The scene shown is 100. Figure 1 As shown, scenario 100 may include a target user 101 , a client 102 , a server 103 , and a network 104 .

[0055] The target user 101 may be the user who triggers pre-training of the large language model. For example, the target user 101 may perform a target operation on the client 102 to trigger pre-training of the large language model. In some embodiments, the target user 101 may upload a pre-training dataset on the client 102 to trigger pre-training of the large language model.

[0056] The client 102 may be an electronic device that provides interactive functions to the target user 101. For example, the client 102 may provide an interactive interface to the target user 101, and the target user 101 may perform interactive operations in the interactive page. In some embodiments, the client 102 executes the pre-training method described in this specification in response to detecting the target operation triggered by the target user 101. At this time, the client 102 may store data or instructions for executing the pre-training method described in this specification, and may execute or be used to execute the data or instructions. In some embodiments, the client 102 may include a hardware device with data information processing capabilities and the necessary programs required to drive the hardware device to operate, so as to execute the pre-training method described in this specification.

[0057] In some embodiments, client 102 may include a mobile device, a tablet computer, a laptop computer, a built-in device in a motor vehicle, or the like, or any combination thereof. In some embodiments, the mobile device may include a smart home device, a smart mobile device, a virtual reality device, an augmented reality device, or the like, or any combination thereof. In some embodiments, the smart home device may include a smart TV, a desktop computer, or the like, or any combination thereof. In some embodiments, the smart mobile device may include a smartphone, a personal digital assistant, a gaming device, a navigation device, or the like, or any combination thereof. In some embodiments, the built-in device in a motor vehicle may include an onboard computer, an onboard television, or the like. In some embodiments, client 102 may include an acquisition device for acquiring a pre-training dataset.

[0058] In some embodiments, client 102 may have one or more applications (APPs) installed. APPs can provide target user 101 with the ability and interface to interact with the outside world via network 104. APPs include, but are not limited to, web browser APPs, search APPs, chat APPs, shopping APPs, video APPs, financial management APPs, instant messaging tools, email clients, social networking platform software, and the like. In some embodiments, client 102 may have a target APP installed. The target APP can collect pre-training datasets for client 102.

[0059] like Figure 1 As shown, client 102 can be communicatively connected to server 103. Server 103 can be communicatively connected to one client 102 or to multiple clients 102. In some embodiments, client 102 can interact with server 103 via network 104 to receive or send messages, etc. For example, client 102 can interact with server 103 via network 104 to send a pre-training dataset to server 103.

[0060] The server 103 may be a server that provides various services. For example, the server 103 may be a cloud server or a local server. The server 103 may be connected to a client 102 and receive data sent by the client 102, or may be connected to multiple clients 102 and receive data sent by each client 102.

[0061] In some embodiments, the pre-training method described herein may be executed on server 103. In this case, server 103 may store data or instructions for executing the pre-training method described herein and may execute or be used to execute the data or instructions. Server 103 may include hardware devices with data information processing capabilities and the necessary programs required to drive the hardware devices.

[0062] The network 104 is a medium for providing a communication connection between the client 102 and the server 103. The network 104 can facilitate the exchange of information or data. Figure 1 As shown, the client 102 and the server 103 can be connected to the network 104 respectively, and transmit information or data to each other through the network 104.

[0063] In some embodiments, the network 104 can be any type of wired or wireless network, or a combination thereof. For example, the network 104 can include a cable network, a wired network, a fiber optic network, a telecommunications network, an intranet, the Internet, a local area network (LAN), a wide area network (WAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a public switched telephone network (PSTN), a Bluetooth™ network, a ZigBee™ network, a near field communication (NFC) network, or the like.

[0064] In some embodiments, network 104 may include one or more network access points. For example, network 104 may include a wired or wireless network access point, such as a base station or an Internet exchange point, through which one or more components of client 102 and server 103 can connect to network 104 to exchange data or information.

[0065] It is worth mentioning that Figure 1The number of clients 102, servers 103, and networks 104 in the example is merely illustrative. Any number of clients 102, servers 103, and networks 104 may be used as needed. Furthermore, the pre-training method provided herein may be executed entirely on the client 102, entirely on the server 103, or partially on the client 102 and partially on the server 103.

[0066] That is to say, Figure 1 and targeting Figure 1 The above description is only used to illustrate the application scenarios to which the pre-training method of this specification may be applicable, and should not be understood as limiting the application scenarios.

[0067] Figure 2 The figure shows a hardware structure diagram of a pre-training system 200 provided according to an embodiment of the present specification. The pre-training system 200 can execute the pre-training method described in this specification. The pre-training method is introduced in other parts of this specification. When the pre-training method is executed on the client 102, the pre-training system 200 can be the client 102. When the pre-training method is executed on the server 103, the pre-training system 200 can be the server 103. When the pre-training method is partially executed on the client 102 and partially executed on the server 103, the pre-training system 200 can be a system including the client 102 and the server 103.

[0068] like Figure 2 As shown, the pre-training system 200 may include at least one storage medium 203 and at least one processor 202. In some embodiments, the pre-training system 200 may further include a communication port 204 and an internal communication bus 201. The pre-training system 200 may further include an I / O component 205.

[0069] The internal communication bus 201 can connect different system components. For example, the internal communication bus 201 can connect the storage medium 203, the processor 202, the communication port 204 and the I / O component 205.

[0070] The I / O component 205 supports input / output between the pre-training system 200 and other components.

[0071] The communication port 204 is used for data communication between the pre-training system 200 and the outside world. For example, the communication port 204 can be used for data communication between the pre-training system 200 and the network 104. The communication port 204 can be a wired communication port or a wireless communication port.

[0072] The storage medium 203 may include a data storage device. The data storage device may be a non-transitory storage medium or a temporary storage medium. For example, the data storage device may include one or more of a disk 2031, a read-only storage medium (ROM) 2032, or a random access storage medium (RAM) 2033. The storage medium 203 also includes at least one instruction set stored in the data storage device. The instruction set includes computer program code, which may include programs, routines, objects, components, data structures, processes, modules, etc. that execute the pre-training method provided in this specification.

[0073] At least one processor 202 can be communicatively connected to at least one storage medium 203. At least one processor 202 is used to execute the at least one instruction set mentioned above. When the pre-training system 200 is running, at least one processor 202 reads the at least one instruction set and, according to the instructions of the at least one instruction set, executes the pre-training method provided in this specification. The processor 202 can perform all steps included in the pre-training method. The processor 202 can be in the form of one or more processors. In some embodiments, the processor 202 can include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application-specific integrated circuit (ASIC), an application-specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physical processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuit or processor capable of performing one or more functions, etc., or any combination thereof.

[0074] For illustrative purposes only, only one processor 202 is shown in the pre-training system 200 in the accompanying drawings. However, it should be noted that the pre-training system 200 in this specification may also include multiple processors. Therefore, the operations and / or method steps disclosed in this specification may be performed by one processor or jointly by multiple processors. For example, if the processor 202 of the pre-training system 200 is described in this specification as performing steps A and B, it should be understood that steps A and B may also be performed jointly or separately by two different processors 202 (for example, the first processor performs step A, the second processor performs step B, or the first and second processors perform steps A and B together).

[0075] See also Figure 3 , Figure 3 This is a flow chart of a method for pre-training a large language model according to one embodiment of this specification. Figure 3The execution subject of the pre-training method shown can be a pre-training system. For the description of the pre-training system, please refer to the above example and will not be repeated here.

[0076] like Figure 3 As shown, the method includes the following S301 and S302:

[0077] S301: Obtain a pre-training dataset. The pre-training dataset is used to learn reasoning ability and general ability. Reasoning ability includes coding ability and mathematical ability. The pre-training dataset includes a first dataset and a second dataset. The first dataset focuses on learning coding ability, while the second dataset focuses on learning mathematical ability.

[0078] A pre-training dataset can be understood as a data set used to pre-train a basic large language model. This dataset primarily helps the basic large language model learn reasoning capabilities in the inference dimension and general capabilities in the general dimension. In other words, a pre-training dataset is a data set that helps the basic large language model learn both reasoning and general capabilities.

[0079] Reasoning ability can be understood as the model's ability to perform logical deduction, problem modeling, and problem solving based on existing information, covering tasks such as mathematical reasoning, code comprehension, and multi-step logical judgment.

[0080] General capabilities can be understood as the model's ability to adapt widely to various natural language processing tasks, such as text understanding, question answering, summary generation, common sense reasoning, etc., and are not limited to specific fields.

[0081] Coding capabilities can be understood as the model's ability to understand and generate program code, including syntax analysis, function calls, debugging suggestions, etc.

[0082] Accordingly, the first dataset used to enable the basic large language model to learn coding skills can be understood as a high-quality code corpus for improving coding skills, such as questions and answers related to coding skills.

[0083] Mathematical ability can be understood as the model's ability to solve mathematical problems, such as algebraic operations, geometric proofs, probability calculations, and calculus applications.

[0084] Accordingly, the second dataset used to enable the basic large language model to learn mathematical skills can be understood as a dataset for improving mathematical skills, such as including mathematics teaching materials, work time database, problem-solving procedures, and mathematics competition questions.

[0085] In addition, the first dataset can also enable the basic large language model to learn general skills and / or mathematical skills. However, relatively speaking, the first dataset allows the basic large language model to learn more about coding skills. Similarly, the second dataset can also enable the basic large language model to learn general skills and / or coding skills. However, relatively speaking, the second dataset allows the basic large language model to learn more about mathematical skills.

[0086] In other words, relatively speaking, the first dataset focuses on coding skills, emphasizing syntax, control flow, function structure, modular programming, etc. The second dataset focuses on mathematical skills, emphasizing symbolic manipulation, multi-step derivation, formula transformation, theorem application, etc.

[0087] The following example can be used to obtain a pre-training dataset:

[0088] In one example, the pre-training system may be connected to a data acquisition device to receive a pre-training data set collected and sent by the data acquisition device.

[0089] In another example, the pre-training system may provide a data loading tool, and the user may transfer the pre-training dataset to the pre-training system through the data loading tool.

[0090] Among them, the tool for loading data can be an interface for connecting to an external device, such as an interface for connecting to other storage devices, through which the pre-training data set transmitted by the external device is obtained; the tool for loading data can also be a display device, such as the pre-training system can output an interface for the data loading function on the display device, and the user can import the pre-training data set into the pre-training system through this interface.

[0091] In some embodiments, the pre-training data set includes mathematical ability training samples, coding ability training samples, and general ability training samples; the first data set and the second data set both include general ability training samples, and the coding ability training samples in the first data set account for the highest proportion, and the mathematical ability training samples in the second data set account for the highest proportion.

[0092] For example, Figure 4 As shown, the pre-training dataset includes math ability training samples, coding ability training samples, and general ability training samples. The math ability training samples are used to enable the basic large language model to learn math ability; the coding ability training samples are used to enable the basic large language model to learn coding ability; and the general ability training samples are used to enable the basic large language model to learn general ability.

[0093] Continue reading Figure 4The first sample data set includes general ability training samples and coding ability training samples, and may also include mathematical ability training samples. However, relatively speaking, in the first sample data set, the coding ability training samples account for the largest proportion of training samples.

[0094] The second sample data set includes general ability training samples and mathematical ability training samples, and may also include coding ability training samples. However, relatively speaking, in the second sample data set, the mathematical ability training samples account for the largest proportion of training samples.

[0095] It is worth noting that the general ability training samples in the first sample dataset and the second sample dataset may be the same or different, and this embodiment does not impose any limitation. Similarly, if the first dataset also includes mathematical ability training samples, the mathematical ability training samples in the second dataset may include at least some of the mathematical ability training samples in the first dataset; if the second dataset also includes coding ability training samples, the coding ability training samples in the first dataset may include at least some of the coding ability training samples in the second dataset.

[0096] Combined with the above analysis, it can be seen that taking the first sample data set as an example, if the first sample data set includes general ability training samples and coding ability training samples, the basic large language model can learn both coding ability and general ability; if the first sample data set includes general ability training samples, coding ability training samples, and mathematical ability training samples, the basic large language model can learn both coding ability, general ability, and mathematical ability.

[0097] It is understood that for the content that is the same or similar to the above content, it will not be repeated again to avoid tedious description. For example, the description of the second sample data set can refer to the description of the first sample data set above, and will not be repeated here.

[0098] Based on the above analysis, we can see that the basic large language model has different emphases when learning based on the first and second sample datasets. The first sample dataset focuses on learning coding skills, while the second sample dataset focuses on learning mathematical skills. Therefore, the basic large language model can focus on learning corresponding skills, so that the pre-trained large language model has coding learning capabilities, general learning capabilities, and mathematical capabilities. This improves the effectiveness and reliability of the pre-trained large language model in subsequent fine-tuning and application.

[0099] S302: Perform multi-stage pre-training on the basic large language model based on the pre-training dataset to obtain a target large language model. The fine-tuned target large language model is used to determine the output answer corresponding to the input question.

[0100] Multi-stage pre-training can be understood as dividing the entire pre-training process into multiple stages, using different data sets for training in each stage to gradually enhance the capabilities of the basic large language model in different aspects.

[0101] Exemplarily, multi-stage pre-training can be a two-stage pre-training. In the first stage, the pre-training system can be performed based on the first dataset to enhance the capabilities of the base large language model in the code dimension. This means that the target large language model has stronger code capabilities. In the second stage, the pre-training system can be performed based on the second dataset to enhance the capabilities of the base large language model in the mathematical dimension. This means that the target large language model has stronger mathematical capabilities. The reverse is also true.

[0102] For example, in some embodiments, S302 may include the following two implementations:

[0103] Implementation method 1: Perform the first stage pre-training on the basic large language model according to the first data set to obtain the coding ability large language model, and perform the second stage pre-training on the coding ability large language model according to the second data set to obtain the target large language model.

[0104] Combined with the above analysis, taking the case where the first data set includes general ability training samples and coding ability training samples, and the second data set includes general ability training samples and mathematical ability training samples as an example:

[0105] First, the pre-training system can perform the first stage of pre-training on the basic large language model based on the general ability training samples and code ability training samples in the first data set, so that the large language model after the first stage of pre-training (such as the code ability large language model) has both general ability and code ability, and the code ability is relatively strong.

[0106] Then, the pre-training system can perform a second-stage pre-training on the large language model that has undergone the first-stage pre-training based on the general ability training samples and mathematical ability training samples in the second data set, so that the large language model that has undergone the second-stage pre-training (such as the target large language model) has both general ability, coding ability, and mathematical ability, and all three abilities are relatively strong.

[0107] In other words, in this implementation, the pre-training system first pre-trains the base large language model from the code dimension to give the target large language model stronger coding capabilities; then pre-trains the base large language model from the mathematics dimension to give the target large language model stronger mathematics capabilities. Ultimately, the target large language model is obtained, which has strong general capabilities, strong coding capabilities, and strong mathematics capabilities.

[0108] Implementation method 2: Perform a first-stage pre-training on the basic large language model based on the second data set to obtain a mathematical ability large language model, and perform a second-stage pre-training on the mathematical ability large language model based on the first data set to obtain a target large language model.

[0109] Combined with the above analysis, taking the case where the first data set includes general ability training samples and coding ability training samples, and the second data set includes general ability training samples and mathematical ability training samples as an example:

[0110] First, the pre-training system can perform the first stage of pre-training on the basic large language model based on the general ability training samples and mathematical ability training samples in the second data set, so that the large language model after the first stage of pre-training (such as the mathematical ability large language model) has both general ability and coding ability, and the mathematical ability is relatively strong.

[0111] Then, the pre-training system can perform a second-stage pre-training on the large language model that has undergone the first-stage pre-training based on the general ability training samples and code ability training samples in the first data set, so that the large language model that has undergone the second-stage pre-training (such as the target large language model) has both general ability, code ability, and mathematical ability, and all three abilities are relatively strong.

[0112] In other words, in this implementation, the pre-training system first pre-trains the base large language model from a mathematical perspective, giving the target large language model stronger mathematical capabilities. It then pre-trains the base large language model from a coding perspective, giving the target large language model stronger coding capabilities. Ultimately, the target large language model is obtained, possessing both strong general capabilities, strong coding capabilities, and strong mathematical capabilities.

[0113] Combining the above implementation methods 1 and 2, it can be seen that in this embodiment, the pre-training system can pre-train the basic large language model from the mathematical ability dimension and the coding ability dimension, or can pre-train the basic large language model from the coding ability dimension and the mathematical ability dimension. The flexibility and diversity of multi-stage pre-training of the basic large language model can be improved. Moreover, the target large language model obtained through multi-stage pre-training can also have strong general capabilities, strong coding capabilities, and strong mathematical capabilities. That is, the balance and improvement of the reasoning ability and general capabilities of the target large language model are achieved.

[0114] Correspondingly, after pre-training is completed and the target large language model is obtained, the pre-training system or other systems can fine-tune the target large language model in a specific task scenario based on the specific application scenario to obtain a fine-tuned large language model.

[0115] At this point, the fine-tuned large language model can be deployed in specific application scenarios to complete specific tasks within those scenarios. For details, please refer to the description of the application scenarios above and will not be repeated here.

[0116] Combined with the above analysis of S301 and S302, it can be seen that in this embodiment, the pre-training system effectively improves the reasoning ability of the target large language model (especially the ability of the code and mathematics dimensions) by pre-training in a phased and focused manner using the corresponding data sets of code and mathematics. In addition, the general language understanding ability of the target large language model is guaranteed by the general ability, thereby ultimately obtaining a target large language model that is both professional and general, that is, achieving a simultaneous improvement in the reasoning ability and general ability of the target large language model.

[0117] Based on the above analysis, it can be seen that in this specification, the pre-training system uses a multi-stage approach to pre-train the basic large language model, so that the target large language model after multi-stage pre-training has strong general capabilities, strong coding capabilities, and strong mathematical capabilities. This embodiment does not limit the order of the multiple stages.

[0118] To facilitate readers' understanding of the pre-training method provided in this specification, the pre-training method provided in this specification is now described in detail in combination with 5. Figure 5 This is a flow chart of a method for pre-training a large language model according to another embodiment of this specification. Figure 5 As shown, the method includes the following S501 and S502:

[0119] S501: Obtain a pre-training dataset. The pre-training dataset is used to learn reasoning and general abilities, including coding and mathematical abilities. The pre-training dataset includes a first dataset, a second dataset, and a third dataset. The first dataset focuses on learning coding, the second dataset focuses on learning mathematical abilities, and the third dataset focuses on learning general abilities.

[0120] Similarly, the content of S501 that has been explained above will not be repeated here.

[0121] In addition, combined with the above description of the first data set and the second data set, it can be seen that the third data set used to enable the basic large language model to learn general capabilities can be understood as a database for improving general capabilities.

[0122] Similarly, in some embodiments, the third dataset can also enable the basic large language model to learn coding skills and / or mathematical skills. However, relatively speaking, through the third dataset, the basic large language model learns more general skills.

[0123] In some embodiments, the third data set includes mathematical ability training samples, coding ability training samples, and general ability training samples, and the general ability training samples in the third data set account for the highest proportion.

[0124] For example, based on the above analysis, the first, second, and third datasets include coding ability training samples, mathematical ability training samples, and general ability training samples, respectively. Training samples of the same type (e.g., coding ability type, mathematical ability type, general ability type) in different datasets may be the same or different.

[0125] Moreover, relatively speaking, the first dataset has the largest number of coding ability training samples; the second dataset has the largest number of mathematical ability training samples; and the third dataset has the largest number of general ability training samples. This allows the basic large language model to learn as many general, coding, and mathematical abilities as possible. This achieves a simultaneous improvement in the target large language model's reasoning and general abilities.

[0126] In some embodiments, in the first data set and the second data set, the mathematical ability training samples and the general ability training samples account for the same proportion; in the third data set, the coding ability training samples and the mathematical ability training samples account for the same proportion.

[0127] For example, in the first dataset, the number of coding skills training samples is the largest, while the number of math skills training samples and general skills training samples is the same. Therefore, the basic large language model can learn strong coding skills and some math skills and general skills.

[0128] In the second dataset, the number of training samples for math skills is the largest, while the number of training samples for coding skills and general skills is equal. Therefore, the basic large language model can learn strong math skills and some coding and general skills.

[0129] In the third dataset, the number of general ability training samples is the largest, while the number of coding ability training samples and mathematical ability training samples is the same. Therefore, the basic large language model can learn strong general abilities and can also learn certain coding and mathematical abilities.

[0130] In other words, with phased pre-training, each dataset not only focuses on its intended learning objective but also cultivates other capabilities. For example, the first dataset not only focuses on learning coding skills, but also on mathematical and general skills. This improves the balance and enhancement of the target large language model's reasoning and general capabilities. Furthermore, because pre-training is phased, learning objectives focused on in the early stages can be continued in the later stages, preventing knowledge omissions. This improves the effectiveness and reliability of pre-training.

[0131] S502: Perform multi-stage pre-training on the basic large language model based on the first data set, the second data set, and the third data set to obtain a target large language model. The fine-tuned target large language model is used to determine an output answer corresponding to the input question.

[0132] For example, the multiple stages may be three stages, and different stages may be implemented using different data sets. For example, the first stage may be implemented based on the first data set, the second data set, or the third data set, which is not limited in this embodiment.

[0133] For example, S502 may include the following six examples:

[0134] Example 1: Multi-stage pre-training is performed on the basic large language model according to the first data set, the second data set, and the third data set to obtain the target large language model.

[0135] For example, Figure 6 As shown, the multi-stage training can be three stages. The first stage is for the base large language model to learn coding skills based on the first data set; the second stage is for learning mathematical skills based on the second data set based on the first stage; and the third stage is for learning general skills based on the third data set based on the second stage. Thus, the target large language model is obtained after multi-stage pre-training.

[0136] Continuing with the above example and Figure 6 The first dataset includes: 60% coding ability training samples, 20% mathematics ability training samples, and 20% general ability training samples. The second dataset includes: 60% mathematics ability training samples, 20% coding ability training samples, and 20% general ability training samples. The third dataset includes: 60% general ability training samples, 20% coding ability training samples, and 20% mathematics ability training samples.

[0137] Continue reading Figure 6In the first phase, the basic large language model is pre-trained based on the first dataset (60% coding ability training samples, 20% math ability training samples, and 20% general ability training samples). This ensures that the large language model obtained through the first phase of pre-training has strong coding ability and certain math and general abilities.

[0138] Continue reading Figure 6 The second phase is carried out on the basis of the first phase. In the second phase, the large language model obtained through the first phase pre-training is pre-trained for the second phase based on the second data set (60% of the mathematical ability training samples, 20% of the coding ability training samples, and 20% of the general ability training samples). This is so that the large language model obtained through the second phase pre-training has stronger mathematical ability, further improves the strong coding ability (at least retains the strong coding ability obtained through the first phase pre-training), and further improves the general ability (at least retains a certain degree of general ability obtained through the first phase pre-training).

[0139] Continue reading Figure 6 , the third stage is carried out on the basis of the second stage. In the third stage, the large language model obtained by pre-training in the second stage is pre-trained in the third stage based on the third data set (60% general ability training samples, 20% coding ability training samples, and 20% mathematical ability training samples). So that the large language model obtained by pre-training in the third stage has stronger general ability, so as to re-enhance the ability of the target large language model to a wide range of language tasks (such as common sense, knowledge question and answer, text understanding, etc.), and avoid loss of generality due to over-specialization. In addition, the stronger coding ability can be further improved (at least the strong coding ability obtained by the second stage pre-training is retained), and the stronger mathematical ability can be further improved (at least the strong mathematical ability obtained by the second stage pre-training is retained).

[0140] Example 2: Multi-stage pre-training is performed on the basic large language model according to the second data set, the first data set, and the third data set to obtain the target large language model.

[0141] Accordingly, Example 2 can be understood as follows: the first stage is for the base large language model to learn mathematical skills based on the second dataset; the second stage is for the base large language model to learn coding skills based on the first dataset, building on the foundation of the first stage; and the third stage is for the base large language model to learn general skills based on the third dataset, building on the foundation of the second stage. This results in a target large language model that has undergone multi-stage pre-training.

[0142] Example 3: Multi-stage pre-training is performed on the basic large language model according to the third data set, the first data set, and the second data set to obtain the target large language model.

[0143] Accordingly, Example 3 can be understood as follows: the first stage is for the base large language model to learn general capabilities based on the third dataset; the second stage is for the base large language model to learn coding capabilities based on the first dataset, building on the first stage; and the third stage is for the base large language model to learn mathematical capabilities based on the second dataset, building on the second stage. This results in a target large language model that has undergone multi-stage pre-training.

[0144] Example 4: Multi-stage pre-training is performed on the basic large language model according to the third data set, the second data set, and the first data set to obtain a target large language model.

[0145] Accordingly, Example 4 can be understood as follows: the first stage is for the base large language model to learn general capabilities based on the third dataset; the second stage is for the base large language model to learn mathematical capabilities based on the second dataset, building on the foundation of the first stage; and the third stage is for the base large language model to learn coding capabilities based on the first dataset, building on the foundation of the second stage. This results in a target large language model that has undergone multi-stage pre-training.

[0146] Example 5: Multi-stage pre-training is performed on the basic large language model according to the first data set, the third data set, and the second data set to obtain the target large language model.

[0147] Accordingly, Example 5 can be understood as follows: the first stage is for the base large language model to learn coding skills based on the first dataset; the second stage is for the base large language model to learn general skills based on the third dataset, building on the first stage; and the third stage is for the base large language model to learn mathematical skills based on the second dataset, building on the second stage. This results in a target large language model that has undergone multi-stage pre-training.

[0148] Example 6: Multi-stage pre-training is performed on the basic large language model according to the second data set, the third data set, and the first data set to obtain the target large language model.

[0149] Accordingly, Example 6 can be understood as follows: the first stage is for the base large language model to learn mathematical skills based on the second dataset; the second stage is for the base large language model to learn general skills based on the third dataset, building on the foundation of the first stage; and the third stage is for the base large language model to learn coding skills based on the first dataset, building on the foundation of the second stage. This results in a target large language model that has undergone multi-stage pre-training.

[0150] Combined with the analysis of Examples 1 to 6 above, it can be seen that in this embodiment, multi-stage pre-training can be performed in different ways to achieve diversity and flexibility. Furthermore, through the multi-stage pre-training described in the above examples, a target large language model with both reasoning capabilities and general capabilities can be obtained. This improves the wide applicability of the target large language model.

[0151] According to another aspect of this specification, this specification also provides an interactive method, which can also be called an application method for a target large language model.

[0152] See also Figure 7 , Figure 7 This is a flow chart of the interactive method provided in the embodiment of this specification. Figure 7 As shown, the method includes the following S701 and S702:

[0153] S701: Obtain input question.

[0154] Similarly, regarding the features that are the same or similar between this embodiment and the above examples, this embodiment will not be repeated. For example, regarding the application scenarios of this embodiment, please refer to the description of the above examples. For another example, the execution subject of this embodiment can be an interactive system, and the interactive system can be the same system as the pre-training system, or it can be a different system. For another example, regarding the way in which the interactive system obtains input questions, please refer to the description of obtaining the pre-training data set in the above examples, or refer to the description of the application scenarios in the above examples.

[0155] S702: Determine an output answer corresponding to the input question based on the fine-tuned target large language model, wherein the target large language model is obtained based on the pre-training method described in the above embodiment.

[0156] Combined with the above analysis, it can be seen that after obtaining the target large language model through the pre-training method described in the above example, the pre-training system or interactive system can fine-tune the target large language model based on the specific application scenario to obtain a fine-tuned large language model.

[0157] Correspondingly, when there is an interaction requirement, such as when the interactive system receives an input question, the output answer can be determined based on the fine-tuned large language model.

[0158] Based on the above analysis, we can see that the target large language model has strong reasoning and general capabilities. Therefore, the fine-tuned large language model also has strong reasoning and general capabilities. Therefore, the output answers determined based on the fine-tuned large language model are highly accurate, reliable, and effective.

[0159] In this specification, the Large Language Model (LLM) may also be referred to as the large model. The large language model is a natural language processing model based on deep learning technology. Its parameter scale usually reaches billions to hundreds of billions or even higher, and it has powerful language understanding and generation capabilities. The large language model can adopt the Transformer architecture or its variants (such as GPT, BERT, etc.). This architecture uses the attention mechanism to achieve global modeling of sequence data, which can efficiently handle long-distance dependencies, thereby performing well in natural language tasks. The large language model learns the statistical characteristics and semantic relevance of language by pre-training on a large-scale corpus, giving it excellent generalization capabilities. The core capabilities of the large language model include but are not limited to: understanding contextual semantics, generating coherent and grammatically correct text, performing logical reasoning, and handling multi-task scenarios. Its usage methods generally include two modes: direct inference and fine-tuning. In direct inference mode, the user guides the large language model to generate specific outputs by designing prompts. The prompts can be textual task descriptions or instructions, stimulating the semantic understanding and generation capabilities of the large language model. In fine-tuning mode, the large language model is further trained on a small-scale dataset in a specific domain to optimize its performance on a specific task. The powerful generalization and flexibility of large language models make them a valuable tool in the field of artificial intelligence technology, providing efficient and accurate solutions for automated text generation and comprehension.

[0160] In some embodiments, the large language model may also have the ability to understand and generate data from other modalities (such as vision, audio, etc.). In this case, the large language model may also be called a multimodal large language model (MLLMs). MLLMs provide a richer and more natural interactive experience by integrating multiple types of inputs and outputs such as text, images, and sounds. The core advantage of MLLMs is that they can process and understand information from different modalities and fuse this information to complete complex tasks. For example, MLLMs can analyze a picture and generate descriptive text, or generate corresponding images based on the text description. This cross-modal understanding and generation capability gives MLLMs broad application prospects in multiple fields.

[0161] It should be noted that the key technologies of large language models can be found in the detailed description in the paper "ASurvey of LargeLanguage Models" (paper number: arXiv:2303.18223v16, public date: March 11, 2025, public link: https: / / doi.org / 10.48550 / arXiv.2303.18223), which will not be repeated in this manual.

[0162] It is worth noting that the above examples are only used to illustrate the possible implementation methods of the pre-training method and the interactive method of this specification, and should not be understood as limiting the implementation methods of the pre-training method and the interactive method of this specification. For example, based on the above technical concept, some of the above technical features can be combined to obtain a new embodiment; new technical features can be added to the above examples to obtain a new embodiment; some technical features can be reduced to obtain a new embodiment based on the above examples; some technical features in the above examples can be replaced with other technical features; some technical features and their order in the above examples can be adjusted to obtain a new embodiment, etc., which will not be listed here one by one.

[0163] Based on the above technical concept, this specification also provides a computer-readable non-transitory storage medium, which stores at least one instruction set. When the at least one instruction set is executed by the processor, the steps of the pre-training method and the interaction method described in this specification are implemented.

[0164] In some possible implementations, various aspects of this specification may also be implemented in the form of a program product that includes program code. Taking the pre-training method as an example, when the program product is run on the pre-training system 200, the program code is used to cause the pre-training system 200 to perform the steps of the pre-training method described in this specification. The program product for implementing the above method may include program code using a portable compact disc read-only memory (CD-ROM) and may be run on the pre-training system 200. However, the program product of this specification is not limited to this. In this specification, a readable storage medium may be any tangible medium that contains or stores a program that can be used by or in combination with an instruction execution system. The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples of computer-readable storage media include: an electrical connection having one or more conductors, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. The computer-readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The computer-readable storage medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing. Program code for performing the operations described herein may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and conventional procedural programming languages ​​such as C or similar programming languages. The program code may be executed entirely on the pre-training system 200, partially on the pre-training system 200, as a stand-alone software package, partially on the pre-training system 200 and partially on a remote pre-training system, or entirely on the remote pre-training system 200.

[0165] It should be noted that the collection, storage, use, processing, transmission, provision and disclosure of user-related information (such as input questions, etc.) involved in the technical solution of this specification are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0166] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0167] In summary, after reading this detailed disclosure, those skilled in the art will appreciate that the foregoing detailed disclosure may be presented by way of example only and may not be limiting. Although not expressly stated herein, those skilled in the art will understand that this specification encompasses various reasonable changes, improvements, and modifications to the embodiments. Such changes, improvements, and modifications are intended to be suggested by this specification and are within the spirit and scope of the exemplary embodiments of this specification.

[0168] Furthermore, certain terms in this specification have been used to describe embodiments of this specification. For example, “one embodiment,” “an embodiment,” and / or “some embodiments” mean that a particular feature, structure, or characteristic described in connection with that embodiment may be included in at least one embodiment of this specification. Therefore, it is emphasized and should be understood that two or more references to “an embodiment,” “one embodiment,” or “an alternative embodiment” in various parts of this specification do not necessarily refer to the same embodiment. Furthermore, particular features, structures, or characteristics may be appropriately combined in one or more embodiments of this specification.

[0169] It should be understood that in the foregoing descriptions of the embodiments of this specification, to facilitate understanding of a feature and to simplify this specification, various features are combined in a single embodiment, figure, or description thereof. However, this does not necessarily mean that these features are combined. When reading this specification, a person skilled in the art may label some of the devices as separate embodiments. In other words, the embodiments of this specification can also be understood as the integration of multiple sub-embodiments. The content of each sub-embodiment is also valid even when it includes fewer than all the features of a single previously disclosed embodiment.

[0170] Each patent, patent application, patent application publication, and other materials, such as articles, books, specifications, publications, documents, and the like, cited herein (excluding any historical review documents related thereto) is hereby incorporated by reference for all purposes relevant to this document, such as within the specification and claims herein. However, if there is any inconsistency or conflict between the descriptions, definitions, and / or terminology of such materials and the descriptions, definitions, and / or terminology used herein, the descriptions, definitions, and / or terminology used herein shall control.

[0171] Finally, it should be understood that the embodiments of the application disclosed herein are illustrative of the principles of the embodiments of this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can adopt alternative configurations based on the embodiments in this specification to implement the application in this specification. Therefore, the embodiments of this specification are not limited to the embodiments precisely described in the application.

Claims

1. A pre-training method for a large language model, comprising: Obtaining a pre-training dataset, wherein the pre-training dataset is used to learn reasoning ability and general ability, the reasoning ability including coding ability and mathematical ability, the pre-training dataset including a first dataset and a second dataset, the first dataset focusing on learning coding ability, and the second dataset focusing on learning mathematical ability; and Performing multi-stage pre-training on the basic large language model according to the pre-training dataset to obtain a target large language model; Among them, the fine-tuned target large language model is used to determine the output answer corresponding to the input question.

2. The method according to claim 1, wherein The pre-training data set includes mathematical ability training samples, coding ability training samples, and general ability training samples; the first data set and the second data set both include general ability training samples, and the coding ability training samples in the first data set account for the highest proportion, and the mathematical ability training samples in the second data set account for the highest proportion.

3. The method according to claim 1 or 2, wherein: The method of performing multi-stage pre-training on the basic large language model according to the pre-training dataset to obtain a target large language model includes: Performing a first-stage pre-training on the basic large language model according to the first data set to obtain a coding ability large language model, and performing a second-stage pre-training on the coding ability large language model according to the second data set to obtain the target large language model; or The basic large language model is pre-trained in the first stage according to the second data set to obtain a mathematical ability large language model, and the mathematical ability large language model is pre-trained in the second stage according to the first data set to obtain the target large language model.

4. The method according to claim 2, wherein: The pre-training dataset further includes a third dataset, and the third dataset is used to learn general capabilities. The multi-stage pre-training of the basic large language model based on the pre-training dataset to obtain the target large language model includes one of the following: Performing multi-stage pre-training on the basic large language model according to the first data set, the second data set, and the third data set in sequence to obtain the target large language model; Performing multi-stage pre-training on the basic large language model according to the second data set, the first data set, and the third data set in sequence to obtain the target large language model; Performing multi-stage pre-training on the basic large language model according to the third data set, the first data set, and the second data set in sequence to obtain the target large language model; Performing multi-stage pre-training on the basic large language model according to the third data set, the second data set, and the first data set in sequence to obtain the target large language model; Performing multi-stage pre-training on the basic large language model according to the first data set, the third data set, and the second data set in sequence to obtain the target large language model; The basic large language model is pre-trained in multiple stages according to the second data set, the third data set, and the first data set to obtain the target large language model.

5. The method according to claim 4, wherein The third data set includes mathematical ability training samples, coding ability training samples, and general ability training samples, and the general ability training samples in the third data set account for the highest proportion.

6. The method according to claim 5, wherein: In the first data set, the proportion of mathematical ability training samples and general ability training samples is the same; in the second data set, the proportion of coding ability training samples and general ability training samples is the same; In the third data set, coding ability training samples and mathematical ability training samples account for the same proportion.

7. An interactive method comprising: Get input questions; An output answer corresponding to the input question is determined based on the fine-tuned target large language model, wherein the target large language model is obtained based on the pre-training method according to any one of claims 1 to 6.

8. A large language model pre-training system, comprising: at least one storage medium storing at least one instruction set for pre-training a large language model; At least one processor is communicatively connected to the at least one storage medium, wherein when the at least one processor is running, it reads the at least one instruction set and executes the pre-training method as described in any one of claims 1 to 6 according to the instructions of the at least one instruction set.

9. An interactive system comprising: at least one storage medium storing at least one set of instructions for interaction; At least one processor is communicatively connected to the at least one storage medium, wherein when the at least one processor is running, it reads the at least one instruction set and executes the interaction method as claimed in claim 7 according to the instructions of the at least one instruction set.

10. A computer-readable non-transitory storage medium, wherein: The computer-readable non-transitory storage medium stores at least one instruction set, and the at least one instruction set is executed by at least one processor to implement the method according to any one of claims 1 to 6, or to execute the interaction method according to claim 7 according to the instructions of the at least one instruction set.