Extensible large model training reasoning method, device, equipment and medium
By introducing autonomous dynamic discriminators and stacked model structures into large models, the number of sub-models of the model is automatically selected, which solves the problem of high resource consumption of large model training and inference, improves efficiency and reduces costs.
Patent Information
- Application Number
- CN202411847995.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-16
AI Technical Summary
Large models require a lot of time and computing resources during training and inference, resulting in inefficiency and waste of resources, especially in resource-constrained environments that are difficult to deploy and use.
The expansion-based large-model training inference method based on autonomous dynamic discriminators is adopted. Through the combination of stacked model structure and autonomous dynamic discriminators, the number of sub-models of the model is automatically selected, and the inference efficiency is improved. The model training and inference process are optimized through the punishment term and gradient stop back-passing technique.
It realizes the efficiency improvement of the large-model inference process, the trade-off between accuracy and speed, saves model inference resource consumption, reduces calculation costs, and is suitable for a wide range of application scenarios.
Smart Images

Figure CN119940525A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an expandable large-model training reasoning method, device, equipment and medium. Background Art
[0002] Large models have demonstrated powerful text generation capabilities in today's life and have been widely used in many fields such as news creation, social media content creation, and intelligent customer service. However, since large models often have large parameters and complex calculations, the training and reasoning processes require a lot of time and computing resources, making them difficult to put into use. At the same time, large models have a large parameter scale and require massive training data and powerful computing resource support. The training process often takes weeks or even months, and consumes huge computing resources. This not only limits the efficiency of large models in practical applications, but also poses a huge challenge to computing power and energy consumption. At the same time, large models usually require longer reasoning time, making their deployment and use complicated and expensive, limiting their application in resource-constrained environments. Therefore, there is an urgent need to study key technologies to accelerate large model reasoning in order to reduce resource consumption and promote the sustainable development of artificial intelligence technology.
[0003] In the field of deep learning, large pre-trained models are widely used due to their powerful performance, but efficiency and resource consumption during inference remain a major challenge. Traditional inference techniques usually rely on full-model inference, which leads to slow response times and high computational costs. In addition, although model compression and distributed inference can improve speed to a certain extent, these methods often sacrifice model performance and increase system complexity. Especially in application scenarios that require real-time response, existing solutions cannot meet the needs, lack flexibility, and are difficult to adapt to rapidly changing application environments. Therefore, it is particularly important to explore new accelerated inference technologies. Summary of the invention
[0004] In order to solve at least one of the technical problems existing in the prior art to a certain extent, the purpose of the present invention is to provide a scalable large model training reasoning method, device, equipment and medium based on an autonomous dynamic discriminator.
[0005] The first technical solution adopted by the present invention is:
[0006] A scalable large model training and reasoning method, comprising the following steps:
[0007] Obtain text data and build a training set;
[0008] Constructing a large model, wherein the large model is a stacked model structure that performs knowledge sharing in a horizontal direction, and the large model includes multiple sub-models;
[0009] Build an autonomous dynamics discriminator. The output of each sub-model will be input into the autonomous dynamics discriminator, and the output of the autonomous dynamics discriminator will be used as the final model prediction.
[0010] The training set is used to train the large model, and the trained large model is used to implement the text generation task.
[0011] Furthermore, the multiple sub-models adopt different encoder architectures.
[0012] Furthermore, the autonomous dynamic discriminator works as follows:
[0013] The autonomous dynamic discriminator records the threshold content, autonomously divides the prediction results according to the number of sub-models, and updates the division threshold in real time according to samples during the inference process.
[0014] Furthermore, the training process of the large model is as follows:
[0015] First, use part of the training data to pre-train the first sub-model separately. The expression is:
[0016] x1=Submodel0(tokenizer(x0))
[0017] In the formula, x0 is the initial input of the model, representing the initial input sequence of the text generation task; tokenizer is a word segmenter, which is used to convert the original input data into a format that can be processed by the model; x1 is the output after passing through the first sub-model, and Submodel0 is the first sub-model;
[0018] The output x1 is further input to the autonomous dynamic discriminator, expressed as:
[0019] y=Discriminator(x1)
[0020] In the formula, Discriminator is the autonomous dynamic discriminator, and y is the final prediction result;
[0021] After pre-training converges, the weight of the first sub-model is shared with the remaining sub-models, effectively avoiding the waste of resources caused by different sub-models redundantly learning shared knowledge from scratch;
[0022] Then the entire model is trained: First, for the initial input x0, each sub-model will get an output x i , where x i+1 =Submodel i (x i ), thus we get x0,…,x i ; After getting each x iAfter that, they will be further input into the autonomous dynamic discriminator to obtain the prediction result y i , and back-propagate to update the parameters of the sub-model. The mathematical expression is:
[0023]
[0024] In the formula, represents the stop_gradient operation, that is, the subsequent sub-model will freeze the previous sub-model when updating the parameters, to prevent the previously learned sub-model parameters from being disturbed by the subsequent sub-model learning, thus preventing the forgetting problem from occurring.
[0025] Furthermore, the process of large model training also includes the steps of autonomous adjustment:
[0026] When the entire model is trained, penalty items of the same form are added to sub-models at different levels. When the confidence of the output result of the latter sub-model is less than that of the previous sub-model, the latter sub-model is penalized in the loss function. In the optimization process, the output capabilities of sub-models at different levels are distinguished, thereby realizing the autonomous model scale selection of the network.
[0027] Furthermore, for the sub-models at different levels, penalty items of the same form are added, including:
[0028] Two loss function penalty terms are added:
[0029] The first is the hierarchical loss, which introduces different weight coefficients for the prediction loss of each layer, so that the top-level encoder has a higher weight in the loss function; the prediction loss of each sub-model is L i , the loss is defined as:
[0030]
[0031] In the formula, w i is the weight;
[0032] The second is the level consistency penalty:
[0033]
[0034] In the formula, p i is the predicted distribution of the i-th sub-model;
[0035] Through these two penalty terms, the confidence of the model output is higher as the sub-models are stacked, and the predictions of the low-level sub-models are closer to the predictions of the high-level sub-models, while ensuring the overall improvement. The final loss is defined as:
[0036] Loss=Layer Loss+Consistency Penalty
[0038] Furthermore, the large model works as follows during the inference phase:
[0039] During the inference phase, the autonomous dynamic discriminator maintains a threshold statistic τ0,…,τ for each sub-model n ,in:
[0040] τ i =μ i +k·σ i
[0041] In the formula, μ i is the mean of the confidence of the ith sub-model, σ i is the standard deviation, and k is an adjustable hyperparameter used to control the strictness of exit;
[0042] When the input data passes through the model, the output confidence c of each layer of sub-model is calculated i , for each layer i, compare its output confidence c i With the preset threshold τ i , if c i >τ i , it is considered that the layer has a sufficiently confident prediction, exits early, and outputs the prediction result; otherwise, the input is passed to the next sub-model until the confidence of a certain layer of sub-model reaches the threshold or reaches the last layer of sub-model;
[0043] Finally, the autonomous dynamic discriminator obtains the output results of multiple sub-models for further fusion to obtain the next token prediction:
[0044]
[0045] In the formula, y i is the output of each sub-model, and L is the number of sub-models;
[0046] Then the predictions of each sub-model can be integrated. The sub-model with higher confidence will contribute more to the final prediction. Finally, the entire text generation task is completed in a cycle to achieve interaction between man and machine.
[0047] The second technical solution adopted by the present invention is:
[0048] A scalable large model training and reasoning device, comprising:
[0049] Data acquisition module, used to obtain text data and build training sets;
[0050] A model building module is used to build a large model, wherein the large model is a stacked model structure that performs knowledge sharing in the horizontal direction, and the large model includes multiple sub-models;
[0051] The discriminator building module is used to build an autonomous dynamic discriminator. The output of each sub-model will be input into the autonomous dynamic discriminator, and the output of the autonomous dynamic discriminator will be used as the final model prediction;
[0052] The model training module is used to train the large model using the training set, and use the trained large model to implement the text generation task.
[0053] The third technical solution adopted by the present invention is:
[0054] An electronic device comprises a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement a scalable large model training inference method as described above.
[0055] The fourth technical solution adopted by the present invention is:
[0056] A computer-readable storage medium, wherein at least one instruction, at least one program, a code set or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement a scalable large model training inference method as described above.
[0057] The fifth technical solution adopted by the present invention is:
[0058] A computer program product or a computer program, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above-mentioned scalable large model training reasoning method.
[0059] The beneficial effects of the present invention are as follows: the present invention aims to realize the autonomous selection of sub-models by the model, and autonomously selects the number of sub-models participating in reasoning through an autonomous dynamic discriminator, thereby improving the efficiency of the large model reasoning process, achieving a balance between accuracy and speed, and saving model reasoning resource consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the embodiments of the present invention or the drawings of related technical solutions in the prior art are introduced below. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0061] Figure 1 It is a flowchart of the steps of a scalable large model training and reasoning method in an embodiment of the present invention;
[0062] Figure 2 It is a structural design diagram of an expandable large model in an embodiment of the present invention. DETAILED DESCRIPTION
[0063] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limitations of the present invention. For the step numbers in the following embodiments, they are only provided for the convenience of explanation, and the order between the steps is not limited in any way, and the execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.
[0064] In the description of the present invention, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., and orientations or positional relationships indicated are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.
[0065] In the description of the present invention, "several" means one or more, "more" means more than two, "greater than", "less than", "exceed" etc. are understood as not including the number itself, and "above", "below", "within" etc. are understood as including the number itself. If there is a description of "first" or "second", it is only used for the purpose of distinguishing the technical features, and cannot be understood as indicating or implying the relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.
[0066] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention based on the specific content of the technical solution.
[0067] Example 1
[0068] like Figure 1 As shown, this embodiment provides a scalable large model training and reasoning method based on an autonomous dynamic discriminator, which solves the problem of huge resource consumption and low efficiency of existing large model training and reasoning, allowing large models to balance accuracy and speed during the training and reasoning process, thereby improving model service efficiency. The method specifically includes the following steps:
[0069] S1. Obtain text data and build a training set.
[0070] S2. Construct a large model, wherein the large model is a stacked model structure that performs knowledge sharing in the horizontal direction, and the large model includes multiple sub-models.
[0071] This step mainly involves designing an extensible model structure to achieve an efficient and extensible model structure, thereby ensuring that the model can balance accuracy and speed during reasoning.
[0072] Specifically, see Figure 2 In order to make the model structure scalable, this embodiment designs a stacked model structure that shares knowledge in the horizontal direction. Each sub-model is responsible for converting the input sequence into sequence features. Different encoder architectures can be used here, common ones include Transformer-based architectures (such as BERT, RoBERTa, BART, etc.). Each sub-model has its own output, and is further output to the autonomous dynamic discriminator to obtain the final prediction result. Exemplarily, in order to maintain the convenience of semantic information transfer between sub-models, all sub-models use exactly the same model architecture and are expanded in the horizontal direction. The sub-models are connected to the parallel modules of the next sub-network through the output of each module of the previous sub-model to transfer information to achieve efficient semantic sharing.
[0073] S3. Build an autonomous dynamic discriminator. The output of each sub-model will be input into the autonomous dynamic discriminator, and the output of the autonomous dynamic discriminator will be used as the final model prediction.
[0074] An autonomous dynamic discriminator is constructed to regulate the confidence levels between sub-models during training to achieve an even distribution of tasks.
[0075] Specifically, in order to enable the large model to select the number of sub-models used in reasoning according to the complexity of the task, this embodiment constructs an autonomous dynamic discriminator. In the training stage, the training sample passes through the model, and each sub-model will have an output, and the number of outputs is equal to the number of sub-models. These outputs are all input into the autonomous dynamic discriminator. The output is the final model prediction, and back propagation is performed according to the data set label. Considering the stacking structure, the subsequent sub-model will affect the parameter update of the previous sub-model during back propagation. Therefore, when the model parameters are updated according to the output of each sub-model, the parameters of the previous sub-model will be frozen. Thus, in the training stage, each sample can update all the parameters of the model and realize that each sub-model can make predictions. In the reasoning stage, for an ideal discriminator, it should be able to divide the set of input space equally into each sub-model to achieve the best allocation requirements. Therefore, the autonomous dynamic discriminator will record the threshold content, autonomously divide the prediction results according to the number of sub-models, and update the threshold of the division in real time according to the sample during the reasoning process.
[0076] S4. Use the training set to train the large model, and use the trained large model to achieve text generation tasks.
[0077] During the model training and reasoning process, due to the large number of sub-models, this embodiment proposes an efficient model training and reasoning algorithm.
[0078] In some embodiments, during the training phase, a portion of the training data is first used to perform a separate pre-training on the first sub-model, and the mathematical expression is:
[0079] x1=Submodel0(tokenizer(x0))
[0080] Among them, x0 is the initial input of the model, representing the initial input sequence of the text generation task, tokenizer is a word segmenter, which is used to convert the original input data into a format that the model can process, x1 is the output after passing through the first sub-model, and Submodel0 is the first sub-model. This output is further input to the autonomous dynamic discriminator:
[0081] y=Discriminator(x1)
[0082] Among them, Discriminator is the autonomous dynamic discriminator, and y is the final prediction result. After the pre-training converges, the weight of the first sub-model is shared with all other sub-models, which effectively avoids the waste of resources caused by different sub-models redundantly learning shared knowledge from scratch.
[0083] Then the entire model is trained. First, for the initial input x0, each sub-network will get an output xi , where x i+1 =Submodel i (x i ), thus we get x0,…,x i , after getting each x i After that, they will be further input into the autonomous dynamic discriminator to obtain the prediction result y i , and back-propagate to update the parameters of the sub-model.
[0084] The mathematical expression is:
[0085]
[0086] in, Represents the stop_gradient operation, that is, the subsequent sub-model will freeze the previous sub-model when updating parameters to prevent the previously learned sub-model parameters from being disturbed by the learning of the subsequent sub-model, thereby preventing the forgetting problem from occurring.
[0087] In some embodiments, in the autonomous adjustment method, the difference in expressive power between different subnets is crucial. Ideally, the prediction results involving multiple subnets need to have higher confidence than the prediction results of a single subnet or a small number of subnets. Therefore, when the entire network is trained, a penalty term of the same form is added to the subnets of different levels. When the confidence of the output result of the next subnet is less than that of the previous subnet, the next subnet is penalized in the loss function, and the purpose of distinguishing the output capabilities of subnets of different levels is achieved during the optimization process, thereby realizing the autonomous model scale selection of the network.
[0088] Specifically, two penalty terms are added to the loss function. The first is the layered loss, which introduces different weight coefficients for the prediction loss of each layer, so that the top-level encoder has a higher weight in the loss function. The prediction loss of each sub-model is L i , the loss is defined as:
[0089]
[0090] where w i is the weight, set to a value w that increases with the number of layers i =exp(i).
[0091] The second is the level consistency penalty:
[0092]
[0093] Among them, p iis the prediction distribution of the ith sub-model. Through these two penalty terms, the confidence of the model output is higher as the sub-models are stacked, and the predictions of the lower-level sub-models are closer to the predictions of the higher-level sub-models, while ensuring the overall improvement.
[0094] The final loss is defined as:
[0095] Loss=Layer Loss+Consistency Penalty
[0096] In some embodiments, during the inference phase, the autonomous dynamics discriminator maintains a threshold statistic τ0,…,τ for each sub-model n , where τ i =μ i +k·σ i , μ i is the mean of the confidence of the ith sub-model, σ i is the standard deviation, and k is an adjustable hyperparameter used to control the strictness of the exit. When used specifically, when the input data passes through the model, the output confidence c of each layer is calculated i , for each layer i, compare its output confidence c i With the preset threshold τ i , if c i >τ i , then it is considered that the layer has a sufficiently confident prediction and can exit early and output the prediction result. Otherwise, the input is passed to the next sub-model until the confidence of a certain layer reaches the threshold or reaches the last layer.
[0097] Finally, the autonomous dynamic discriminator obtains the output results of multiple sub-models for further fusion to obtain the next token prediction:
[0098]
[0099] in Then the predictions of each sub-model can be integrated. The sub-model with higher confidence will contribute more to the final prediction. Finally, the entire text generation task is completed in a cycle to achieve interaction between man and machine.
[0100] In general, existing model acceleration reasoning solutions such as model compression and distributed reasoning can improve model reasoning performance to a certain extent. However, these methods often sacrifice model performance, or have certain hardware requirements that further increase the complexity of the system. In response to these problems, the present invention adopts reasonable initialization and gradient stop return techniques to effectively improve the efficiency of the model in training. At the same time, the present invention aims to enable the model to autonomously select sub-models. Through a series of model structure designs and training methods, the model can perceive the accuracy of the reasoning results, and autonomously select output results or further predictions, greatly reducing unnecessary waste of resources, thereby achieving efficient human-computer interaction while ensuring accuracy and improving the user's interactive experience.
[0101] In summary, compared with the prior art, the method of the present invention has at least the following advantages and beneficial effects:
[0102] (1) In large model training reasoning, the present invention allows the large model to autonomously select the number of sub-models involved in reasoning according to the difficulty of the input sequence, thereby improving the efficiency of the large model reasoning process, achieving a trade-off between accuracy and speed, and saving model reasoning resource consumption.
[0103] (2) The present invention reduces the complexity and time-consuming training of the model from scratch through a reasonable model initialization method, and also reduces the required computing resources and costs.
[0104] (3) The present invention provides a training reasoning paradigm with wide application adaptability, which is applicable to the current stacked large model structure and is not limited to the text generation model mentioned in the present invention. Therefore, the scheme of the present invention has high technical advantages and wide practical application value.
[0105] Example 2
[0106] This embodiment provides a scalable large model training and reasoning device, including:
[0107] Data acquisition module, used to obtain text data and build training sets;
[0108] A model building module is used to build a large model, wherein the large model is a stacked model structure that performs knowledge sharing in the horizontal direction, and the large model includes multiple sub-models;
[0109] The discriminator building module is used to build an autonomous dynamic discriminator. The output of each sub-model will be input into the autonomous dynamic discriminator, and the output of the autonomous dynamic discriminator will be used as the final model prediction;
[0110] The model training module is used to train the large model using the training set, and use the trained large model to implement the text generation task.
[0111] Since the device is an expandable large-model training and inference device of an embodiment of the present invention, and the principle of solving the problem by the device is similar to that of the method, the implementation of the device can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0112] Example 3
[0113] An embodiment of the present invention further provides an electronic device, the electronic device comprising a processor and a memory, the memory storing at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, the at least one program, the code set or the instruction set being loaded and executed by the processor to implement the following Figure 1 A scalable large model training inference method is shown.
[0114] It is understandable that the memory may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory may be used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data created according to the use of the server, etc.
[0115] The processor may include one or more processing cores. The processor uses various interfaces and lines to connect the various parts of the entire server, and executes various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory, and calling data stored in the memory. Optionally, the processor can be implemented in at least one hardware form of digital signal processing (DSP), field programmable gate array (FPGA), and programmable logic array (PLA). The processor can integrate one or a combination of a central processing unit (CPU) and a modem. Among them, the CPU mainly processes the operating system and application programs; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor, but implemented separately through a chip.
[0116] Since the electronic device is an electronic device corresponding to a scalable large model training inference method of an embodiment of the present invention, and the principle of solving the problem by the electronic device is similar to that of the method, the implementation of the electronic device can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0117] Example 4
[0118] The embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, the at least one program, the code set or instruction set is loaded and executed by a processor to implement the following Figure 1 A scalable large model training inference method is shown.
[0119] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable rewritable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.
[0120] Since the storage medium is a storage medium corresponding to a scalable large model training inference method in an embodiment of the present invention, and the principle of solving the problem by the storage medium is similar to that of the method, the implementation of the storage medium can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0121] Example 5
[0122] In some possible implementations, various aspects of the method of the embodiment of the present invention may also be implemented in the form of a program product, which includes a program code, and when the program product is run on a computer device, the program code is used to cause the computer device to execute the steps of a scalable large model training reasoning method according to various exemplary embodiments of the present application described above in this specification. Among them, the executable computer program code or "code" for executing each embodiment can be written in a high-level programming language such as C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, structured query language (e.g., Transact-SQL), Perl, or in various other programming languages.
[0123] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0124] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0125] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable ordinary technicians in the field to understand the content of the present invention and implement it accordingly, and they cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made based on the essence of the content of the present invention should be included in the protection scope of the present invention.
Claims
1. A scalable large model training and reasoning method, characterized in that: The following steps are involved: Obtain text data and build a training set; Constructing a large model, wherein the large model is a stacked model structure that performs knowledge sharing in a horizontal direction, and the large model includes multiple sub-models; Build an autonomous dynamics discriminator. The output of each sub-model will be input into the autonomous dynamics discriminator, and the output of the autonomous dynamics discriminator will be used as the final model prediction. The training set is used to train the large model, and the trained large model is used to implement the text generation task.
2. The scalable large model training inference method according to claim 1, characterized in that: The multiple sub-models adopt different encoder architectures.
3. The scalable large model training inference method according to claim 1, characterized in that: The working mode of the autonomous dynamic discriminator is: The autonomous dynamic discriminator records the threshold content, autonomously divides the prediction results according to the number of sub-models, and updates the division threshold in real time according to samples during the inference process.
4. The scalable large model training inference method according to claim 1, characterized in that: The training process of the large model is as follows: First, use part of the training data to pre-train the first sub-model separately. The expression is: x1=Submodel0(tokenizer(x0)) In the formula, x0 is the initial input of the model, representing the initial input sequence of the text generation task; tokenizer is a word segmenter, which is used to convert the original input data into a format that can be processed by the model; x1 is the output after passing through the first sub-model, and Submodel0 is the first sub-model; The output x1 is further input to the autonomous dynamic discriminator, expressed as: y=Discriminator(x1) In the formula, Discriminator is the autonomous dynamic discriminator, and y is the final prediction result; After pre-training converges, the weights of the first sub-model are shared with the remaining sub-models; Then the entire model is trained: First, for the initial input x0, each sub-model will get an output x i , where x i+1 =Submodel i (x i ), thus we get x0,…,x i ; After getting each x i After that, they will be further input into the autonomous dynamic discriminator to obtain the prediction result y i , and back-propagate to update the parameters of the sub-model. The mathematical expression is: In the formula, Represents the stop_gradient operation, that is, the subsequent sub-model will freeze the previous sub-model when updating parameters to prevent the previously learned sub-model parameters from being disturbed by the learning of the subsequent sub-model, thereby preventing the forgetting problem from occurring.
5. The scalable large model training inference method according to claim 1, characterized in that: The process of large model training also includes the steps of autonomous adjustment: When the entire model is trained, penalty items of the same form are added to sub-models at different levels. When the confidence of the output result of the latter sub-model is less than that of the previous sub-model, the latter sub-model is penalized in the loss function. In the optimization process, the output capabilities of sub-models at different levels are distinguished, thereby realizing the autonomous model scale selection of the network.
6. The scalable large model training inference method according to claim 5, characterized in that: For sub-models at different levels, penalty items of the same form are added, including: Two loss function penalty terms are added: The first is the hierarchical loss, which introduces different weight coefficients for the prediction loss of each layer, so that the top-level encoder has a higher weight in the loss function; the prediction loss of each sub-model is L i , the loss is defined as: In the formula, w i is the weight; The second is the level consistency penalty: In the formula, p i is the predicted distribution of the i-th sub-model; The final loss is defined as: Loss=Layer Loss+Consistency Penalty.
7. The scalable large model training inference method according to claim 1, characterized in that: The large model works as follows during the inference phase: During the inference phase, the autonomous dynamic discriminator maintains a threshold statistic τ0,…,τ for each sub-model n ,in: t i =μ i +k·s i In the formula, μ i is the mean of the confidence of the ith sub-model, σ i is the standard deviation, k is an adjustable hyperparameter; When the input data passes through the model, the output confidence c of each layer of sub-model is calculated i , for each layer i, compare its output confidence c i With the preset threshold τ i , if c i >τ i , it is considered that the layer has a sufficiently confident prediction, exits early, and outputs the prediction result; otherwise, the input is passed to the next sub-model until the confidence of a certain layer of sub-model reaches the threshold or reaches the last layer of sub-model; Finally, the autonomous dynamic discriminator obtains the output results of multiple sub-models for further fusion to obtain the next token prediction: In the formula, y i is the output of each sub-model, and L is the number of sub-models; This will then realize the integration of the predictions of each sub-model. The sub-model with higher confidence will contribute more to the final prediction, and finally the entire text generation task will be completed in a cycle.
8. A scalable large model training and reasoning device, characterized in that: include: Data acquisition module, used to obtain text data and build training sets; A model building module is used to build a large model, wherein the large model is a stacked model structure that performs knowledge sharing in the horizontal direction, and the large model includes multiple sub-models; The discriminator building module is used to build an autonomous dynamic discriminator. The output of each sub-model will be input into the autonomous dynamic discriminator, and the output of the autonomous dynamic discriminator will be used as the final model prediction; The model training module is used to train the large model using the training set, and use the trained large model to implement the text generation task.
9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Neural network reasoning method and device based on layered loading and medium
CN118278524A
Method for generating pre-trained model, electronic device and storage medium
US20220335711A1
Method, system and apparatus for federated learning
WO2022169136A1