Large model parameter optimization method and system based on adaptive learning rate adjustment

Through the adaptive learning rate adjustment method, the problems of large resource consumption and difficult parameter tuning during the training of large models are solved, and efficient and accurate model training is achieved, which is suitable for various deep learning models.

CN120509447APending Publication Date: 2025-08-19SUZHOU COLLABORATIVE INNOVATION INTELLIGENT MFG EQUIP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510573551.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

During the training of large models, computing resources are consumed, training time is long, and parameter tuning is difficult. Traditional methods are inefficient and difficult to adapt to the needs of large-scale data sets and complex models.

Method used

Adaptive learning rate adjustment method is adopted, by initializing the model parameters and learning rate, the loss value on the training set is calculated, the learning rate is dynamically adjusted to adapt to the change of the loss value, and the model parameters are updated until the stop condition is met.

Benefits of technology

It improves the efficiency and accuracy of large-scale model training, and is suitable for deep learning models of various scales and complexities, avoids the cumbersomeness of manual parameter adjustment, and dynamically adjusts the learning rate to accelerate the training process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120509447A_ABST
    Figure CN120509447A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a large model parameter optimization method based on adaptive learning rate adjustment, comprising the following steps: S1, initializing model parameters and a learning rate; s2, calculating the loss of the model on the training set according to the current learning rate and the model parameters; s3, dynamically adjusting the learning rate according to the change condition of the loss value; the invention also provides a large model parameter optimization system based on adaptive learning rate adjustment, which comprises a system module assembly, and the system module assembly comprises a data preprocessing module, a model training module, a learning rate adjustment module and a parameter updating module. According to the method, the learning rate is dynamically adjusted according to the change condition of the loss value, adaptive adjustment of the learning rate is realized, complexity and uncertainty of manual parameter adjustment are avoided, and the method is suitable for deep learning models of various scales and complexity and has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a large model parameter optimization method and system based on adaptive learning rate adjustment. Background Art

[0002] Large models, such as large language models, and their application in artificial intelligence have become a global research hotspot. Large language models are deep learning models trained on massive amounts of text data. They can not only generate natural language text, but also deeply understand the meaning of the text and handle various natural language tasks such as text summarization, question answering, and translation.

[0003] With the rapid development of deep learning technology, large models are increasingly being used in various fields. However, the training process of large models often faces problems such as high computing resource consumption, long training time, and difficulty in parameter tuning. Traditional parameter tuning methods, such as grid search and random search, while able to find optimal parameter combinations to a certain extent, are often inefficient and difficult to adapt to the needs of large datasets and complex models. Therefore, it is of great significance to develop an efficient and adaptive parameter optimization method for large models. Summary of the Invention

[0004] (1) Technical problems solved

[0005] In response to the deficiencies of the existing technology, the present invention provides a large model parameter optimization method and system based on adaptive learning rate adjustment, which dynamically adjusts the learning rate, accelerates the model training process, and improves the accuracy and generalization ability of the model.

[0006] (2) Technical solution

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] The large model parameter optimization method based on adaptive learning rate adjustment includes the following steps:

[0009] S1: Initialize model parameters and learning rate;

[0010] S2: Calculate the loss of the model on the training set based on the current learning rate and model parameters;

[0011] S3: Dynamically adjust the learning rate according to the change of loss value;

[0012] S4: Update the model parameters and repeat steps 2 and 3 until the stopping condition is met.

[0013] As a further solution of the present invention, in S1, normal distribution is used to initialize model parameters and learning rate, and the pytorch deep learning framework and Transformer model structure are selected.

[0014] Furthermore, in S2, the calculation model calculates the loss value using cross entropy loss on the training set based on the current learning rate and model parameters.

[0015] Based on the above solution, the method for dynamically adjusting the learning rate in S3 includes:

[0016] a. Calculate the difference between the current loss value and the loss value at the previous moment;

[0017] b. Determine whether the current learning rate is appropriate based on the size and direction of the difference;

[0018] c. If the current learning rate is too large and causes the loss value to fluctuate, reduce the learning rate;

[0019] d. If the current learning rate is too small and the training speed is too slow, increase the learning rate.

[0020] Furthermore, the current loss value L is calculated in S2 t and the loss value L at the previous moment t-1 The difference between the two values is used to determine whether the current learning rate is appropriate. The difference calculation formula is ΔL = L t -L t-1 , the loss increases, indicating that the learning rate is too high or the model is unstable; the loss decreases, and the learning direction is basically correct. The larger the difference, the more significant the impact of the single-step update on the model. In S3, if the current learning rate is too large and the loss value fluctuates, the learning rate is reduced according to the preset reduction ratio; if the current learning rate is too small and the training speed is too slow, the learning rate is increased according to the preset increase ratio.

[0021] The present invention also proposes a large model parameter optimization system based on adaptive learning rate adjustment, including a system module assembly, which includes a data preprocessing module, a model training module, a learning rate adjustment module and a parameter update module. The data preprocessing module is connected to a data collection module and a data initialization module. The data initialization module includes a deep learning module and a model module. The deep learning module selects the pytorch deep learning framework, and the model module selects the Transformer model structure.

[0022] Based on the above scheme, the model training module is connected to the learning rate adjustment module, which includes a difference calculation module and a judgment module. The judgment module judges whether the current learning rate is appropriate based on the size and direction of the difference.

[0023] Furthermore, the parameter updating module is connected to the learning rate adjusting module.

[0024] (3) Beneficial effects

[0025] Compared with the prior art, the present invention provides a large model parameter optimization method and system based on adaptive learning rate adjustment, which has the following beneficial effects:

[0026] 1. In the present invention, the loss of the model on the training set is calculated based on the current learning rate and model parameters, which improves the efficiency and accuracy of large model training.

[0027] 2. In the present invention, the learning rate is dynamically adjusted according to the change of the loss value, thereby realizing adaptive adjustment of the learning rate and avoiding the tediousness and uncertainty of manual parameter adjustment. It is suitable for deep learning models of various scales and complexities and has broad application prospects.

[0028] 3. In the present invention, the learning rate can be dynamically adjusted to accelerate the model training process, while improving the accuracy and generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 This is a schematic flow chart of the steps of the large model parameter optimization method based on adaptive learning rate adjustment proposed in the present invention.

[0030] Figure 2 This is a schematic diagram of the framework of the large model parameter optimization system based on adaptive learning rate adjustment proposed by the present invention. DETAILED DESCRIPTION

[0031] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0032] Example 1

[0033] Reference Figure 1-2 , a large model parameter optimization method based on adaptive learning rate adjustment, comprising the following steps:

[0034] Step 1: Initialize model parameters and learning rate using normal distribution, select the pytorch deep learning framework and Transformer model structure;

[0035] Step 2: Based on the current learning rate and model parameters, calculate the loss value of the model using cross entropy loss on the training set;

[0036] Step 3: Calculate the difference between the current loss value and the loss value at the previous moment, and determine whether the current learning rate is appropriate based on the size and direction of the difference;

[0037] Difference calculation:

[0038] If the loss increases, the learning rate may be too high or the model is unstable. If the loss decreases, the learning direction is basically correct. The larger the difference, the more significant the impact of the single-step update on the model.

[0039] Step 4: If the current learning rate is too large and causes the loss value to fluctuate, reduce the learning rate according to the preset reduction ratio; if the current learning rate is too small and causes the training speed to be too slow, increase the learning rate according to the preset increase ratio;

[0040] Step 5: Update the model parameters and repeat steps 2 to 4 until the stopping conditions are met (such as loss convergence, reaching the preset number of training rounds, etc.).

[0041] The present invention also proposes a large model parameter optimization system based on adaptive learning rate adjustment, including a system module assembly, which includes a data preprocessing module, a model training module, a learning rate adjustment module and a parameter update module. The data preprocessing module is connected to a data collection module and a data initialization module. The data initialization module includes a deep learning module and a model module. The deep learning module selects the pytorch deep learning framework, and the model module selects the Transformer model structure.

[0042] Example 2

[0043] Reference Figure 1-2 , a large model parameter optimization method based on adaptive learning rate adjustment, comprising the following steps:

[0044] Step 1: Initialize model parameters and learning rate using normal distribution, select the pytorch deep learning framework and Transformer model structure;

[0045] Step 2: Based on the current learning rate and model parameters, calculate the loss value of the model using cross entropy loss on the training set;

[0046] Step 3: Calculate the difference between the current loss value and the loss value at the previous moment, and determine whether the current learning rate is appropriate based on the size and direction of the difference;

[0047] Difference calculation:

[0048] If the loss increases, the learning rate may be too high or the model is unstable. If the loss decreases, the learning direction is basically correct. The larger the difference, the more significant the impact of the single-step update on the model.

[0049] Step 4: If the current learning rate is too large and causes the loss value to fluctuate, reduce the learning rate according to the preset reduction ratio; if the current learning rate is too small and causes the training speed to be too slow, increase the learning rate according to the preset increase ratio;

[0050] Step 5: Update the model parameters and repeat steps 2 to 4 until the stopping conditions are met (such as loss convergence, reaching the preset number of training rounds, etc.).

[0051] The present invention also proposes a large model parameter optimization system based on adaptive learning rate adjustment, including a system module assembly, which includes a data preprocessing module, a model training module, a learning rate adjustment module and a parameter update module. The data preprocessing module is connected to a data collection module and a data initialization module. The data initialization module includes a deep learning module and a model module. The deep learning module selects the pytorch deep learning framework, and the model module selects the Transformer model structure.

[0052] In particular, the model training module is connected to the learning rate adjustment module, the learning rate adjustment module includes a difference calculation module and a judgment module, the judgment module judges whether the current learning rate is appropriate based on the size and direction of the difference, and the parameter update module is connected to the learning rate adjustment module.

[0053] In the description herein, it should be noted that relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Furthermore, the terms "include," "comprise," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0054] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A large model parameter optimization method based on adaptive learning rate adjustment, characterized in that: The following steps are involved: S1: Initialize model parameters and learning rate; S2: Calculate the loss of the model on the training set based on the current learning rate and model parameters; S3: Dynamically adjust the learning rate according to the change of loss value; S4: Update the model parameters and repeat steps 2 and 3 until the stopping condition is met.

2. The large model parameter optimization method based on adaptive learning rate adjustment according to claim 1, characterized in that: In S1, the normal distribution is used to initialize the model parameters and learning rate, and the pytorch deep learning framework and Transformer model structure are selected.

3. The large model parameter optimization method based on adaptive learning rate adjustment according to claim 2 is characterized in that: In S2, the calculation model calculates the loss value using cross entropy loss on the training set based on the current learning rate and model parameters.

4. The large model parameter optimization method based on adaptive learning rate adjustment according to claim 1, characterized in that The method for dynamically adjusting the learning rate in S3 includes: a. Calculate the difference between the current loss value and the loss value at the previous moment; b. Determine whether the current learning rate is appropriate based on the size and direction of the difference; c. If the current learning rate is too large and causes the loss value to fluctuate, reduce the learning rate; d. If the current learning rate is too small and the training speed is too slow, increase the learning rate.

5. The large model parameter optimization method based on adaptive learning rate adjustment according to claim 1, characterized in that: The current loss value L is calculated in S2 t and the loss value L at the previous moment t-1 The difference between the two values is used to determine whether the current learning rate is appropriate. The difference calculation formula is ΔL = L t -L t-1 , the loss increases, indicating that the learning rate is too high or the model is unstable; the loss decreases, and the learning direction is basically correct. The larger the difference, the more significant the impact of the single-step update on the model. In S3, if the current learning rate is too large and the loss value fluctuates, the learning rate is reduced according to the preset reduction ratio; if the current learning rate is too small and the training speed is too slow, the learning rate is increased according to the preset increase ratio.

6. A large model parameter optimization system based on adaptive learning rate adjustment, characterized in that: It includes a system module assembly, which includes a data preprocessing module, a model training module, a learning rate adjustment module and a parameter update module. The data preprocessing module is connected to a data collection module and a data initialization module. The data initialization module includes a deep learning module and a model module. The deep learning module selects the pytorch deep learning framework, and the model module selects the Transformer model structure.

7. The large model parameter optimization system based on adaptive learning rate adjustment according to claim 6, characterized in that: The model training module is connected to the learning rate adjustment module, which includes a difference calculation module and a judgment module. The judgment module judges whether the current learning rate is appropriate based on the size and direction of the difference.

8. The large model parameter optimization method based on adaptive learning rate adjustment according to claim 6, characterized in that: The parameter updating module is connected to the learning rate adjusting module.