Model training method and system

Through data preprocessing, distributed data set division, mixed precision training, adaptive learning rate and model compression, the model training process is optimized, and the problems of high computing power demand and low efficiency in the model training process are solved, and efficient and stable model training is achieved.

CN120387489APending Publication Date: 2025-07-29INSPUR FINANCIAL INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510358209.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

During the training process of existing models, with the expansion of model scale and the growth of data sets, the demand for computing power has increased sharply, resulting in high demand for hardware resources, increased training costs, and low training efficiency, making it difficult to improve training efficiency without sacrificing model performance.

Method used

Through data preprocessing and feature selection, distributed data set division and cache, hybrid precision training framework, adaptive learning rate adjustment, model compression and pruning, and optimization process control based on reinforcement learning, the model training process is optimized, the amount of data is reduced, the data access speed is improved, the hardware computing power is fully utilized, and the training process is automatically optimized.

Benefits of technology

It significantly reduces the computing power consumption of model training, improves training efficiency and stability, saves computing power resources and time costs, and promotes the widespread application of artificial intelligence technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387489A_ABST
    Figure CN120387489A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method and system, and the method comprises the steps: employing a pruning algorithm to recognize and remove connections or neurons, which have small contribution to a prediction result, in a model in a model training process or after the training is completed; meanwhile, converting model parameters from high-precision representation to low-precision representation by adopting a model quantification method; a reinforcement learning agent is constructed, each decision point in the model training process is used as an action space of the agent, the agent learns an optimal decision sequence in different training states through interaction with a training environment, and the whole model training process is automatically optimized; according to the method, the original data volume can be reduced through data preprocessing and feature selection, and the input scale of model training is remarkably reduced; a distributed data set division and caching strategy improves the data access speed and reduces the I / O overhead; the hybrid precision training framework makes full use of the hardware computing power under the condition of not influencing the model precision, and the computing power demand is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of model training, and particularly to a model training method and system. Background Art

[0002] With the rapid development of artificial intelligence technology, model training has been widely applied in many fields, such as image recognition and natural language processing in deep learning. However, the current model training process faces many challenges. On the one hand, with the continuous expansion of model scale and the rapid growth of datasets, the computing power required for model training increases exponentially, which poses extremely high requirements on hardware resources, such as high-performance GPU clusters and large-scale data centers, resulting in a substantial increase in training costs. Many enterprises and research institutions are restricted by computing power resources and it is difficult to carry out large-scale and efficient model training. On the other hand, traditional model training methods lack effective optimization mechanisms. When dealing with complex tasks and large-scale data, the training efficiency is low, the convergence speed is slow, which further exacerbates the waste of resources and the extension of training time. How to improve the model training efficiency and significantly reduce the computing power consumption without sacrificing the model performance has become one of the key problems restricting the further promotion and application of artificial intelligence technology. Summary of the Invention

[0003] The purpose of the present invention is to provide a model training method and system for the above problems in the prior art, so as to solve all or one of the above problems existing in the prior art.

[0004] To solve the above technical problems, the specific technical solutions of the present invention are as follows: On the one hand, the present invention provides a model training method, including: Data preprocessing and feature selection: performing multi-dimensional cleaning on the original data to remove noise, duplicate, and incomplete data, and using a feature importance evaluation algorithm to screen out a few feature subsets that have a key impact on the model training results, so as to reduce the amount of data to be processed in subsequent model training; Distributed dataset partitioning and caching strategy: partitioning the filtered dataset into multiple subsets according to the distributed principle, and allocating each subset to different computing nodes for processing. Moreover, an intelligent caching mechanism is introduced to locally store some common data subsets on the computing nodes, so as to reduce the frequent loading of data from storage devices to computing units and improve the data access speed; Construction of a mixed-precision training framework: During the model training process, adopting a mixed-precision data representation, that is, preferentially using a low-precision data type during the calculation process, and using a high-precision data type for some key calculation steps with extremely high precision requirements; automatically adapting the parameters of mixed-precision calculation for different hardware platforms, so as to make the most of the hardware computing power to the greatest extent without affecting the model convergence and reduce the computing power consumption; Adaptive learning rate dynamic adjustment: Monitor the training status of the model in real time based on information such as the loss function value and gradient change during model training, and adopt an adaptive learning rate algorithm to dynamically adjust the learning rate according to the training status. Use a larger learning rate at the beginning of training to converge quickly, and gradually reduce the learning rate when approaching convergence to improve the stability and training efficiency of the model; Model compression and pruning: During or after model training, use pruning algorithms to identify and remove connections or neurons in the model that contribute less to the prediction results, reducing the complexity and computational cost of the model; At the same time, adopt model quantization methods to convert model parameters from high-precision representations to low-precision representations, further reducing model storage requirements and computational overhead; Optimization process control based on reinforcement learning: Construct a reinforcement learning agent, take each decision point during model training as the action space of the agent, and the agent learns the optimal decision sequence in different training states through interaction with the training environment, automatically optimizing the entire model training process to improve training efficiency and computing power utilization.

[0005] As an improved solution, the multi-dimensional cleaning includes: separately processing various situations such as numerical anomalies, null values, duplicate records, and format errors in the data.

[0006] As an improved solution, the feature importance evaluation algorithm includes: algorithms based on information gain, Gini index, and random forest feature importance ranking.

[0007] As an improved solution, when dividing the filtered data set into multiple subsets according to the distributed principle, perform balanced allocation based on the distribution characteristics of the data and the performance parameters of the computing nodes to ensure load balancing of each computing node.

[0008] As an improved solution, during the model training process, automatically detect the accuracy requirements of each calculation step, and dynamically switch to use low-precision or high-precision data types according to the detection results.

[0009] As an improved solution, the adaptive learning rate algorithm includes: optimized versions of the Adagrad, Adadelta, and Adam algorithms.

[0010] As an improved solution, in the model compression and pruning steps, adopt a strategy combining step-by-step pruning and global pruning. First, perform step-by-step pruning to determine key connections and neurons, and then perform global pruning to further improve the model compression effect.

[0011] As an improved solution, the reinforcement learning agent adopts a deep reinforcement learning algorithm to estimate state values and action values through a multi-layer neural network to more accurately learn the optimal decision-making strategy.

[0012] On the other hand, the present invention also provides a model training system, including: A data preprocessing module for cleaning and feature selection of the original data, and screening out a key feature subset; A distributed processing module responsible for distributed partitioning of the data set and implementing a data caching strategy to provide efficient data access services for different computing nodes; A mixed-precision training module for constructing a mixed-precision training framework and automatically matching and adjusting mixed-precision calculation parameters according to the hardware platform; A learning rate adjustment module for adjusting the learning rate according to the model training state by using an adaptive learning rate algorithm to optimize the model training process; A model compression module for performing model pruning and quantization operations to reduce the complexity and computational overhead of the model; A reinforcement learning optimization module including a reinforcement learning agent, which learns and executes an optimal training process optimization strategy by interacting with the training environment.

[0013] As an improved solution, the model training system further includes: A communication module for realizing efficient data transmission and instruction interaction among the data preprocessing module, the distributed processing module, the mixed-precision training module, the learning rate adjustment module, the model compression module and the reinforcement learning optimization module to ensure the collaborative work and stability of the entire training system.

[0014] The beneficial effects of the technical solution of the present invention are as follows: The present invention can reduce the amount of original data through data preprocessing and feature selection, significantly reducing the input scale of model training; the distributed data set partitioning and caching strategy improve the data access speed and reduce the I / O overhead; the mixed-precision training framework makes full use of the hardware computing power without affecting the model accuracy, effectively reducing the computing power requirements; the adaptive learning rate dynamically adjusts to optimize the stability of the training process and speeds up the convergence rate; model compression and pruning further reduce the amount of calculation; the optimization process control based on reinforcement learning realizes the intelligent optimization of the entire training process, enabling each strategy to work together, comprehensively improving the model training efficiency and significantly reducing the computing power consumption. On the premise of maintaining the stable performance of the model, the technology of the present invention can save a large amount of computing power resources and time costs for enterprises and research institutions, and strongly promote the wide application and popularization of artificial intelligence technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0016] Figure 1 is a schematic flowchart of the model training method described in Embodiment 1 of the present invention; Figure 2 is a schematic architecture diagram of the model training system described in Embodiment 2 of the present invention. Specific Embodiments

[0017] The following will elaborate on the preferred embodiments of the present invention in conjunction with the drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.

[0018] In the description of the present invention, it should be noted that the embodiments described in the present invention are part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0019] The terms "first", "second", etc. in the specification and claims of this article and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product, or equipment that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or equipment. Embodiment 1

[0020] This embodiment provides a model training method, as Figure 1 shown, including: Data preprocessing and feature selection: perform multi-dimensional cleaning on the original data to remove noise, duplicate, and incomplete data, and use a feature importance evaluation algorithm to screen out a few feature subsets that have a key impact on the model training results, so as to reduce the amount of data to be processed in subsequent model training; Distributed dataset partitioning and caching strategy: The filtered dataset is partitioned into multiple subsets according to the distributed principle, and each subset is assigned to a different computing node for processing. An intelligent caching mechanism is introduced to locally store some commonly used data subsets on the computing node to reduce the frequent loading of data from the storage device to the computing unit and improve the data access speed; Construction of a mixed-precision training framework: During the model training process, mixed-precision data representation is adopted, that is, a low-precision data type is preferentially used during the calculation process, and a high-precision data type is used for some key calculation steps with extremely high precision requirements; The parameters of mixed-precision calculation are automatically adapted to different hardware platforms to maximize the use of hardware computing power and reduce computing power consumption without affecting the convergence of the model; Adaptive learning rate dynamic adjustment: According to information such as the loss function value and gradient change during the model training process, the training state of the model is monitored in real time, and an adaptive learning rate algorithm is adopted to dynamically adjust the learning rate according to the training state. A larger learning rate is used at the beginning of training to quickly converge, and the learning rate is gradually reduced when approaching convergence to improve the stability and training efficiency of the model; Model compression and pruning: During or after the model training process, pruning algorithms are used to identify and remove connections or neurons in the model that contribute less to the prediction results, reducing the complexity and computational amount of the model; At the same time, the model quantization method is used to convert the model parameters from high-precision representation to low-precision representation to further reduce the model storage requirements and computational overhead; Optimization process control based on reinforcement learning: Build a reinforcement learning agent, take each decision point in the model training process as the action space of the agent, and the agent learns the optimal decision sequence in different training states through interaction with the training environment, automatically optimizing the entire model training process and improving the training efficiency and computing power utilization rate.

[0021] As an improved solution, the multi-dimensional cleaning includes: separately processing various situations such as numerical anomalies, null values, duplicate records, and format errors in the data.

[0022] As an improved solution, the feature importance evaluation algorithm includes: algorithms based on information gain, Gini index, and random forest feature importance ranking.

[0023] As an improved solution, when the filtered dataset is partitioned into multiple subsets according to the distributed principle, it is evenly distributed according to the distribution characteristics of the data and the performance parameters of the computing nodes to ensure the load balance of each computing node.

[0024] As an improved solution, during the model training process, the precision requirements of each calculation step are automatically detected, and the low-precision or high-precision data type is dynamically switched according to the detection results.

[0025] As an improved solution, the adaptive learning rate algorithm includes optimized versions of Adagrad, Adadelta, and Adam algorithms.

[0026] As an improved solution, in the model compression and pruning step, a strategy combining step-by-step pruning and global pruning is adopted. First, step-by-step pruning is performed to determine key connections and neurons, and then global pruning is carried out to further improve the model compression effect.

[0027] As an improved solution, the reinforcement learning agent adopts a deep reinforcement learning algorithm, and estimates state values and action values through a multi-layer neural network to more accurately learn the optimal decision-making strategy.

[0028] It should be noted that the examples here are only for explaining the present invention and should not limit the protection scope of the present invention. Embodiment 2

[0029] Based on the same inventive concept as the model training method described in Embodiment 1, this embodiment provides a model training system, as Figure 2 shown, including the following steps: A data preprocessing module, used to clean and perform feature selection on the original data, and screen out the key feature subset; A distributed processing module, responsible for distributing the data set in a distributed manner and implementing a data caching strategy to provide efficient data access services for different computing nodes; A mixed-precision training module, constructing a mixed-precision training framework, and automatically matching and adjusting the mixed-precision calculation parameters according to the hardware platform; A learning rate adjustment module, adjusting the learning rate according to the model training status by using the adaptive learning rate algorithm to optimize the model training process; A model compression module, performing model pruning and quantization operations to reduce the complexity and computational overhead of the model; A reinforcement learning optimization module, including a reinforcement learning agent, which learns and executes the optimal training process optimization strategy by interacting with the training environment.

[0030] As an improved solution, the model training system further includes: A communication module, used to achieve efficient data transmission and instruction interaction between the data preprocessing module, the distributed processing module, the mixed-precision training module, the learning rate adjustment module, the model compression module, and the reinforcement learning optimization module to ensure the collaborative work and stability of the entire training system.

[0031] It should be understood that in various embodiments herein, the sequence numbers of the above processes do not imply the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments herein.

[0032] It should also be understood that in the embodiments herein, the term "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this text generally represents an "or" relationship between the associated objects before and after.

[0033] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this article.

[0034] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific logical processes of the above-described methods can refer to the corresponding working processes of the systems, devices, and units in the foregoing method embodiments, and will not be elaborated herein.

[0035] In the several embodiments provided herein, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection to each other can be an indirect coupling or communication connection through some interfaces, devices, or units, and can also be in an electrical, mechanical, or other form of connection.

[0036] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiments herein.

[0037] In addition, each functional unit in the various embodiments of this document may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0038] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, the essence of the technical solution in this document, or the part that contributes to the prior art, or all or part of this technical solution may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this document. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0039] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A model training method, characterized in that: It includes the following steps: Data preprocessing and feature selection: Perform multi-dimensional cleaning on the original data to remove noise, duplicate, and incomplete data, and use a feature importance evaluation algorithm to screen out a few feature subsets that have a key impact on the model training results, so as to reduce the amount of data to be processed in subsequent model training; Distributed dataset partitioning and caching strategy: Partition the filtered dataset into multiple subsets according to the distributed principle, and each subset is assigned to a different computing node for processing. At the same time, an intelligent caching mechanism is introduced to locally store some commonly used data subsets on the computing node to reduce the frequent loading of data from the storage device to the computing unit and improve the data access speed; Construction of a mixed-precision training framework: During the model training process, adopt a mixed-precision data representation, that is, preferentially use a low-precision data type during the calculation process, and use a high-precision data type for some key calculation steps with extremely high precision requirements; Automatically adapt the parameters of the mixed-precision calculation for different hardware platforms to make the most of the hardware computing power without affecting the model convergence and reduce the computing power consumption; Adaptive learning rate dynamic adjustment: Monitor the training status of the model in real time according to information such as the loss function value and gradient change during the model training process, and use an adaptive learning rate algorithm to dynamically adjust the learning rate according to the training status. Use a larger learning rate at the beginning of training to quickly converge, and gradually reduce the learning rate when approaching convergence to improve the stability and training efficiency of the model; Model compression and pruning: During or after the model training process, use a pruning algorithm to identify and remove the connections or neurons in the model that contribute less to the prediction results, reducing the complexity and computational amount of the model; At the same time, adopt a model quantization method to convert the model parameters from a high-precision representation to a low-precision representation, further reducing the model storage requirements and computational overhead; Optimization process control based on reinforcement learning: Build a reinforcement learning agent, use each decision point in the model training process as the action space of the agent, and the agent learns the optimal decision sequence under different training states through interaction with the training environment, automatically optimizing the entire model training process and improving the training efficiency and computing power utilization rate.

2. The model training method according to claim 1, wherein: The multi-dimensional cleaning includes: separately processing various situations such as numerical anomalies, null values, duplicate records, and format errors in the data.

3. The model training method according to claim 1, wherein: The feature importance evaluation algorithm includes: algorithms based on information gain, Gini index, and random forest feature importance ranking.

4. The model training method according to claim 1, wherein: When partitioning the filtered dataset into multiple subsets according to the distributed principle, perform balanced allocation based on the distribution characteristics of the data and the performance parameters of the computing nodes to ensure the load balance of each computing node.

5. The model training method according to claim 1, wherein: During the model training process, automatically detect the precision requirements of each calculation step and dynamically switch to use a low-precision or high-precision data type according to the detection results.

6. The model training method according to claim 1, wherein: The adaptive learning rate algorithm includes optimized versions of Adagrad, Adadelta, and Adam algorithms.

7. The model training method according to claim 1, wherein: In the model compression and pruning step, a strategy combining step-by-step pruning and global pruning is adopted. First, step-by-step pruning is performed to determine key connections and neurons, and then global pruning is performed to further improve the model compression effect.

8. The model training method according to claim 1, wherein: The reinforcement learning agent adopts a deep reinforcement learning algorithm to estimate state values and action values through a multi-layer neural network to more accurately learn the optimal decision-making strategy.

9. A model training system, characterized in that: It includes: A data preprocessing module for cleaning and feature selection of the original data to screen out a key feature subset; A distributed processing module responsible for distributed partitioning of the data set and implementing a data caching strategy to provide efficient data access services for different computing nodes; A mixed-precision training module that constructs a mixed-precision training framework and automatically matches and adjusts mixed-precision calculation parameters according to the hardware platform; A learning rate adjustment module that adjusts the learning rate according to the model training status using an adaptive learning rate algorithm to optimize the model training process; A model compression module that performs model pruning and quantization operations to reduce the complexity and computational overhead of the model; A reinforcement learning optimization module that includes a reinforcement learning agent and learns and executes an optimal training process optimization strategy by interacting with the training environment.

10. The model training system according to claim 9, wherein: The model training system further includes: A communication module for realizing efficient data transmission and instruction interaction among the data preprocessing module, the distributed processing module, the mixed-precision training module, the learning rate adjustment module, the model compression module, and the reinforcement learning optimization module to ensure the collaborative work and stability of the entire training system.