Automatic learning engine device based on packaging large model training platform
Through the automatic learning engine device based on the encapsulated large model training platform, the unified YAML configuration template and multiple hardware device support is used, combined with the parallel computing framework and the Master-Driver-Work architecture, the complexity and resource consumption of large model fine-tuning deployment solutions are solved, and efficient and reliable model training and deployment are achieved.
Patent Information
- Application Number
- CN202510543460.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The existing large model fine-tuning deployment solutions are complex and resource consumption, resulting in long development cycles, duplicate codes, and difficulty in maintaining. It lacks a unified standard configuration method, making it difficult to achieve smooth switching between models, poor scheduling flexibility in training resource, and lack of systematic solutions for exception handling and status monitoring.
An automatic learning engine device based on the encapsulated large model training platform is proposed. The training parameters are defined through a unified YAML configuration template, and a variety of hardware devices such as CPU, GPU, and NPU are supported. Combined with parallel computing frameworks such as Accelerate and DeepSpeed, the Master-Driver-Work three-layer architecture is used to implement task allocation, status monitoring and exception handling, and automatically select the optimal parallel computing strategy.
It significantly simplifies the development workflow, improves training efficiency and resource utilization, lowers the technical threshold, and allows non-professional personnel to complete complex model training tasks, improves the reliability and fault tolerance of the system, and realizes reliable evaluation of model quality and smooth expansion of the system.
Smart Images

Figure CN120066523A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to an automatic learning engine device based on an encapsulated large model training platform. Background Art
[0002] In the current field of artificial intelligence, especially with the rapid development of natural language processing (NLP) technology, large language models (LLMs) have become the core engines driving technological innovation and industrial transformation. With the emergence of open-source or semi-open-source large language models such as GPT, LLaMA, and Qwen, these models with powerful language understanding and generation capabilities are being widely applied in various industries. However, these pre-trained models are often general-purpose and need to be fine-tuned for specific domains or tasks to achieve optimal performance. The complexity, resource consumption, and technical threshold of the fine-tuning process have become key factors restricting the widespread implementation of large models. Various open-source large models in the market have unique interface designs, parameter configurations, and training requirements, and this diversity makes it extremely difficult to efficiently access and adapt these models in an enterprise's private environment. At the same time, the computing resources required for large model training are often extremely large, and how to achieve efficient training under limited hardware conditions is also an urgent problem to be solved.
[0003] Traditional large model fine-tuning deployment solutions usually adopt the method of developing and adapting specialized code for a single model. Technical teams need to deeply study the architecture characteristics, interface designs, and training methods of each model, and write a large amount of customized code to implement model loading, training, and inference. This method not only leads to a long development cycle but also causes problems such as code duplication and difficult maintenance. When new models need to be supported or existing models need to be updated, it is often necessary to rewrite or significantly modify the code. The lack of a unified standard configuration method makes it difficult to achieve smooth switching between models. In terms of training resource scheduling, existing solutions generally lack flexible multi-device support capabilities and are difficult to automatically select the optimal parallel computing strategy according to task characteristics and hardware conditions. In addition, there are also no systematic and standardized solutions for key links such as exception handling, status monitoring, and result evaluation during the fine-tuning training process, relying on a large amount of manual intervention and empirical judgment.
[0004] These technical problems have caused serious consequences: First, the application threshold of large model technology remains high, and small and medium-sized enterprises and research institutions are difficult to bear the high technology R & D costs, resulting in the hindrance of the popularization of artificial intelligence technology; Second, the development efficiency is low and the maintenance cost is high. The technical team needs to invest a lot of time and energy in repetitive engineering adaptation work, rather than business innovation; Third, computing resources cannot be fully utilized. The low training efficiency not only prolongs the model development cycle, but also increases the hardware investment cost; Fourth, the model quality is difficult to guarantee. The lack of a standardized evaluation process leads to unstable fine-tuning effects; Finally, the system reliability is insufficient, and the frequent failure of training tasks reduces the success rate of large model project implementation. These problems together constitute a bottleneck restricting the in-depth application and industrialization of artificial intelligence technology. Summary of the Invention
[0005] To overcome the deficiencies of the prior art, the present invention proposes an automatic learning engine device based on an encapsulated large model training platform. By defining training parameters through a unified YAML configuration template, the access process of different large models is standardized, enabling developers to avoid writing specialized code for each model and significantly simplifying the development workflow.
[0006] To achieve the above object, the present invention proposes an automatic learning engine device based on an encapsulated large model training platform, including: A unified access specification module, which standardizes the training and inference parameter configuration through a YAML configuration template file. The YAML configuration template file follows the naming rule of "ai.config.train.<model name encoding>.<scenario name>.yaml", where the model name encoding is the encoding corresponding to the large model selected by the user, and the scenario name is fine-tuning (sft), pre-training (pt), or reward model training (rm); A multi-device support and parallel computing framework module, which is compatible with three hardware devices: CPU, GPU, and NPU, and supports two parallel computing frameworks: Accelerate and DeepSpeed. It automatically generates corresponding training startup commands according to the selected running type and parallel framework by the user; A training engine module, which adopts a three-layer architecture of Master-Driver-Work to implement task allocation, status monitoring, and exception handling. Among them, the Master is responsible for receiving task messages and starting containers, the Driver is responsible for listening to the Work status and writing back results, and the Work is responsible for executing fine-tuning training tasks; A training algorithm framework module, which is used to parse the ai.envs.train.yaml running configuration file, perform dataset splitting, data format conversion, support multiple fine-tuning training methods, and evaluate the model.
[0007] Furthermore, the YAML configuration template file in the unified access specification module includes the following five parts: The runtime part is used to describe the algorithm running environment configuration, specifically including: The type field indicates the running type, which can be single machine single card (SMSG), single machine multi card (SMMG), or multi machine multi card (MMMG); The parallel_framework field indicates the parallel computing framework, which can be accelerate or deepspeed; The workspace field indicates the working directory path; The cmd field indicates the algorithm running command list; The options part is used to describe the algorithm running hyperparameters, including multiple hyperparameter objects, and each object has: The code field indicates the hyperparameter code; The name field indicates the hyperparameter display name; The type field indicates the hyperparameter data type; The display field indicates whether the hyperparameter is displayed; The default field indicates the hyperparameter default value; The desc field indicates the hyperparameter description; The rule field is used to define the enumerated values or value ranges of the hyperparameters; The inputs part is used to define the input resources required by the algorithm, including training data, evaluation data, and the base model. Each input resource has: The name field indicates the resource name; The code field indicates the resource encoding; The oid field indicates the original encoding / ID; The type field indicates the resource type, and the value is dataset, datasource, or model; The label field indicates the resource label; The accessType field indicates the access type, and the value is local or remote; The uri field indicates the uniform resource identifier; The evalScale field indicates the evaluation data ratio; The "outputs" section is used to define the output results during the training process, including the model (main), TensorBoard visualization results, checkpoint, and task_result. Each output item has fields such as name, code, oid, type, accessType, and uri; The "logs" section is used to define the storage information of the logs, including fields such as name, code, type, accessType, and uri.
[0008] Furthermore, the unified access specification module also includes a YAML configuration template file parser, which is specifically used for: Reading the YAML file and parsing it into a nested dictionary and list structure through the yaml library in Python, and parsing the five sections of runtime, options, inputs, outputs, and logs layer by layer; Verifying the data type and value range of each parameter to ensure that the parameters meet the expectations, including ensuring that runtime.type must be one of SMSG, SMMG, or MMMG, and parallel_framework must be accelerate or deepspeed; Automatically filling default values for parameters not explicitly specified in the YAML file, including the default values of hyperparameters in options; When the user creates a fine-tuning task, according to the model and scenario selected by the user, read the corresponding YAML template file, and dynamically generate the runtime YAML file ai.envs.train.yaml, and merge the user-defined hyperparameters with other configuration items in the template file; Ensure that the generated runtime YAML file is correctly loaded into the container internally by mounting for use by the fine-tuning algorithm when the task starts.
[0009] Furthermore, the multi-device support and parallel computing framework module specifically includes: Device adaptation component, which is used to build corresponding container images for three different hardware devices of CPU, GPU, and NPU, and configure the image address in the YAML template file to ensure efficient operation in different running environments; Parallel computing support component, specifically supporting: Accelerate framework, which provides four functions: device management, mixed-precision training, distributed training, and gradient accumulation; DeepSpeed framework, which provides six functions: device management, mixed-precision training, gradient accumulation, zero-redundancy optimizer (ZeRO), model parallelism, and checkpoint; A running type selector for automatically constructing corresponding training startup commands according to the running type and parallel framework selected by the user: When single machine single card (SMSG) is selected, directly use Python to run the fine-tuning task; When single machine multi card (SMMG) or multi machine multi card (MMMG) is selected, use accelerate or deepspeed to construct the startup command according to the parallel_framework parameter.
[0010] Furthermore, the Master-Driver-Work three-layer architecture of the training engine module specifically includes: The Master component is specifically responsible for: Receiving task messages from the management end, adding the task ID to the Redis message queue for consumption; Querying the task configuration details according to the task ID during consumption and generating a Kubernetes standard yaml file; Invoking the Kubernetes API to create Driver containers and Work containers; Receiving the service registration event of the Driver and setting up the startup listening event according to the configured time interval and timeout; Monitoring whether the Driver is alive and restarting the Driver when the listening event times out; Destroying the Driver and Work containers to release resources after the task is completed; The Driver component is specifically responsible for: Invoking the Master service registration API for registration after successful startup; Receiving the service registration event of the Work and starting the listening event to monitor whether the Work is alive; When a certain Work fails, reassigning tasks and starting a new Work; Receiving the training completion notification sent by the Work and writing back the training results; The Work component is specifically responsible for: Invoking the Driver service registration API for registration after successful startup; Dynamically loading the dataset, large model, and configuration file through mounting; Executing the fine-tuning training task; Invoking the Driver-side event listening API to write back the training results after the training is completed.
[0011] Furthermore, the task communication mechanism of the training engine module is specifically implemented as: Task allocation mechanism: After the user starts a task, the backend calls the task startup API interface of the Master and adds the task ID to the Redis message queue; The Master obtains the task ID from the message queue, queries the task configuration details, and generates a Kubernetes standard yaml file; The Master first calls the Kubernetes API to create a Driver container. After the Driver container starts successfully, it creates a Work container; Container creation uses the Kubernetes API and is compatible with existing container orchestration systems; Status monitoring mechanism: After the Driver container starts successfully, it calls the Master service registration API for registration. After the Master receives the registration event, it starts listening for events; After the Work container starts successfully, it calls the Driver service registration API for registration. After the Driver receives the registration event, it starts listening for events; The Master detects whether the Driver is alive by periodically sending heartbeat packets. If there is no response after the timeout, it is considered that the Driver has crashed and is restarted; The Driver detects whether the Work is alive by periodically sending heartbeat packets. If there is no response after the timeout, it is considered that the Work has crashed and is restarted; Result writing back mechanism: After the Work completes the training, it calls the event listening API provided by the Driver to write back the training results; After the Driver receives the training results, it saves the results and updates the task status; The Master periodically checks the task status. When all Work is completed, it marks the task as completed and releases resources.
[0012] Furthermore, the exception handling mechanism of the training engine module specifically includes: Task failure retry mechanism: When the fine-tuning training task fails, the system automatically retries the task. The number of retries is specified by the retry_count parameter in the YAML configuration file; The retry interval is specified by the retry_interval parameter in the YAML configuration file, and the unit is seconds; If the task still fails after the configured number of retries, the system marks the task as failed and records the error log; Container health check and automatic restart mechanism: The Master and Driver containers perform health checks regularly, and the frequency of the health checks is specified by the health_check_interval parameter; If the container does not respond within the specified time (specified by the health_check_timeout parameter), the system will automatically restart the container; When restarting the container, the system will retain the configuration and status information of the original container to ensure that the restarted container can continue to execute the original tasks; Task status monitoring and recovery mechanism: The Master monitors the status of the Driver and Work containers in real time and triggers the corresponding recovery process when an anomaly is detected; If a Work container fails, the Driver will reassign tasks and start a new Work container to take over the unfinished training; If the Driver crashes, after the Master detects it, it will start a new Driver container, and the new Driver will re-register with the Master and take over the unfinished tasks; If the Master crashes, the system will automatically start a new Master container and resume the task execution status from the most recent checkpoint; All recovery operations are based on the most recent checkpoint to ensure that the task can continue to execute from the breakpoint and reduce repeated calculations.
[0013] Furthermore, the training algorithm framework module specifically includes: A configuration parsing component for parsing the running configuration ai.envs.train.yaml file, extracting the values of each parameter, and converting them into the parameter format required for Llama-Factory fine-tuning, including three hyperparameters: learning rate, batch size, and number of training epochs; A dataset splitting component for randomly splitting the dataset, and the specific implementation is as follows: Determine the ratio of the training set and the evaluation set according to the value of the evalScale parameter in the ai.envs.train.yaml file; Use a random splitting algorithm to ensure the uniformity of data distribution and prevent data bias; Save the split dataset as two files: the training set and the evaluation set respectively; A data format conversion component for converting datasets in various formats into the Llama-Factory standard format, and the specific implementation is as follows: Parse the original data format and extract the key fields; Reorganize the data according to the format required by Llama-Factory; Ensure that the converted data is fully compatible with Llama-Factory; The fine-tuning training component supports three fine-tuning training methods, which are specifically determined by the train_method parameter in the ai.envs.train.yaml file; The evaluation component is used to evaluate the model and outputs four evaluation metrics; The intermediate process monitoring component is used to monitor the task iteration progress in real time and write the real-time results to the path specified by output.task_result in the ai.envs.train.yaml file in JSON file format, and supports the client to receive and display the real-time training status through the WebSocket protocol.
[0014] Furthermore, the fine-tuning training component specifically supports the following three fine-tuning methods: The LoRA fine-tuning method, and its specific implementation is as follows: Approximate the update of the model weights through low-rank decomposition, and add a pair of trainable low-rank decomposition matrices to each weight matrix that needs to be updated; Assume the original weight matrix is W, with dimension d×k, introduce matrices A (dimension d×r) and B (dimension r×k) with rank r, and the fine-tuned weight matrix becomes W + AB; During the training process, only the two low-rank matrices A and B are trained, and the original weight matrix W remains fixed; This method is especially suitable for scenarios with limited computing resources, suitable for fine-tuning on consumer-grade GPUs, and suitable for scenarios where different fine-tuning directions need to be quickly tried; The full-scale fine-tuning method, and its specific implementation is as follows: All parameters of the pre-trained model are regarded as trainable parameters, including three types of parameters: the embedding layer, the multi-head attention layer, and the feed-forward neural network layer; Use the backpropagation algorithm to update each parameter in the model, and the update step size is controlled by the optimization algorithm (Adam or SGD algorithms) according to the learning rate; This method is suitable for scenarios with a large amount of high-quality labeled data and sufficient computing resources, and can maximize the adaptation of the model to specific tasks; The freezing fine-tuning method, and its specific implementation is as follows: Only adjust a part of the layers in the pre-trained model, while the parameters of other layers remain fixed; Select to freeze the early layers close to the input end and only fine-tune the later layers close to the output end; This method is suitable for scenarios where the task is similar to the pre-training task but has certain differences, and can adapt to specific task requirements while retaining the pre-training knowledge; The fine-tuning method automatic mechanism, and its specific details are as follows: Select a fine-tuning method according to the user's task requirements, dataset size, and hardware resources; When computing resources are limited or when you need to quickly try different fine-tuning directions, use LoRA fine-tuning; When there is a large amount of high-quality labeled data and sufficient computing resources, use full fine-tuning; When the task is similar to the pre-training task but has certain differences, use freeze fine-tuning; 1. Recommended conditions for LoRA fine-tuning (Low-Rank Adaptation) Applicable scenarios: Limited computing resources or the need for rapid iterative experiments.
[0015] Judgment conditions: Hardware resources: GPU video memory ≤ 32GB Available training time < 4 hours (need to quickly verify the effect).
[0016] Data scale: The amount of labeled data ≤ 100,000.
[0017] The data quality is medium or there is noise (LoRA is less sensitive to noise).
[0018] Task requirements: The difference between the task and the pre-training task is medium (e.g., text classification → sentiment analysis).
[0019] 2. Recommended conditions for full fine-tuning (Full Fine-Tuning) Applicable scenarios: Sufficient resources and high data quality.
[0020] Judgment conditions: Hardware resources: GPU video memory > 32GB (such as A100, H100).
[0021] Available training time > 4 hours.
[0022] Data scale: The amount of labeled data > 500,000 and the labeling consistency is high (e.g., the manual review pass rate ≥ 95%).
[0023] Task requirements: The difference between the task and the pre-training task is significant (e.g., pre-training is for general text → downstream is medical entity recognition).
[0024] All model parameters need to be adjusted to maximize performance.
[0025] 3. Recommended conditions for freeze fine-tuning (Freeze-Tuning) Applicable scenarios: The task is similar to the pre-training task but requires minor adaptation.
[0026] Judgment conditions: Hardware resources: GPU video memory ≤ 24GB.
[0027] Available training time < 1 hour.
[0028] Data scale: Annotated data volume ≤ 50,000 (small sample adaptation).
[0029] Task requirements: The domain overlap between the pre-training task and the downstream task ≥ 70% (e.g., BERT pre-training → news classification) Users can manually select the fine-tuning method when creating a fine-tuning task, or the system can automatically select according to the rules.
[0030] Furthermore, the evaluation component is specifically used to calculate the following four evaluation metrics: Perplexity, precision, recall, and loss value; The evaluation process is specifically implemented as follows: Use the number of iterations (num_train_epochs) specified in the options of the ai.envs.train.yaml file to train the model; After training is completed, save the model to the path specified by outputs.main in the ai.envs.train.yaml file; When evaluating, read the model under this path and use the split test set for evaluation; Write the evaluation results to the path specified by output.task_result in the ai.envs.train.yaml file in JSON file format.
[0031] Furthermore, the intermediate process monitoring component is specifically implemented as follows: Adopt the callback function mechanism to trigger the callback function at the end of each iteration step of model training; The callback function collects training information in real time, including three parameters: the current iteration number, training loss value, and learning rate; Write the collected information to the path specified by output.task_result in the ai.envs.train.yaml file in JSON format; The JSON file structure contains five fields: task ID, current iteration number, total iteration number, training loss value, and learning rate; The backend service uses the file monitoring mechanism to detect changes in the JSON file and reads the latest content when the file is updated; The backend service pushes the latest training status to the connected clients via the WebSocket protocol; After receiving the pushed data, the client updates two visualization components, namely the training progress bar and the loss curve graph, in real time, providing an intuitive training process monitoring experience.
[0032] Furthermore, the implementation process of the device includes the following steps: The user selects a basic large model and a training scenario in the management interface, and the system reads the corresponding YAML configuration template file according to the selection; The system parses the YAML configuration template file, extracts the options part and displays it on the interface, and the user can adjust the hyperparameters as needed; The user uploads or selects an existing training dataset and specifies the evaluation data ratio; After the user submits the task, the system generates a running-state YAML file ai.envs.train.yaml, and sends it together with the training dataset and the basic model to the Master via the start task API; The Master adds the task to the queue and starts the Driver and Work containers to execute the training task; During the training process, the system monitors the training status in real time and pushes the status to the client via the WebSocket protocol; After the training is completed, the system evaluates the model, outputs the evaluation metrics, and saves the trained model to the specified path; The user views the training results, evaluation metrics, and visualization charts of the training process in the management interface, and can also download the trained model for inference.
[0033] Compared with the prior art, the beneficial effects of the present invention are: 1. The present invention provides an automatic learning engine device based on an encapsulated large model training platform, which supports various hardware devices such as CPUs, GPUs, and NPUs, and combines parallel computing frameworks such as Accelerate and DeepSpeed, enabling efficient utilization of computing resources in different scenarios such as single machine with single card, single machine with multiple cards, and multiple machines with multiple cards.
[0034] 2. The present invention provides an automatic learning engine device based on an encapsulated large model training platform. The Master-Driver-Work architecture provides a perfect exception handling mechanism, including task failure retry, container health check and automatic restart, and task status monitoring and recovery, greatly improving the reliability and fault tolerance of the system.
[0035] 3. The present invention provides an automatic learning engine device based on an encapsulated large model training platform, which automatically selects the most suitable fine-tuning method (LoRA, full-scale or frozen fine-tuning) according to task requirements, dataset scale, and hardware resources, avoiding the uncertainty of manual selection and improving the training effect.
[0036] 4. The present invention provides an automatic learning engine device based on an encapsulated large model training platform. Through functions such as configuration parsing, automatic splitting of datasets, and automatic conversion of data formats, manual intervention is reduced, and the full process cycle from data preparation to model deployment is significantly shortened; the overall solution lowers the technical threshold for fine-tuning large models, enabling non-professionals to complete complex model training tasks through simple configuration, promoting the popularization and application of large model technology.
[0037] 5. The present invention provides an automatic learning engine device based on an encapsulated large model training platform. It uses a callback function method to monitor the training progress in real time and timely feedback on state changes, enabling users to always keep track of the training situation and adjust strategies in advance when necessary. A standardized model performance evaluation system is established through standard metrics such as perplexity, precision, recall, and loss value, facilitating comparison and evaluation between different models. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 It is a schematic diagram of the device process; Figure 2 It is a flowchart of the engine training; Figure 3 It is an effect diagram of the interface display. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] The following will more clearly and completely elaborate on the technical solutions of the present invention by describing the preferred embodiments of the present invention in conjunction with the drawings.
[0041] Term Explanation: Technical Term Explanation of the Large Model Training Platform Llama-Factory: An open-source large model training toolbox; YAML Configuration: A human-readable data serialization format used in this system to define training parameters, environment configurations, and resource allocations; Fine-tuning (SFT): Further training the pre-trained model with task-specific data to adapt the model to a specific domain or task; Pre-training (PT): The process of initial training of the model on a large-scale general corpus to learn the basic laws and knowledge of the language; Reward Model Training (RM): The model training process used to evaluate the quality of generated content in reinforcement learning; LoRA Fine-tuning: A parameter-efficient fine-tuning method; Accelerate: A lightweight distributed training library; DeepSpeed: A deep learning optimization library; Mixed Precision Training: Calculating using different numerical precisions simultaneously to balance training speed and precision; Gradient Accumulation: Accumulating gradients over multiple small batches and then updating the model; Zero Redundancy Optimizer (ZeRO): A key technology in DeepSpeed that reduces memory usage through hierarchical optimization; Master-Driver-Work Architecture: A hierarchical task management architecture where the Master is responsible for global scheduling, the Driver monitors the work progress, and the Work performs specific computational tasks; Kubernetes: A container orchestration platform for automating the deployment, scaling, and management of containerized applications; MLflow / Kubeflow: MLOps tools for managing machine learning experiments and model deployment; Perplexity: A metric for evaluating the prediction ability of a language model, where a lower value indicates more accurate model predictions; Precision: The proportion of samples predicted as positive classes that are actually positive classes, reflecting the prediction accuracy; Recall: The proportion of samples that are actually positive classes and are correctly identified, reflecting the model's capture ability; Loss Value: A numerical value quantifying the degree of model prediction error, usually calculated using the cross-entropy loss function.
[0042] Such as Figure 1As shown in the figure, the top layer is the MaaS platform, which serves as the control center of the overall system. It distributes tasks to the Master component through the task message mechanism. After receiving the tasks, the Master component is responsible for starting the Driver container. Subsequently, the Driver container starts the Work containers, forming a container orchestration chain. As the executor of the actual workload, the Work containers process various computing tasks. It will execute the specific business logic stored in the code framework, configuration framework, or large model configuration file. After the task is completed, a completion signal is sent. At the same time, the system has a perfect resource release mechanism, which is represented by the curve feedback path on the right, ensuring that the Driver container and Work containers can return the computing resources to the resource pool in a timely manner after completing their respective responsibilities, avoiding resource waste, and forming an efficient closed-loop of task processing and resource management overall.
[0043] As a specific implementation manner, the specific implementation manner of the automatic learning engine device for encapsulating the large model training platform based on Llama-Factory The present invention provides an automatic learning engine device for encapsulating a large model training platform based on Llama-Factory. The device mainly includes four main parts: a unified access specification module, a multi-device support and parallel computing framework module, a training engine module, and a training algorithm framework module, which can realize the efficient management, deployment, and execution of large model training.
[0044] In this implementation manner, the YAML configuration template file is the core configuration unit of the system and is used to standardize the parameter configuration for training and inference. These files follow specific naming rules: ai.config.train.<model name encoding>.<scenario name>.yaml. Among them, the model name encoding is the encoding corresponding to the large model selected by the user (for example, the encoding corresponding to the Qianwen 7B large model is Qwen-7B-Chat), and the scenario name is fine-tuning (sft), pre-training (pt), or reward model training (rm).
[0045] For example, for the fine-tuning scenario of the Qianwen 7B large model, the template file name is ai.config.train.Qwen-7B-Chat.sft.yaml. The function of this file is: when the user adds a new fine-tuning task and selects the Qianwen 7B base model, the system backend will read this yaml template file according to the model encoding and scenario, and parse out the hyperparameters (options) part to be displayed on the front-end page for the user to customize and adjust the parameters. When the task starts, except for the hyperparameters (options) part that obtains the user-defined configuration, other configuration items are obtained from the yaml template, and a runtime ai.envs.train.yaml file is generated for the task to run.
[0046] The YAML configuration template file is mainly divided into five parts: the runtime part, the options part, the inputs part, the outputs part, and the logs part. The runtime part is used to describe the algorithm runtime environment configuration, including the type field (type) indicating the runtime type, which can be single machine single card (SMSG), single machine multi card (SMMG), or multi machine multi card (MMMG); the parallel computing framework field (parallel_framework) indicating the parallel computing framework, which can be accelerate or deepspeed; the workspace field (workspace) indicating the working directory path; and the command field (cmd) indicating the algorithm runtime command list. The options part is used to describe the algorithm runtime hyperparameters, containing multiple hyperparameter objects, each object having a code field (code), a name field (name), a type field (type), a display field (display), a default value field (default), a description field (desc), and a rule field (rule) for defining the enumeration values or value ranges of the hyperparameters. The inputs part is used to define the input resources required by the algorithm, including training data, evaluation data, and the base model. Each input has a name field (name), a code field (code), an original ID field (oid), a type field (type) indicating the resource type such as dataset, datasource, or model, a label field (label), an access type field (accessType) indicating local or remote access, a uniform resource identifier field (uri), and an evaluation ratio field (evalScale) for training data. The outputs part is used to define the output results during the training process, such as the model (main), the tensorboard visualization results, the checkpoint, and the task_result. Each output item contains the name, code, oid, type, accessType, and uri fields. The logs part is used to define the storage information of the logs, including the name, code, type, accessType, and uri fields.
[0047] The parsing of the YAML configuration template file is implemented through a custom parser, which can parse layer by layer according to the hierarchical structure of the YAML file and convert it into an internal data structure. The system first reads the YAML file and uses the Python yaml library to load it as a dictionary object. The parser parses the five parts of runtime, options, inputs, outputs, and logs layer by layer according to the hierarchical structure of the YAML file. The parser validates each parameter to ensure that its data type and value range meet the expectations. If some parameters are not explicitly specified in the YAML file, the parser will automatically fill in the default values. When the user creates a fine-tuning task, the system will read the corresponding YAML template file according to the selected model and scenario, and dynamically generate the runtime YAML file ai.envs.train.yaml, merging the user-defined hyperparameters with other configuration items in the template file.
[0048] When the user creates a fine-tuning task, the system will read the corresponding YAML template file according to the selected large model and scenario, and dynamically display the hyperparameter configuration (options section). The user can modify the hyperparameters as needed. When saving the task, the system will generate the runtime YAML file ai.envs.train.yaml. This file is dynamically mounted inside the container for the fine-tuning algorithm to use when the task is started. The runtime YAML file contains the same five parts as the template file (runtime, options, inputs, outputs, and logs), but the options section will contain the user-defined parameter values.
[0049] This embodiment is fully compatible with three types of hardware devices, including CPU, GPU, and NPU, ensuring efficient operation in different operating environments. Corresponding images are built according to different hardware environments, and the image addresses are configured in the yaml template file. At the same time, it supports two parallel computing frameworks: Accelerate and DeepSpeed, to greatly improve the training efficiency. The Accelerate framework provides device management functions (automatically manages device allocation, supports CPU, GPU, and NPU), mixed precision functions (supports automatic mixed precision training, reduces memory occupancy and accelerates training), distributed training functions (supports multi-GPU and multi-node distributed training, automatically processes data parallelism and model parallelism), and gradient accumulation functions (supports gradient accumulation, can train larger models with limited GPU memory). The DeepSpeed framework provides device management functions (supports multi-GPU and multi-node distributed training, automatically manages device allocation), mixed precision functions (supports automatic mixed precision training, reduces memory occupancy and accelerates training), gradient accumulation functions (supports gradient accumulation, can train larger models with limited GPU memory), zero redundant optimizer function (ZeRO) (through hierarchical optimization technology, reduces memory occupancy, supports training of larger-scale models), model parallelism functions (supports model parallelism, can split large models for training on multiple GPUs), and checkpoint functions (supports automatic saving and restoring of training status, facilitating interrupted and resumed training).
[0050] In this embodiment, when the user selects single machine and single card, the runtime.type parameter in the ai.envs.train.yaml file is SMSG, and the runtime.parallel_framework parameter is invalid (the option will be hidden on the page), and the fine-tuning task is directly run using Python; when the user selects single machine and multi-card or multi-machine and multi-card, the runtime.type parameter in the ai.envs.train.yaml file is SMMG or MMMG, and the runtime.parallel_framework parameter can be selected as accelerate or deepspeed. When starting the task, the startup command is built according to the parallel_framework configuration for startup.
[0051] The training engine architecture adopts a three-layer structure of Master-Driver-Work, and the responsibilities of each component are as follows: The Master is responsible for receiving task messages from the management end, then starting a Driver container and one or more Work containers according to the configuration, and at the same time monitoring the status of the Driver side. After the task is completed, the Driver and Work containers are destroyed to release resources; the Driver container is responsible for monitoring the status of the Work containers and writing back the training results; the Work containers are specifically responsible for executing the fine-tuning training tasks. The required data sets, large models, and configuration files are dynamically loaded into the containers by mounting. After the task is completed, it will call the Driver to write back the training results and status.
[0052] The specific implementation of the communication mechanism between Master, Driver, and Work is as follows: After the user starts a task, the backend calls the start task API interface of the Master, and then adds the task ID to the Redis message queue to wait for consumption; when consuming, it will query the task configuration details according to the task ID and generate Kubernetes standard yaml (including Driver yaml and Work yaml), and then call the Kubernetes API to pass in the Driver yaml to create the Driver container; after the Driver container starts successfully, it will pass in the Work yaml to create the Work container; after the Driver-side service starts successfully, it will call the Master service registration API for registration. After the Master receives the registration event, it will set up monitoring events according to the time interval and timeout time in the configuration file to monitor whether the Driver is alive; similarly, after the Work container is created successfully, it will also call the Driver-side service registration API for registration. After the Driver receives the registration event, it will also start monitoring events according to the settings in the configuration file to monitor whether the Work is alive. After the Work is registered successfully, it will start the fine-tuning task for training. After the training is completed, it will send a notification to call the Driver-side event monitoring API to write back the training results.
[0053] The exception handling mechanism of the training engine specifically includes: the task failure retry mechanism. When a certain task fails, the system will automatically retry the task. The number of retries and the interval time can be configured through the retry_count and retry_interval parameters in the YAML configuration file. If the task still fails after multiple retries, the system will mark the task as failed and record detailed error logs; container health check and automatic restart. The Master and Driver containers will perform health checks regularly. If a certain container does not respond within the specified time, the system will automatically restart the container. The frequency and timeout of the health check can be configured through the health_check_interval and health_check_timeout parameters in the YAML configuration file; task status monitoring and recovery. The Master will monitor the status of the Driver and Work containers in real time. If a certain Work container fails, the Driver will reallocate tasks and start a new Work container to take over the unfinished training. If the Master or Driver crashes, the system will automatically start a new Master or Driver container and resume task execution from the most recent checkpoint.
[0054] As Figure 2 shown, it is the flowchart of the training engine. After the system starts, the Master component receives task messages from the MaaS management terminal; then the Master conducts intelligent scheduling analysis based on task priorities and the management terminal task list; subsequently, the system will determine whether the task is schedulable, which is a key decision point. If it is not schedulable, the task will be marked as invalid and the processing will be terminated. If it is schedulable, the system will continue to execute according to the established strategy; for schedulable tasks, the system starts a dedicated Driver container and triggers the task processing logic; the Driver container will dynamically start one or more Work containers according to task requirements to achieve parallel processing of tasks; the Work containers take over and execute specific scheduling tasks to complete the actual business processing; after the task is executed, the system records the completion result and enters the resource recovery stage, destroying the unnecessary Driver and Work containers in sequence; finally, the system releases all relevant resources, marking the completion of the entire task processing cycle, forming a complete closed-loop process from task reception, scheduling, execution to resource recovery.
[0055] The configuration parsing component of the training algorithm framework can parse the running configuration ai.envs.train.yaml file and construct the parameters required for Llama-Factory fine-tuning; the dataset splitting component supports random splitting of the dataset, splitting the original fine-tuning text dataset into a training set and an evaluation set according to a ratio, and the splitting ratio is subject to the evalScale parameter in the ai.envs.train.yaml configuration file; the data format conversion component converts the dataset format into the Llama-Factory standard format to ensure data consistency and compatibility, and has no impact on the input and output formats of the model.
[0056] The fine-tuning training component supports three fine-tuning training methods, namely LoRA fine-tuning, full fine-tuning, and freezing fine-tuning. Which method to use specifically is determined by the fine-tuning training method (train_method value lora / full / freeze) selected by the user when adding a new task, and the selection result will be written into the train_method parameter of the ai.envs.train.yaml file. The core idea of LoRA fine-tuning is to approximate the update of the model weights through low-rank decomposition. On the basis of the original pre-trained model, a pair of trainable low-rank decomposition matrices are added to each weight matrix that needs to be updated. Assuming the original weight matrix is W with dimensions d×k, by introducing matrices A (dimension d×r) and B (dimension r×k) with rank r, the fine-tuned weight matrix becomes W + AB. During the training process, only the two low-rank matrices A and B are trained, while the original weight matrix W remains fixed. Full fine-tuning is the most direct fine-tuning method. On the basis of the pre-trained model, all the parameters of the model are used as trainable parameters. During the training process, according to the loss function of the task, the backpropagation algorithm is used to update every parameter in the model, including the parameters of all layers such as the embedding layer, multi-head attention layer, and feed-forward neural network layer. Freezing fine-tuning means only adjusting a part of the layers in the pre-trained model, while the parameters of other layers remain fixed. Generally speaking, the layers closer to the input end capture more general features, while the layers closer to the output end focus more on task-specific information. Therefore, it is usually chosen to freeze the early layers and only fine-tune the later layers. The system also has an automatic fine-tuning method mechanism, which automatically selects a suitable fine-tuning method according to the user's task requirements, dataset size, and hardware resources: when the user's computing resources are limited or they want to quickly try different fine-tuning directions, the system will use LoRA fine-tuning; when the user has a large amount of high-quality labeled data and the computing resources allow for long-term training of the entire model, the system will use full fine-tuning; when the user's task is relatively similar to the pre-trained task but has certain differences, the system will use freezing fine-tuning.
[0057] The evaluation component has the function of evaluating the model and can output four evaluation metrics: Perplexity measures how well a probability model predicts samples and is commonly used in NLP to evaluate the quality of language models. The calculation method is , where N is the number of words in the sentence, is the probability of the entire sequence, and the lower the ideal value, the better. Ideally, it is close to 1; Precision refers to the proportion of samples actually belonging to the positive class among all samples predicted as the positive class. The calculation method is , where TP is the number of true positives and FP is the number of false positives. A high precision means the model has high accuracy when predicting a certain class; Recall refers to the proportion of samples actually belonging to the positive class that are correctly identified. The calculation method is , where FN is the number of false negatives. A high recall indicates that the model can capture more real positive examples; Loss is a way to quantify the degree of model prediction error. In this system, the cross-entropy loss function is used for calculation, and the formula is , where N is the number of samples, is the true label (0 or 1) of sample i, is the probability that the model predicts sample i as 1. During the training process, the loss value should gradually decrease, indicating that the model is learning and improving its performance. The evaluation process is as follows: Iterative training is performed using the number of iterations (num_train_epochs) in the options of the ai.envs.train.yaml file. After training is completed, the model is written to the output path specified in the ai.envs.train.yaml file outputs.main according to the model output path. During evaluation, the output model is read and evaluated using the test set, and the evaluation results are written in JSON file format to the path specified by ai.envs.train.yaml file output.task_result.
[0058] The intermediate process monitoring component uses the callback function method to monitor the task iteration progress in real time, timely feedback the training status, and write the real-time results in JSON file format to the corresponding path according to the configuration of the ai.envs.train.yaml file output.task_result. The management client uses the WebSocket protocol to connect to the backend service. The backend service monitors the changes in the result file and pushes the data to the client for display, enabling users to monitor the training progress in real time, as Figure 3 shown.
[0059] The specific system operation process is as follows: The user selects a basic large model and a training scenario in the management interface, and the system reads the corresponding YAML configuration template file according to the selection; the system parses the YAML configuration template file, extracts the options section and displays it on the interface, and the user can adjust the hyperparameters as needed; the user uploads or selects an existing training dataset and specifies the evaluation data ratio; after the user submits the task, the system generates a running-state YAML file ai.envs.train.yaml, and sends it together with the training dataset and the basic model to the Master through the Start Task API; the Master adds the task to the queue and starts the Driver and Work containers to execute the training task; during the training process, the system monitors the training status in real time and pushes the status to the client through the WebSocket protocol; after the training is completed, the system evaluates the model, outputs four evaluation metrics: perplexity, precision, recall, and loss value, and saves the trained model to the specified path; the user can view the training results, evaluation metrics, and visualization charts of the training process in the management interface, and can also download the trained model for inference.
[0060] Through innovative designs such as unified access specifications, multi-device support and parallel computing frameworks, a three-layer Master-Driver-Work training engine, a perfect exception handling mechanism, support for multiple fine-tuning methods, and real-time monitoring, this embodiment constructs an easy-to-use, stable, and efficient large model training platform, which is applicable to a variety of large model training scenarios and greatly improves the efficiency and success rate of large model training. Compared with the existing MLOps solutions, this device has stronger flexibility, scalability, and automation, and can better meet the complex and changing large model training requirements.
[0061] The above specific embodiments only describe the preferred embodiments of the present invention, and do not limit the protection scope of the present invention. Without departing from the design concept and spirit scope of the present invention, various deformations, substitutions, and improvements made by those of ordinary skill in the art to the technical solutions of the present invention based on the written description and drawings provided by the present invention shall fall within the protection scope of the present invention. The protection scope of the present invention is determined by the claims.
Claims
1. An automatic learning engine device based on a packaged large model training platform, characterized in that: include: Unified access specification module, which standardizes training and inference parameter configuration through YAML configuration template files. The YAML configuration template files follow the naming convention of "ai.config.train.<model name encoding>.<scenario name>.yaml", where the model name encoding is the encoding corresponding to the large model selected by the user, and the scenario name is fine-tuning, pre-training, or reward model training; Multi-device support and parallel computing framework module, which is compatible with three hardware devices: CPU, GPU, and NPU, and supports two parallel computing frameworks: Accelerate and DeepSpeed. It automatically generates corresponding training start commands according to the operation type and parallel framework selected by the user; The training engine module adopts the Master-Driver-Work three-layer architecture to implement task allocation, status monitoring and exception handling, where the Master is responsible for receiving task messages and starting containers, the Driver is responsible for monitoring the Work status and writing back the results, and the Work is responsible for executing fine-tuning training tasks; The training algorithm framework module is used to parse the ai.envs.train.yaml operation configuration file, split the data set, convert the data format, support multiple fine-tuning training methods and evaluate the model.
2. The automatic learning engine device based on the encapsulated large model training platform according to claim 1 is characterized in that: The YAML configuration template file in the unified access specification module consists of the following five parts: The runtime part is used to describe the algorithm runtime environment configuration, including: The type field indicates the operation type; The parallel_framework field indicates the parallel computing framework; The workspace field indicates the working directory path; The cmd field indicates the algorithm running command list; The options section is used to describe the algorithm running hyperparameters. It contains multiple hyperparameter objects, each of which has: The code field indicates the hyperparameter code; The name field indicates the display name of the hyperparameter; The type field indicates the hyperparameter data type; The display field indicates whether the hyperparameters are displayed; The default field indicates the default value of the hyperparameter; The desc field indicates the hyperparameter description; The rule field is used to define the enumeration value or value range of the hyperparameter; The inputs section is used to define the input resources required by the algorithm, including training data, evaluation data, and basic models. Each input resource has: The name field indicates the resource name; The code field indicates the resource code; The oid field indicates the original code / ID; The type field indicates the resource type; The label field indicates the resource label; The accessType field indicates the access type, and the value is local or remote; The uri field indicates a uniform resource identifier; The evalScale field indicates the evaluation data scale; The outputs section is used to define the output results of the training process, including models, tensorboard visualization results, checkpoint checkpoints, and task_result task results. Each output item has name, code, oid, type, accessType, and uri fields; The logs section is used to define the storage information of the logs, including the name, code, type, accessType, and uri fields.
3. The automatic learning engine device based on the encapsulated large model training platform according to claim 1 is characterized in that: The unified access specification module also includes a YAML configuration template file parser, which is specifically used to: Read the YAML file and parse it into nested dictionary and list structures through Python's yaml library, parsing the runtime, options, inputs, outputs, and logs parts layer by layer; Verify the data type and value range of each parameter to ensure that the parameter meets expectations; Automatically fill in default values for parameters not explicitly specified in the YAML file, including default values for hyperparameters in options; When a user creates a fine-tuning task, the corresponding YAML template file is read according to the model and scenario selected by the user, and the running YAML file ai.envs.train.yaml is dynamically generated to merge the user-defined hyperparameters with other configuration items in the template file; Ensure that the generated runtime YAML file is correctly loaded into the container by mounting when the task is started for use in the fine-tuning algorithm.
4. The automatic learning engine device based on the encapsulated large model training platform according to claim 1, characterized in that: The multi-device support and parallel computing framework modules specifically include: The device adapter component is used to build corresponding container images for three different hardware devices: CPU, GPU, and NPU, and configure the image address in the YAML template file to ensure efficient operation in different operating environments; Parallel computing support components, specifically supporting: Accelerate framework, which provides four functions: device management, mixed precision training, distributed training, and gradient accumulation; DeepSpeed framework, which provides six functions: device management, mixed precision training, gradient accumulation, zero-redundancy optimizer, model parallelism, and checkpoints; The run type selector is used to automatically build the corresponding training launch command based on the run type and parallel framework selected by the user: When a single machine and single card is selected, use Python directly to run the fine-tuning task; When you select single machine with multiple cards or multiple machines with multiple cards, choose to use accelerate or deepspeed to build the startup command based on the parallel_framework parameter.
5. The automatic learning engine device based on the encapsulated large model training platform according to claim 1, characterized in that: The Master-Driver-Work three-layer architecture of the training engine module specifically includes: The Master component is responsible for: Receive task messages from the management end and add the task ID to the Redis message queue to wait for consumption; When consuming, query the task configuration details according to the task ID and generate a Kubernetes standard YAML file; Call the Kubernetes API to create the Driver container and the Work container; Receives the service registration event of the Driver and starts listening for events according to the configured time interval and timeout settings; Monitor whether the Driver is alive and restart the Driver when the monitoring event times out; After the task is completed, destroy the Driver and Work containers to release resources; Driver component, specifically responsible for: After successful startup, call the Master service registration API to register; Receive the service registration event of the Work and start the monitoring event to monitor whether the Work is alive; When a Work fails, reassign tasks and start a new Work; Receive the training completion notification sent by the Worker and write back the training results; Work component, specifically responsible for: After successful startup, call the Driver service registration API to register; Dynamically load data sets, large models, and configuration files through mounting; Perform fine-tuning training tasks; After the training is completed, call the Driver-side event monitoring API to write back the training results.
6. The automatic learning engine device based on the encapsulated large model training platform according to claim 5, characterized in that: The task communication mechanism of the training engine module is as follows: Task allocation mechanism: After the user starts the task, the backend calls the Master's start task API interface and adds the task ID to the Redis message queue; The Master obtains the task ID from the message queue, queries the task configuration details and generates a Kubernetes standard YAML file; The Master first calls the Kubernetes API to create a Driver container. After the Driver container is successfully started, it creates a Work container. Container creation uses the Kubernetes API and is compatible with existing container orchestration systems; Status monitoring mechanism: After the Driver container is successfully started, it calls the Master service registration API to register. After receiving the registration event, the Master starts listening for events. After the Work container is successfully started, the Driver service registration API is called to register. After receiving the registration event, the Driver starts listening for events. The Master periodically sends heartbeat packets to detect whether the Driver is alive. If there is no response after the timeout, the Driver is considered to be down and restarted; The driver periodically sends heartbeat packets to detect whether the Work is alive. If there is no response after the timeout, the Work is considered to be down and restarted; Result write-back mechanism: After the Worker completes the training, it calls the event listening API provided by the Driver to write back the training results; After receiving the training results, the Driver saves the results and updates the task status; The Master periodically checks the task status. When all Work is completed, it marks the task as completed and releases resources.
7. The automatic learning engine device based on the encapsulated large model training platform according to claim 5, characterized in that: The exception handling mechanism of the training engine module specifically includes: Task failure retry mechanism: When a fine-tuning training task fails, the system automatically retries the task. The number of retries is specified by the retry_count parameter in the YAML configuration file. The retry interval is specified by the retry_interval parameter in the YAML configuration file, in seconds; If the task still fails after the configured number of retries, the system marks the task as failed and records an error log; Container health check and automatic restart mechanism: The Master and Driver containers perform health checks regularly. The frequency of health checks is specified by the health_check_interval parameter. If the container does not respond within the specified time, the system automatically restarts the container; When you restart a container, the system will retain the configuration and status information of the original container to ensure that the restarted container can continue to execute the original task; Task status monitoring and recovery mechanism: The Master monitors the status of the Driver and Work containers in real time and triggers the corresponding recovery process when an abnormality is detected; If a Work container fails, the Driver will reallocate tasks and start a new Work container to take over the unfinished training; If the Driver crashes, the Master will detect it and start a new Driver container. The new Driver will re-register with the Master and take over the unfinished tasks. If the Master fails, the system will automatically start a new Master container and restore the task execution status from the most recent checkpoint; All recovery operations are based on the most recent checkpoint, ensuring that the task can continue from the interruption point and reducing repeated calculations.
8. The automatic learning engine device based on the encapsulated large model training platform according to claim 1, characterized in that: The training algorithm framework module specifically includes: Configuration parsing component, used to parse the run configuration ai.envs.train.yaml file, extract the parameter values and convert them into the parameter format required for Llama-Factory fine-tuning, including three hyperparameters: learning rate, batch size, and number of training rounds; The dataset splitting component is used to randomly split the dataset, specifically: Determine the ratio of the training set and the evaluation set according to the evalScale parameter value in the ai.envs.train.yaml file; Use random splitting algorithm to ensure uniformity of data distribution and prevent data bias; Save the split data set into two files: training set and evaluation set; The data format conversion component is used to convert data sets in various formats into the Llama-Factory standard format. The specific implementation is as follows: Parse the original data format and extract key fields; Reorganize the data according to the format required by Llama-Factory; Ensure that the converted data is fully compatible with Llama-Factory; Fine-tuning training component, supports three fine-tuning training methods, which are determined by the train_method parameter in the ai.envs.train.yaml file; Evaluation component, used to evaluate the model and output four evaluation indicators; The intermediate process monitoring component is used to monitor the task iteration progress in real time, and write the real-time results in JSON file format to the path specified by output.task_result in the ai.envs.train.yaml file. It supports the client to receive and display the real-time training status through the WebSocket protocol.
9. The automatic learning engine device based on the encapsulated large model training platform according to claim 8, characterized in that: The fine-tuning training component specifically supports the following three fine-tuning methods: The LoRA fine-tuning method is as follows: Approximate the update of model weights through low-rank decomposition, and add a pair of trainable low-rank decomposition matrices for each weight matrix that needs to be updated; Assume that the original weight matrix is W, with a dimension of d×k, and introduce matrices A (with a dimension of d×r) and B (with a dimension of r×k) of rank r. The fine-tuned weight matrix becomes W+AB. During the training process, only the two low-rank matrices A and B are trained, and the original weight matrix W remains fixed; The full amount of fine-tuning method is as follows: All parameters of the pre-trained model are used as trainable parameters, including three types of parameters: embedding layer, multi-head attention layer, and feedforward neural network layer; Use the back-propagation algorithm to update each parameter in the model, and the update step size is controlled by the optimization algorithm according to the learning rate; The freeze fine-tuning method is as follows: Only some of the layers in the pre-trained model are adjusted, while the parameters of other layers remain fixed; Choose to freeze early layers near the input and only fine-tune later layers near the output; The automatic mechanism of fine-tuning method is as follows: Select a fine-tuning method based on the user's task requirements, dataset size, and hardware resources; When computing resources are limited or you need to quickly try different fine-tuning directions, use LoRA fine-tuning. The criteria are as follows: Hardware resources: GPU memory ≤ 32GB; available training time < 4 hours; Data scale: the amount of annotated data is ≤ 100,000, and the data quality is medium or noisy; Task requirements: The task is moderately different from the pre-training task; When there is a large amount of high-quality labeled data and sufficient computing resources, full-scale fine-tuning is used. The criteria are as follows: Hardware resources: GPU memory > 32GB; available training time > 4 hours; Data scale: the amount of labeled data is > 500,000 and the labeling consistency is high; Task requirements: The task is significantly different from the pre-training task; When the task is similar to the pre-training task but has some differences, use frozen fine-tuning. The criteria are as follows: Hardware resources: GPU memory ≤ 24GB, available training time < 1 hour; Data scale: the amount of annotated data is ≤ 50,000; Task requirements: The domain overlap between the pre-training task and the downstream task is ≥ 70%; Users can manually select a fine-tuning method when creating a fine-tuning task, or the system can automatically select a fine-tuning method based on rules.
10. The automatic learning engine device based on the encapsulated large model training platform according to claim 8, characterized in that: The evaluation component is specifically used to calculate the following four evaluation indicators: Perplexity, precision, recall, and loss; The evaluation process is specifically implemented as follows: Use the number of iterations specified in the ai.envs.train.yaml file options to train the model; After training is completed, save the model to the path specified by outputs.main in the ai.envs.train.yaml file; During evaluation, read the model under the path and use the split test set for evaluation; Write the evaluation results in JSON file format to the path specified by output.task_result in the ai.envs.train.yaml file.
Citation Information
Patent Citations
AI model training system and method based on fragmented automatic learning
CN110659741A
Container-based machine learning process training task execution method and system
CN112418438A
Model management method and device, electronic equipment and storage medium
CN114003248A
Model training method and device
CN116450156A
Model training deployment system and method
CN119005356A
Cited By
Multi-modal visual analysis system for teaching
CN120339924A
Electronic archive filing method and device based on multi-thread fair scheduling
CN120743478A