Training method and device of vertical field large model

By constructing the initial training dataset and performing pre-training and supervising fine-tuning, combining actual unit testing and reward model training, the problem that the general big model cannot fit vertical field projects is solved, improving the output quality and coverage of the model, and meeting the business needs of specific fields.

CN120338030APending Publication Date: 2025-07-18广州宸祺出行科技有限公司
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510438724.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing general-purpose big models cannot meet the needs of vertical fields projects, the output format does not match the business requirements, and the unit test coverage is low.

Method used

The initial training data set is constructed by collecting project code, a database that stores travel APP data, and manually written data, and pre-training and supervising fine-tuning of the initial large model, combining actual unit testing and reward model training, optimize the model to meet the data format and business needs of the vertical field.

Benefits of technology

It improves the output quality and unit test coverage of large models in vertical fields, reduces the time consumption of manual completion unit tests, and improves the generation accuracy and efficiency of the model in specific fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338030A_ABST
    Figure CN120338030A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a vertical field large model training method, which comprises the steps of S10, collecting an initial training data set; s20, pre-training an initial large model based on the project code until the initial large model learns a style and a data format of the project code; s30, performing supervised fine tuning on the initial large model through an initial training data set to obtain a supervised fine tuning stage large model meeting the requirements of the vertical field; s40, constructing a reward model training data set by screening part of data used for supervising and fine tuning and combining output data of the large model in the supervising and fine tuning stage; s50, according to the reward model training data set, reward model training is carried out on the large model in the supervision fine tuning stage, and a reward model which is subjected to sorting scoring training is obtained; and S60, through the reward model and all the collected data, training the large model in the supervision fine tuning stage to obtain a special large model in the vertical field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a training method and device for a large model in a vertical domain. Background Art

[0002] Currently, the explosion of AIGC technology is no less significant than the profound impact and far-reaching changes brought about by computers to society, and it has also promoted major developments in all industries. AIGC (AI-Generated Content) is based on technologies such as deep learning and natural language processing. By training models to imitate human creation and generate content that meets specific requirements, it can generate various forms of content using artificial intelligence technology, such as articles, music, paintings, videos, etc.

[0003] The quality of software directly affects the reliability, usability, and user satisfaction of software products. Among them, the code coverage rate of unit tests is an important indicator to measure the quality of software testing. However, although the general-purpose large models on the market have certain reasoning capabilities, due to the lack of data in vertical domains such as internal design documents and product codes, the output result data is not ideal when directly used. General large models are trained on large-scale and diverse general data, with the goal of having broad knowledge and general capabilities to handle various common scenarios and tasks. However, this generality also makes them have obvious limitations when facing vertical domain projects. Vertical domain projects often have unique business logics, professional terms, and data format requirements.

[0004] In addition, for the test evaluation of mainstream general large models, the test coverage rate of the data generated by different general models can only reach about 30%, which is still a significant gap from manual testing. In order to make the model capabilities meet the production requirements, we need to fine-tune and train the basic model using its own data to improve the intelligence level of the large model in the vertical domain. Summary of the Invention

[0005] The technical problem to be solved by the present invention is: to solve the problem that the existing general large models cannot meet the needs of vertical domain projects, and the output format of the general large models does not match the business requirements.

[0006] To solve the above technical problem, the present invention provides a training method for a large model in a vertical domain, and the method includes:

[0007] S10, collecting an initial training data set, which is obtained by collecting project codes, a database storing travel APP data, and manually written data.

[0008] S20. Pre-train the initial large model based on the project code until the initial large model learns the style and data format of the project code;

[0009] S30. Perform supervised fine-tuning on the initial large model using the initial training dataset to obtain a large model in the supervised fine-tuning stage that meets the requirements of the vertical domain. The large model in the supervised fine-tuning stage can basically output data that conforms to the format;

[0010] S40. By screening part of the data used for supervised fine-tuning and combining the data output by the large model in the supervised fine-tuning stage itself, conduct actual unit tests, calculate the coverage rate, and sort and score according to the coverage rate to construct a reward model training dataset;

[0011] S50. Train the large model in the supervised fine-tuning stage according to the reward model training dataset to obtain a reward model trained by sorting and scoring;

[0012] S60. Train the large model in the supervised fine-tuning stage through the reward model and all the collected data to obtain a dedicated large model for the vertical domain.

[0013] Furthermore, in S10, the initial training dataset further includes collecting ES logs, and the sorting method further includes data cleaning and formatting.

[0014] Furthermore, in S10, the method includes:

[0015] Write a script to collect datasets from project codes, databases, and ES logs;

[0016] Screen and optimize the dataset and perform formatting processing;

[0017] Save the processed dataset in JSON format to a file to form an initial training dataset.

[0018] Furthermore, in S40, the method includes:

[0019] Connect to an open-source large model, which is used to generate unit test data and calculate the coverage rate.

[0020] Furthermore, the method includes:

[0021] Connect to a training framework and write a corresponding supervised fine-tuning training script;

[0022] Load the initial training dataset to the GPU and perform supervised fine-tuning training based on the training script;

[0023] During the training process, monitor the training loss value and filter the checkpoint according to the loss value;

[0024] Load the model parameters corresponding to the checkpoint for inference ability evaluation;

[0025] After the supervised fine-tuning training is completed, load the parameters of the dedicated large model in the vertical domain to perform secondary inference evaluation and verify the accuracy of the generated content of the model.

[0026] Furthermore, the method includes

[0027] S70, optimize and apply the dedicated large model in the vertical domain to the vertical domain.

[0028] Furthermore, optimize and apply the dedicated large model in the vertical domain to the vertical domain; including the following steps:

[0029] S71, load the dedicated large model in the vertical domain and connect it to the project to generate unit test data;

[0030] S72, obtain the code call chain and calculate the coverage rate;

[0031] S73, summarize the unit test data, sort out the path data, and write and generate content in the vertical domain.

[0032] According to another aspect of the present invention, there is provided a training device for a large model in a vertical domain. When using the training device for the large model in the vertical domain to perform model training, it includes the above-mentioned training method for the large model in the vertical domain.

[0033] Furthermore, the device includes:

[0034] A data collection module, which is used to collect an initial training set, and the initial training set is compiled by collecting project codes, a database storing travel APP data, and collecting manually written data;

[0035] A pre-training module, which is used to pre-train the initial large model based on the project code until the initial large model learns the style and data format of the project code;

[0036] A supervised fine-tuning module, which performs supervised fine-tuning on the initial large model through the initial training data set to obtain a large model in the supervised fine-tuning stage that meets the requirements of the vertical domain, and the large model in the supervised fine-tuning stage can basically output data that meets the format;

[0037] Reward model training dataset construction module, which is used to construct a reward model training dataset by screening part of the data used for supervised fine-tuning, combining the output data of the large model itself in the supervised fine-tuning stage, performing actual unit tests, calculating the coverage rate, and sorting and scoring according to the coverage rate;

[0038] Reward model construction module, which is used to train the large model in the supervised fine-tuning stage according to the reward model training dataset to obtain a reward model trained by sorting and scoring;

[0039] Reinforcement learning training module, which is used to train the large model in the supervised fine-tuning stage through the reward model and all the collected data to obtain a dedicated large model for the vertical domain.

[0040] Compared with the prior art, the beneficial effects of a training method for a large model in a vertical domain provided by the present invention are as follows:

[0041] The present invention constructs an initial training dataset by collecting project codes, databases storing travel APP data, and manually written data, ensuring that the large model in the vertical domain can learn the specific styles and data formats in the vertical domain. After pre-training and supervised fine-tuning, the model can basically output data that conforms to the business format in the vertical domain, thus meeting the requirements of specific projects. The present invention can guide the large model in the vertical domain to pay more attention to the quality of the output data and the matching degree with business requirements during the training process through the reward model, thereby improving the output quality of the large model in the vertical domain. The large model in the vertical domain of the present invention also improves the coverage rate of unit tests during the training process and reduces the time consumed by manual completion of unit tests. Description of the Drawings

[0042] Figure 1 is a flowchart of the training method for the large model in the vertical domain provided by the embodiment of the present invention;

[0043] Figure 2 is another flowchart of the training method for the large model in the vertical domain provided by the embodiment of the present invention;

[0044] Figure 3 is a flowchart of the training method for the large model in the vertical domain provided by another embodiment of the present invention;

[0045] Figure 4 is a schematic diagram of the training device for the large model in the vertical domain provided by the embodiment of the present invention;

[0046] In the figure, 10 is the data collection module; 20 is the pre-training module; 30 is the supervised fine-tuning module; 40 is the reward model training dataset construction module; 50 is the reward model construction module; 60 is the reinforcement learning training module. Specific Embodiments

[0047] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present invention. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted below.

[0048] As Figure 1-2 shown, in an alternative embodiment of the present invention, the training method of the vertical domain large model includes:

[0049] S10, collecting an initial training dataset, which is compiled by collecting project codes, a database storing travel APP data, and collecting manually written data;

[0050] S20, pre-training the initial large model based on the project codes until the initial large model learns the style and data format of the project codes;

[0051] S30, performing supervised fine-tuning on the initial large model through the initial training dataset to obtain a large model in the supervised fine-tuning stage that meets the requirements of the vertical domain, and the large model in the supervised fine-tuning stage can basically output data that conforms to the format;

[0052] S40, by screening part of the data used for supervised fine-tuning, combining the data output by the large model in the supervised fine-tuning stage itself, performing actual unit tests, calculating the coverage rate, and sorting and scoring according to the coverage rate to construct a reward model training dataset;

[0053] S50, training the large model in the supervised fine-tuning stage according to the reward model training dataset to obtain a reward model trained by sorting and scoring;

[0054] S60, training the large model in the supervised fine-tuning stage through the reward model and all the collected data to obtain a dedicated large model for the vertical domain.

[0055] Among them, the initial training dataset is a dataset compiled from project codes, databases storing travel APP data, and manually written data, and is used for the initial training and fine-tuning of the model. The initial large model refers to a general large model that has not been trained with specific domain data and has basic language understanding and generation capabilities. The large model in the supervised fine-tuning stage refers to the model after being supervised and fine-tuned with the initial training dataset, and can basically output data that conforms to the vertical domain format. The reward model training dataset refers to the dataset constructed by screening part of the data used in supervised fine-tuning, combining the data output by the model itself for actual unit testing, calculating the coverage rate, and sorting and scoring according to the coverage rate, and is used for the training of the reward model. The reward model refers to the model trained by sorting and scoring, and is used to guide the further training of the large model in the supervised fine-tuning stage and optimize its output quality and business matching degree. The dedicated large model in the vertical domain refers to the dedicated large model obtained after being trained with the reward model and all the collected data, and can highly meet the project requirements of the vertical domain.

[0056] Specifically, in S10, its role is to provide a data basis for the initial training and fine-tuning of the model, ensuring that the model can learn the specific styles and data formats in the vertical domain. By collecting project codes, databases storing travel APP data, and manually written data, and organizing them, an initial training dataset is formed. In S20, its role is to enable the initial large model to learn the styles and data formats of project codes, and the model can initially understand the language characteristics and data representation methods in the vertical domain. By using project codes to pre-train the initial large model until the model learns the styles and data formats of project codes. In S30, its role is to make the model more in line with the requirements of the vertical domain and obtain a large model in the supervised fine-tuning stage that can basically output data in a compliant format. Use the initial training dataset to perform supervised fine-tuning on the initial large model, adjust the parameters of the model to make its output more in line with the business format of the vertical domain. In S40, its role is to provide data support for the training of the reward model, ensuring that the reward model can accurately evaluate the output quality of the large model in the supervised fine-tuning stage. Screen a part of the data used in the supervised fine-tuning, combine the output data of the large model itself in the supervised fine-tuning stage, conduct actual unit tests, calculate the coverage rate, and sort and score according to the coverage rate to construct a reward model training dataset. In S50, its role is to obtain a reward model that can guide the further training of the large model in the supervised fine-tuning stage. The reward model can optimize the output quality and business matching degree of the large model in the supervised fine-tuning stage. Use the reward model training dataset to perform reward model training on the large model in the supervised fine-tuning stage to obtain a reward model trained through sorting and scoring. In S60, its role is to further train the large model in the supervised fine-tuning stage through the reward model and all the collected data. Obtain a dedicated large model that can highly fit the requirements of vertical domain projects. Use the reward model and all the collected data to train the large model in the supervised fine-tuning stage, adjust the parameters of the model to make its output more in line with the business requirements and format requirements of the vertical domain.

[0057] Illustrated with an example Figure 2 The embodiments of the present invention can be further divided into five stages:

[0058] Training data collection:

[0059] Extract a large amount of data from several places such as project codes, databases storing travel APP data, and ES logs. At the same time, combine the data collected and written manually, perform cleaning and formatting, and organize them into an initial training dataset.

[0060] Pre-training (PT):

[0061] Perform pre-training through project codes to enable the large model to basically learn the styles, data formats, etc. of project codes.

[0062] Supervised fine-tuning (SFT):

[0063] Using the initial training dataset, the large model is supervised and fine-tuned to obtain a large model that meets the requirements of the vertical domain. The large model at this stage can basically output data that conforms to the format, but there is a certain gap in the accuracy of the output data. Therefore, subsequent reinforcement training is required on this basis.

[0064] Reward Model Training (RW):

[0065] By screening part of the data used in the SFT stage and combining it with the data output by the AI large model itself, actual unit tests are conducted, the coverage rate is calculated, and sorting and scoring are performed according to the coverage rate to construct a reward model training dataset. And this training set is used to perform RW training on the large model obtained in the SFT stage to obtain a reward model trained through sorting and scoring, which is used as the scoring model for the next stage.

[0066] Proximal Policy Optimization (PPO) Training:

[0067] Combining the reward model obtained in the RW stage and all the collected data, proximal policy optimization training is performed on the large model obtained in the SFT stage. After proximal policy optimization training, the large model can better judge the accuracy of the data, and the accuracy of the generated data is greatly improved. The large model that finally meets the actual needs of the vertical domain can be obtained at this stage.

[0068] In the embodiment of the present invention, an initial training dataset is constructed by collecting project codes, databases storing travel APP data, and manually written data, ensuring that the large model in the vertical domain can learn the specific style and data format of the vertical domain. After pre-training and supervised fine-tuning, the model can basically output data that conforms to the business format of the vertical domain, thereby meeting the requirements of specific projects. The present invention can guide the large model in the vertical domain to pay more attention to the quality of the output data and the matching degree with business requirements during the training process through the reward model, thereby improving the output quality of the large model in the vertical domain. The large model in the vertical domain of the present invention also improves the coverage rate of unit tests during the training process and reduces the time consumed by manual completion of unit tests.

[0069] The embodiment of the present invention can combine with actual projects to write a scoring script for the collected data. By collecting the call chain of the request code, calculating the code coverage rate of the actual request of the data, obtaining the score of the data through feedback, and sorting through the score, the dataset for proximal policy optimization training is finally obtained. Based on the above fine-tuning method, a dedicated large model for the vertical domain is trained, which can further improve the accuracy of data generation on the basis of the original fine-tuning, be applied to actual projects more efficiently and accurately, further improve the development efficiency and productivity, and further improve the ability of intelligent office.

[0070] In an alternative embodiment of the present invention, in S10, the initial training dataset further includes collecting ES logs, and the sorting method further includes data cleaning and formatting.

[0071] Specifically, in step S10, in addition to collecting project codes, the database storing travel APP data, and manually written data, the initial training dataset also includes the collection of ES logs. ES logs, as key records during the system operation, contain rich information, such as user behavior, system exceptions, performance data, etc. These log data are of great significance for model training because they can reflect the real situation of the system during actual operation and help the model learn data features and patterns closer to the actual business scenarios.

[0072] Specifically, to ensure the quality and consistency of the initial training dataset, the sorting method further includes data cleaning and formatting. For example, ES logs may contain a large amount of useless or incorrect information, such as duplicate log entries, logs with incorrect formats, etc. By data cleaning, these noisy data can be removed to improve the quality of the dataset.

[0073] In the embodiment of the present invention, by introducing ES log data, the model can learn more diverse data features and business scenarios, thereby enhancing its generalization ability in actual applications. Data cleaning and formatting can remove noisy data, handle missing values, and standardize data, thereby improving the quality and consistency of the dataset. This helps the model more accurately learn the patterns and rules in the data and improve the accuracy of the model output.

[0074] In an alternative embodiment of the present invention, in S10, the method includes:

[0075] Writing a script to collect datasets from project codes, databases, and ES logs;

[0076] Screening and optimizing the dataset, and performing formatting processing;

[0077] Saving the processed dataset in JSON format to a file to form an initial training dataset.

[0078] Among them, JSON (JavaScript Object Notation) is a lightweight data exchange format, which is easy for humans to read and write, and is also easy for machines to parse and generate. Selecting the JSON format can facilitate subsequent data processing and model training.

[0079] In the embodiment of the present invention, the data exchange efficiency of the JSON format is high, and the dataset can be quickly transmitted and shared between different systems, which helps to accelerate the model training and development process.

[0080] In an alternative embodiment of the present invention, the method in S40 includes:

[0081] Access an open-source large model, which is used to generate unit test data and calculate code coverage.

[0082] Among them, select a suitable open-source large model, such as GPT series, Codex, etc. These models perform well in natural language processing and code generation and can generate high-quality unit test data. Connect the open-source large model to the training process through API calls or local deployment to ensure that the model can receive input and generate corresponding output. Prepare the code snippets or functions for which unit test data needs to be generated as input. Call the generation function of the open-source large model to generate corresponding unit test data according to the input code snippets or functions. These data may include test cases, expected outputs, etc. Select a suitable code coverage calculation tool, such as JaCoCo, Cobertura, etc. These tools can analyze the execution paths and coverage of the code. Use the selected coverage calculation tool to calculate the code coverage of the generated unit test data. By calculating the coverage, the coverage degree of the test cases for the code paths can be evaluated, thereby guiding the generation and optimization of subsequent test cases.

[0083] Such as Figure 3 shown, in an alternative embodiment of the present invention, the method includes:

[0084] Access the training framework and write the corresponding supervised fine-tuning training script;

[0085] Load the initial training dataset onto the GPU and perform supervised fine-tuning training based on the training script;

[0086] Monitor the loss value during training and select checkpoints according to the loss value;

[0087] Load the model parameters corresponding to the checkpoint for inference ability evaluation;

[0088] After the supervised fine-tuning training is completed, load the parameters of the dedicated large model in the vertical domain to perform secondary inference evaluation and verify the accuracy of the generated content of the model.

[0089] In an alternative embodiment of the present invention, the method includes:

[0090] S70, optimize and apply the dedicated large model in the vertical domain.

[0091] In an alternative embodiment of the present invention, optimize and apply the dedicated large model in the vertical domain; the method includes the following steps:

[0092] S71. Load the dedicated large model for the vertical domain and connect it to the project to generate unit test data.

[0093] S72. Obtain the code call chain and calculate the coverage rate.

[0094] S73. Summarize the unit test data, sort out the path data, and write and generate content for the vertical domain.

[0095] Specifically, obtaining the code call chain is to obtain the call chain information of the project code through static analysis or dynamic tracking technology, which helps to understand the execution flow and dependency relationship of the code. Coverage rate calculation refers to selecting a suitable code coverage rate calculation tool (such as JaCoCo, Cobertura, etc.) to calculate the coverage rate of the generated unit test data. These tools can analyze the execution path and coverage of the code and provide a detailed coverage report.

[0096] In the embodiment of the present invention, by optimizing and applying the dedicated large model for the vertical domain, the performance and accuracy of the model in a specific domain can be further improved; the embodiment of the present invention improves the coverage rate of unit tests and reduces the time consumed by manual completion of unit tests.

[0097] According to another aspect of the present invention, there is provided a training device for a large model in a vertical domain. When using the training device for the large model in the vertical domain for model training, it includes the above-mentioned training method for the large model in the vertical domain.

[0098] As Figure 4 shown, in an optional embodiment of the present invention, the device includes:

[0099] A data collection module 10, which is used to collect an initial training set, and the initial training set is compiled by collecting project codes, a database storing travel APP data, and manually written data.

[0100] A pre-training module 20, which is used to pre-train an initial large model based on the project code until the initial large model learns the style and data format of the project code.

[0101] A supervised fine-tuning module 30, which performs supervised fine-tuning on the initial large model through an initial training data set to obtain a large model in the supervised fine-tuning stage that meets the requirements of the vertical domain, and the large model in the supervised fine-tuning stage can basically output data that conforms to the format.

[0102] The reward model training dataset construction module 40 is used to construct a reward model training dataset by screening part of the data used for supervised fine-tuning, combining the output data of the large model itself in the supervised fine-tuning stage, performing actual unit tests, calculating the coverage rate, and sorting and scoring according to the coverage rate;

[0103] The reward model construction module 50 is used to train a reward model for the large model in the supervised fine-tuning stage according to the reward model training dataset, and obtain a reward model trained by sorting and scoring;

[0104] The reinforcement learning training module 60 is used to train the large model in the supervised fine-tuning stage through the reward model and all the collected data to obtain a dedicated large model for the vertical domain.

[0105] It should be understood that various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present invention can be achieved, and no limitations are imposed herein.

[0106] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A training method for a large model in a vertical domain, characterized in that, The method includes: S10. Collect an initial training data set, which is compiled by collecting project codes, a database storing travel APP data, and manually written data; S20. Pre-train an initial large model based on the project codes until the initial large model learns the style and data format of the project codes; S30. Perform supervised fine-tuning on the initial large model through the initial training data set to obtain a large model in the supervised fine-tuning stage that meets the requirements of the vertical domain. The large model in the supervised fine-tuning stage can basically output data in line with the format; S40. By screening some of the data used in the supervised fine-tuning, combining the output data of the large model in the supervised fine-tuning stage itself, perform actual unit tests, calculate the coverage rate, and sort and score according to the coverage rate to construct a reward model training data set; S50. Train the large model in the supervised fine-tuning stage according to the reward model training data set to obtain a reward model trained by sorting and scoring; S60. Train the large model in the supervised fine-tuning stage through the reward model and all the collected data to obtain a dedicated large model for the vertical domain.

2. The training method of the large model in the vertical field according to claim 1, characterized in that, In S10, the initial training data set also includes collecting ES logs, and the sorting method also includes data cleaning and formatting.

3. The training method of the large model in the vertical domain according to claim 2, characterized in that, In S10, the method includes: Write a script to collect data sets from project codes, databases, and ES logs; Screen and optimize the data sets and perform formatting processing; Save the processed data sets in JSON format to a file to form an initial training data set.

4. The training method of the large model in the vertical domain according to claim 1, characterized in that In S40, the method includes: Connect to an open-source large model, which is used to generate unit test data and calculate the coverage rate.

5. The training method of the large model in the vertical field according to claim 1, characterized in that The method includes: Connect to a training framework and write a corresponding supervised fine-tuning training script; Load the initial training data set to the GPU and perform supervised fine-tuning training based on the training script; Monitor the loss value of the training during the training process and screen the checkpoint according to the loss value; Load the model parameters corresponding to the checkpoint for inference ability evaluation; After the supervised fine-tuning training is completed, load the parameters of the dedicated large model for the vertical domain to perform secondary inference evaluation and verify the accuracy of the generated content of the model.

6. The training method of the large model in the vertical domain according to claim 1, characterized in that, The method includes S70. Optimize and apply the dedicated large model for the vertical domain.

7. The training method of the large model in the vertical domain according to claim 6, characterized in that In S70, optimize and apply the dedicated large model for the vertical domain; it includes the following steps: S71. Load the dedicated large model for the vertical domain and connect it to the project to generate unit test data; S72. Obtain the code call chain and calculate the coverage rate; S73. Summarize the unit test data, sort out the path data, and write to generate content for the vertical domain.

8. A training device for a large model in a vertical domain, characterized in that, When using the training device for the vertical domain large model to perform model training, it includes the training method of the vertical domain large model according to any one of claims 1-7.

9. The training device for the large model in the vertical domain according to claim 8, characterized in that, The device includes: A data collection module, which is used to collect an initial training set. The initial training set is compiled by collecting project codes, a database storing travel APP data, and manually written data. A pre-training module, which is used to pre-train an initial large model based on the project codes until the initial large model learns the style and data format of the project codes. A supervised fine-tuning module, which performs supervised fine-tuning on the initial large model through the initial training data set to obtain a large model in the supervised fine-tuning stage that meets the requirements of the vertical domain. The large model in the supervised fine-tuning stage can basically output data that conforms to the format. A reward model training data set construction module, which is used to construct a reward model training data set by screening part of the data used in the supervised fine-tuning, combining the data output by the large model in the supervised fine-tuning stage itself, performing actual unit tests, calculating the coverage rate, and sorting and scoring according to the coverage rate. A reward model construction module, which is used to train the large model in the supervised fine-tuning stage according to the reward model training data set to obtain a reward model trained through sorting and scoring. A reinforcement learning training module, which is used to train the large model in the supervised fine-tuning stage through the reward model and all the collected data to obtain a dedicated large model for the vertical domain.

Citation Information

Cited By

  • Large code model training method and electronic equipment

    CN121187569A

  • Method for training a code large model and electronic device

    CN121187569B

  • Offshore wind power and ocean engineering large model application method, device, equipment and medium

    CN122114190A