Large language model training method and device and electronic equipment

By setting digital labels and sources in the large language model and using reward value optimization for training, the problem of inaccurate digital output is solved and the model's digital generation accuracy is improved.

CN120632448AActive Publication Date: 2025-09-12BEIJING SANKUAI CLOUD COMPUTING TECH CO LTD

Patent Information

Application Number
CN202510705088.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-12
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

Large language models have the problem of hallucination in numerical output. The generated numbers are irrelevant or inaccurate to the original text, and the accuracy of the numbers may be sacrificed during the reinforcement learning alignment process.

Method used

By setting the system prompt words of the large language model, its output carries preset digital labels and digital sources, and using proximal strategy optimization or group relative strategy optimization during the training process, combined with the first reward value and the second reward value for reinforcement learning training, the accuracy of the digital output is optimized.

Benefits of technology

The accuracy of the output numbers of large language models is improved, the occurrence of hallucination problems is reduced, and low-cost digital output improvements are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632448A_ABST
    Figure CN120632448A_ABST
Patent Text Reader

Abstract

The invention provides a large language model training method and device and electronic equipment. The method comprises the steps that control information is set for a large language model through system cue words of the large language model; inputting training data to the large language model, obtaining output data of the large language model, and extracting N digits from the output data; obtaining a first reward value according to the number M of the numbers carrying the preset number label and the number source in the N numbers; in M numbers carrying preset number labels and number sources, standard values corresponding to the numbers are determined according to the number sources, and second reward values are obtained according to comparison results of the standard values and the numbers; and performing reinforcement learning training on the large language model by adopting a near-end strategy optimization mode or a group relative strategy optimization mode, and forming a training reward value in the near-end strategy optimization mode or the group relative strategy optimization mode according to the first reward value and the second reward value. According to the embodiment of the invention, the accuracy of digits generated by the large language model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and more specifically, to a large language model training method, device, and electronic device. Background Art

[0002] With the widespread adoption of large language models, the problem of hallucination in numerical output (generated numbers that are irrelevant or inaccurate to the original text) has attracted widespread attention. Because large language models generate content based on the probability distribution of text, their architecture inherently lacks arithmetic logic units and relies on probabilistic predictions rather than logical calculations, making them prone to deviations when faced with complex numerical relationships. Furthermore, during reinforcement learning alignment, large language models may sacrifice numerical accuracy in favor of text fluency.

[0003] Therefore, there is a need to propose a low-cost solution to the problem of hallucinations in the digital output of large language models.

[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the Invention

[0005] The purpose of the present disclosure is to provide a large language model training method, device and electronic device for solving the problem of hallucination in the numerical output of the large language model at a low cost.

[0006] According to a first aspect of an embodiment of the present disclosure, a large language model training method is provided, comprising: setting control information for the large language model through a system prompt word of the large language model, the control information being used to set a number output by the large language model to carry a preset digital label and a digital source; obtaining training data for the large language model, the training data including digital information; inputting the training data into the large language model, obtaining output data of the large language model, and extracting N numbers from the output data, where N≥0; obtaining a first reward value based on the number M of numbers in the N numbers that carry the preset digital label and digital source, where 0≤M≤N; determining a standard value corresponding to the number among the M numbers that carry the preset digital label and digital source based on the digital source, and obtaining a second reward value based on a comparison result between the standard value and the number; and performing reinforcement learning training on the large language model using a proximal strategy optimization method or a group relative strategy optimization method, wherein a training reward value in the proximal strategy optimization method or the group relative strategy optimization method is formed based on the first reward value and the second reward value.

[0007] According to a second aspect of an embodiment of the present disclosure, a large language model training device is provided, comprising: an output format setting module, configured to set control information for the large language model through a system prompt word of the large language model, wherein the control information is used to set the numbers output by the large language model to carry preset digital labels and digital sources; input the training data into the large language model, obtain the output data of the large language model, and extract N numbers from the output data, where N≥0; a training data acquisition module, configured to obtain the training data of the large language model, wherein the training data includes digital information; an output digital extraction module, configured to input the training data into the large language model, obtain the output data of the large language model, and extract N numbers from the output data; Extract N numbers, N≥0; a format evaluation module, configured to obtain a first reward value based on the number M of numbers carrying the preset digital label and the digital source in the N numbers, 0≤M≤N; an accuracy evaluation module, configured to determine the standard value corresponding to the number according to the digital source among the M numbers carrying the preset digital label and the digital source, and obtain a second reward value based on the comparison result between the standard value and the number; a feedback training module, configured to use a proximal strategy optimization method or a group relative strategy optimization method to perform reinforcement learning training on the large language model, wherein the training reward value in the proximal strategy optimization method or the group relative strategy optimization method is formed according to the first reward value and the second reward value.

[0008] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a memory; and a processor coupled to the memory, wherein the processor is configured to execute any one of the above methods based on instructions stored in the memory.

[0009] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a program is stored. When the program is executed by a processor, the large language model training method as described in any one of the above items is implemented.

[0010] The disclosed embodiment sets the digital output format of the large language model by using the system prompt words of the large language model, so that the large language model carries preset digital labels and digital sources on the output numbers, and at the same time supervises the numbers in the output data of the large language model, and forms a first reward value and a second reward value based on whether they carry the preset digital labels and digital sources and whether they are accurate. When the large language model is subjected to reinforcement learning training using a proximal strategy optimization method or a group relative strategy optimization method, the first reward value and the second reward value are used to form the training reward value in the proximal strategy optimization method or the group relative strategy optimization method, so that the trained large language model can learn the habit of digital output with a basis, while improving the accuracy of the numbers output by the large language model, and reducing the occurrence of hallucination problems in the digital output of the large language model at a low cost.

[0011] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0013] Figure 1 4 is a flowchart of a large language model training method in an exemplary embodiment of the present disclosure.

[0014] Figure 2 is a schematic diagram of the PPO algorithm in an exemplary embodiment of the present disclosure.

[0015] Figure 3 is a schematic diagram of the GRPO algorithm in an exemplary embodiment of the present disclosure.

[0016] Figure 4 2 is a schematic diagram of calculating the training reward value in the exemplary embodiment of the present invention.

[0017] Figure 5 It is a block diagram of a large language model training device in an exemplary embodiment of the present disclosure.

[0018] Figure 6 is a block diagram of an electronic device in an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0019] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the present disclosure will be more comprehensive and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or that other methods, components, devices, steps, etc. may be employed. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present disclosure.

[0020] The accompanying drawings are merely schematic illustrations of the present disclosure. Identical reference numerals in the drawings denote identical or similar components, and thus their repeated descriptions will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0021] Hereinafter, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.

[0022] Figure 1 4 is a flowchart of a large language model training method in an exemplary embodiment of the present disclosure.

[0023] refer to Figure 1 , the large language model training method 100 may include:

[0024] Step S1, setting control information for the large language model using the system prompt word of the large language model, wherein the control information is used to set the numbers output by the large language model to carry a preset digital label and a digital source;

[0025] Step S2, obtaining training data for a large language model, wherein the training data includes digital information;

[0026] Step S3, inputting the training data into the large language model, obtaining output data of the large language model, and extracting N numbers from the output data, where N≥0;

[0027] Step S4, obtaining a first reward value according to the number M of numbers carrying the preset digital label and the digital source in the N numbers, 0≤M≤N;

[0028] Step S5, among the M numbers carrying the preset digital label and the digital source, determine the standard value corresponding to the number according to the digital source, and obtain the second reward value according to the comparison result between the standard value and the number.

[0029] Because large language models perform reasoning based on language and vocabulary, there are issues with hallucinations when generating numbers, which can lead to errors in some numerical indicators. For example, a 5.5% increase in the data table may become 5.4% in the generated report. This is due to the probabilistic decoding method of large language models (such as top-psampling).

[0030] The disclosed embodiment sets the digital output format of the large language model by using the system prompt words of the large language model, so that the large language model carries preset digital labels and digital sources on the output numbers, and at the same time supervises the numbers in the output data of the large language model, and forms a first reward value and a second reward value based on whether they carry the preset digital labels and digital sources and whether they are accurate. When the large language model is subjected to reinforcement learning training using a proximal strategy optimization method or a group relative strategy optimization method, the first reward value and the second reward value are used to form the training reward value in the proximal strategy optimization method or the group relative strategy optimization method, so that the trained large language model can learn the habit of digital output with a basis, while improving the accuracy of the numbers output by the large language model, and reducing the occurrence of hallucination problems in the digital output of the large language model at a low cost.

[0031] Below, each step of the large language model training method 100 is described in detail.

[0032] In step S1, control information is set for the large language model through the system prompt word of the large language model, and the control information is used to set the numbers output by the large language model to carry preset digital labels and digital sources.

[0033] In interactions with a large language model (LLM), unlike the question information (prompt) input by users to the large language model, system prompts are global instructions used to set model behavior rules, background information, or output format. They are usually injected at the beginning of a conversation. By inputting specific control instructions or information into the large language model, they impose structural constraints and regulations on the output content of the large language model, guiding the large language model to respond to subsequent queries in a specific way.

[0034] In an exemplary embodiment, the system prompt word can be passed by calling the system field of the API interface of the large language model (such as OpenAI's Function Calling and Anthropic's system_prompt parameter).

[0035] When the embodiment of the present disclosure starts training the large language model, the output format of the numbers output by the large language model is controlled through system prompt words to achieve training of the large language model's ability to output numbers.

[0036] In an exemplary embodiment, the preset digital tag may be an HTML digital tag (eg <number> …< / number> "tag), so that the large language model can accurately output the required numbers with the preset digital tags through its good HTML format understanding ability. In other embodiments, the preset digital tags can also be other types of tags that can mark the output content as digital, including but not limited to the following forms:

[0037] 1. Use a custom XML tag format, such as <numeric> ...< / numeric> or <data-value> ...< / data-value> , the semantics can be further extended through namespaces and attributes, such as<numerictype="currency"source="financial_db"> 123.45.

[0038] 2. Use Markdown's custom markup syntax, such as wrapping numbers with double tilde lines: ~~25~~, or marking the source in footnote form: 25^[weather_api], and cooperate with the parser to extract tags and sources.

[0039] 3. Add specific prefixes or suffixes before and after numbers to implement natural language markup. For example, use [Number:25], #25#, or Number{25}. This is suitable for scenarios where the natural flow of text needs to be maintained and facilitates regular expression recognition.

[0040] 4. Wrap numbers in a JSON object and implement JSON embedded tags, for example: {"type":"number","value":25,"source":"sensor_data"}. This is particularly suitable for scenarios where API interface return values ​​or structured processing are required.

[0041] 5. In academic or technical documentation, you can use LaTeX math mode markers, such as $25$ or \num{25}, combined with package configuration to implement source annotation:

[0042] \num[source=experimental]{25}.

[0043] 6. Use HTML5 data-* attributes to embed digital information and semantic HTML attributes, for example:<spandata-type="temperature"data-source="weather_station"> 25℃, obtain numbers and metadata through JavaScript parsing.

[0044] 7. Use Unicode private area characters or special symbol combinations as markers to encode special characters, such as encoding numbers as 25 (U+24Dd Surrounded Digit n), or use Emoji symbols to represent numbers and their origins.

[0045] 8. In Word or RTF documents, you can use custom styles or field codes to mark numbers and output rich text format tags, such as {FIELDNUMERIC25\*MERGEFORMAT}, and use the document processing library to extract tag information.

[0046] 9. In specific field communication protocols, prefix identifiers can also be used to mark numbers to implement custom protocol headers, such as NUM@25; SOURCE=GPS, which is suitable for IoT device data transmission or industrial control systems.

[0047] 10. In scenarios where tamper-proofing is required, the digital tag and source information are encapsulated as a blockchain transaction, and a verifiable hash value is generated as the tag, i.e., a blockchain-verifiable tag. For example:

[0048] <numberhash="0xabc123..."> 25. Verify the authenticity of the source through smart contracts.

[0049] These tag types can be selected and combined based on the needs of specific application scenarios. The key is to ensure that the large language model can understand the tagging rules and accurately output content that meets the format requirements, while also facilitating the subsequent system's automated recognition and processing of numbers and their sources. Those skilled in the art can configure these based on actual circumstances.

[0050] In an exemplary embodiment, the source of a number includes the position of the number in the training data, or the calculation process of the number and the position of the number involved in the calculation process in the training data, thereby making the source of the number traceable.

[0051] In an exemplary embodiment, the control information may include textual description information and format examples. The textual description information is used to explain the control purpose, namely, the aforementioned requirement that the large language model output numbers carry a preset numerical label and numerical source. The format examples are used to help the large language model better understand this requirement.

[0052] In an exemplary embodiment, the format example includes an HTML paradigm containing the preset digital tag, and the preset digital tag includes an HTML digital tag. Correspondingly, the format example is an HTML paradigm. For example, the HTML paradigm example requires the large language model to enclose each output number in " <number> …< / number> "tag, where the tag will include attributes such as id, source, equation, etc. depending on the specific situation. The format example is as follows:

[0053] Example of non-traceable numbers: Overall business sales increased by 10%.

[0054] Digital traceability example content: Overall business improved <numberid=”1”source=”table-1”equation=”df[“2024”][1]–df[“2023”][1]”format=”.2f%”> 10.00%.

[0055] This format example is only for the large language model to understand the format (what kind of preset digital labels, how to display the digital source), and does not restrict the large language model to operate according to this example. Therefore, it is best to carry text information explaining this purpose in the control information so that the large language model can better understand the control information. The specific content of the control information can be set according to needs, and this disclosure does not impose too many restrictions on this.

[0056] In step S2, training data of a large language model is obtained, where the training data includes digital information.

[0057] In order to optimize the data output capability of the large language model, the training data acquired in the embodiments of the present disclosure are all training data including digital information, such as various tables, statistical reports, financial statements, scientific experimental data, sensor acquisition sequences, etc.

[0058] In the process of acquiring training data, the acquired training data can be set to cover various forms of structured and unstructured data, such as numerical data in tables, trend figures in statistical reports, income and expenditure details in financial statements, measurement values ​​in scientific experiments, and time series data continuously recorded by sensors, etc., to train the model's ability to process multiple data types.

[0059] Furthermore, when selecting training data, you can select numerical information from different fields and at different scales. Furthermore, to help the model better understand the logical relationships between numbers, you can also configure the training data to include text related to numerical calculations, comparisons, and predictions, such as "Predict next year's revenue growth based on the past three years' sales."

[0060] After acquiring training data, it can be cleaned, labeled, and preprocessed. Numerical data that contains noise, errors, or incomplete information can be specifically processed and labeled. This allows the large language model to learn to identify and correct data issues during training, thereby improving its processing capabilities and output accuracy when faced with complex numerical information in real-world scenarios.

[0061] In step S3, the training data is input into the large language model, output data of the large language model is obtained, and N numbers are extracted from the output data, where N≥0.

[0062] During the training process, a training data is input into the large language model, the output data of the large language model is collected, and the numbers are extracted from the output data.

[0063] For example, based on a certain training data, the output data of the large language model is:

[0064] Net profit for the quarter reached

[0065] <numberid="profit-2024Q2"source="financial-system"equation="revenue-cost"format=",.2f"> 12,580,320.50 yuan,

[0066] Growth compared to the previous quarter

[0067] <numberid="growth-QoQ"source="financial-system"equation="(Q2_2024 / Q1_2024-1)*100"format=".2f%"> 12.35%.

[0068] At this time, it is recognized that there are two numbers, 12,580,320.50 and 12.35%, in the output data, and N=2.

[0069] In step S4, a first reward value is obtained according to the number M of numbers carrying the preset digital label and the digital source in the N numbers, 0≤M≤N.

[0070] After obtaining the number, the preset digital label and the format of the digital source (refer to the above format example) are used to determine whether the number meets the format characteristics to form a first reward value (reward1), which is used to evaluate the large language model's ability to comply with the format requirements. Using this first reward value to train the large language model can improve the large language model's ability to comply with the format requirements.

[0071] For example, for the above-mentioned input and output data, identify the characters before and after the two numbers, and determine through regular expression and other methods that the two numbers meet the preset digital labels and digital sources required in the above-mentioned format example, and are not simple numbers, then M is recorded as 2.

[0072] In an exemplary embodiment, the first reward value may be determined according to M / N. In the above example, M=N=2, and the first reward value is 1, which is the highest score.

[0073] If there is another number in the output data, for example, 10%, which does not have the above-mentioned preset number labels and number sources before and after it, and does not meet the format characteristics, then N=3, M=2, and the first reward value is 2 / 3=0.66. And so on.

[0074] In addition to determining the reward value in the above manner, the reward value may also be determined specifically based on the degree to which the large language model complies with format requirements.

[0075] For example, if a number in the output data completely matches the preset label structure (such as <number> ...< / number> ), the number is awarded 20 points; if the tag exists but lacks required attributes (such as missing source), the number is awarded 10 points; if the tag format is incorrect (such as not closed or the tag name is misspelled), the number is awarded 10 points; if the tag format is missing or incorrect, and there is no source, the number is awarded 0 points. According to the above scoring logic, each number is scored, and the average score of N numbers is finally formed into the first reward value to provide feedback to the large language model on its compliance with format requirements.

[0076] There can be various specific scoring formats for the first reward value, as long as it can evaluate the degree to which the large language model complies with the format requirements.

[0077] In step S5, among the M numbers carrying the preset digital label and the digital source, the standard value corresponding to the number is determined according to the digital source, and the second reward value is obtained according to the comparison result between the standard value and the number.

[0078] Next, among the numbers that meet the format requirements, determine whether the numbers are correct to obtain a second reward value.

[0079] Because large language models reason based on textual information, they are limited by various factors, such as the way text data is decoded and the different precision of numbers. Even when asked to extract a number directly from the training data, errors may occur. The error rate is even higher when asked to generate new numbers based on the numbers provided in the training data.

[0080] Therefore, in order to optimize the output accuracy of the large language model for numbers, the second reward value is used as the training basis of the large language model to provide feedback to the large language model on the accuracy of its digital output.

[0081] Since each of the M numbers that meet the format requirements has a digital source, such as the position in the training data (a row or column in a file), or the numbers generated by the large language model will give a calculation process, and the numbers used in the calculation process will be marked with the digital source, such as annual growth rate - January 2024 data (data source 1) - January 2023 data (data source 2) = 10%, therefore, in the process of forming the second reward value, the value logic corresponding to each number can be obtained first (for example, directly extracted from the training data, or calculated based on the content of the training data). At the same time, according to the digital source corresponding to the number, the value logic is applied to recalculate to obtain the standard value.

[0082] Next, the calculated standard value of the number is compared with the number output by the large language model to determine whether the number is accurate. If the number is consistent with its standard value, T is increased by 1. T is the number of accurate numbers, or the number of numbers that are consistent with its corresponding standard value. The initial value of T is 0.

[0083] After accuracy checking the M numbers, we get a value T, and then we get a second reward value based on T / M. This second reward value can reflect the accuracy of the numbers output by the large language model.

[0084] In step S6, reinforcement learning training is performed on the large language model using a proximal strategy optimization method or a group relative strategy optimization method, wherein a training reward value in the proximal strategy optimization method or the group relative strategy optimization method is formed according to the first reward value and the second reward value.

[0085] Next, the first reward value and the second reward value that reflect the digital output capability of the large language model can be used as immediate reward signals for reinforcement learning. The model parameters can be updated using a policy gradient algorithm (such as PPO, GRPO, etc.) to perform digital output reinforcement training on the large language model.

[0086] Figure 2 is a schematic diagram of the PPO algorithm in an exemplary embodiment of the present disclosure.

[0087] refer to Figure 2 ,PPO (Proximal Policy Optimization) is a reinforcement learning algorithm designed to solve the training instability problem in the policy gradient method.

[0088] The PPO algorithm inputs the input data q into the model to be trained, and the output result O of the model to be trained is passed to the Reference Model, Reward Model and Value Model at the same time.

[0089] The Reference Model calculates the difference between the output and the policy model, measured using KL divergence (a metric that measures the difference between two probability distributions). This difference information is combined with the reward value r output by the Reward Model. The Reward Model awards rewards based on the performance of the policy model's output in the actual task. This training reward value provides quantitative feedback on the quality of the policy model's decisions. The Value Model, on the other hand, outputs a value v, which is used to assess the value of the current state.

[0090] PPO improves training stability and efficiency by limiting the magnitude of policy updates, ensuring that each update does not deviate too far from the current policy. A key difference between PPO and simplified algorithms such as Determined Policy Algorithm (DPO) lies in its reward model. In PPO, the reward model assigns a corresponding categorical reward value to each decision step of the policy model based on environmental feedback and model output, ultimately forming a training reward value. This training reward value serves as a key signal for evaluating the quality of the model's decisions, guiding the policy model to learn and adjust towards optimal results.

[0091] For example, when training a large language model to process digital formats, the reward model of the embodiment of the present disclosure gives a corresponding classification reward value, i.e., a first reward value, based on whether the number conforms to the preset format characteristics. A high reward is given if the conformity is high, and a low reward is given if the conformity is low. A second reward value is given based on whether the number is accurate. The reward model can give different classification reward values ​​for different decision factors, and finally form a fusion of different reward values. Figure 2 The training reward value r is shown.

[0092] Continue to refer Figure 2 In PPO, the advantage value A is calculated using the generalized advantage estimation (GAE) based on the training reward value r and the value evaluation value v. The advantage value A reflects the degree of advantage of taking a certain action compared to the average action in the current state, and guides the trained model (the large language model in this application) on how to update its parameters so that subsequent decisions can obtain higher rewards. In this way, the PPO algorithm limits the range of policy updates and improves the stability and efficiency of training.

[0093] Figure 3 is a schematic diagram of the GRPO algorithm in an exemplary embodiment of the present disclosure.

[0094] refer to Figure 3The core idea of ​​the GRPO (Grouped Proximal Policy Optimization) algorithm is to estimate the baseline through relative rewards within a group, thus avoiding the use of an additional value function model (critic model). Traditional PPO algorithms require training a value function to estimate the advantage function, while GRPO replaces this process by calculating the average reward from multiple outputs of the same problem, significantly reducing memory and computing resource consumption.

[0095] refer to Figure 3 In the GRPO algorithm, the input data q is also input into the PolicyModel. Unlike PPO, the GRPO algorithm will output multiple observation values ​​O1, O2...OG, which are respectively input into the Reference Model and the Reward Model to calculate multiple KL divergence values ​​and multiple training reward values ​​r1, r2...rG. Then, these training reward values ​​are used to obtain multiple advantage values ​​A1, A2...AG through Group Computation. GRPO abandons the Value Model in PPO and estimates the baseline through group scores, which significantly reduces the resources required for training and provides a more efficient training idea in resource-constrained scenarios. The two algorithms provide different solutions for policy optimization in reinforcement learning tasks through different architectures and calculation methods.

[0096] In an exemplary embodiment, a standard PPO / GRPO approach can be used to optimize an LLM (Large Language Model) using RL (Reinforcement Learning). The only modification is the reward calculation formula (training reward value). In this scenario, the reward is defined as a reward signal for "numeric accuracy." Thanks to the traceability of each numeric metric, a verifiable numeric accuracy reward signal can be easily obtained.

[0097] Regardless of the method, in the embodiment of the present disclosure, a training reward value of the large language model can be first formed according to the first reward value and the second reward value, and then the large language model can be optimized according to the training reward value.

[0098] In an exemplary embodiment, the training reward value can be directly formed according to the weighted value of the first reward value and the second reward value, that is, reward = αFormat Reward + βAccuracy Reward. Where reward is the training reward value, that is Figure 2 r in (corresponding to PPO) or Figure 3In the example, r1 to rG (corresponding to GRPO); Format Reward is the first reward value, Accuracy Reward is the second reward value, and α and β are weights. In some embodiments, α<β to improve the training weight for the accuracy of the numbers output by the large language model.

[0099] In some other embodiments, as long as the training method is based on reinforcement learning, the first reward value and the second reward value can both be involved in the calculation of its training reward value.

[0100] In an exemplary embodiment, the training reward value may be formed according to a weighted value of the first reward value, the second reward value, and the third reward value, and the third reward value may include a custom reward value.

[0101] For example, a large language model training method incorporates rewards for optimizing the depth and breadth of the analysis output by the model. To incorporate these rewards, the formula for calculating the training reward is modified to: reward = depth reward + breadth reward + αFormat reward + βAccuracy reward. Both the depth reward and breadth reward are third-order reward values.

[0102] Based on the above example using HTML tags, the process of calculating the training reward value can be summarized as:

[0103] (1) Locate all digital indicators

[0104] After LLM outputs the report (i.e. outputs data), the report content is parsed in HTML to obtain all the <number>The digital index of the package, which includes the attributes of each tag (eg, id, source, equation, etc.). Since LLM has a certain probability of not being generated according to the digital HTML format in the prompt, it is necessary to use rules to determine whether there are any omissions (not included in the <number>The numerical indicators in the tag are numbers.

[0105] (2) Calculate the Format Reward (first reward value) of the digital indicator

[0106] The format reward value (i.e., the first reward value) is set to a range of 0 to 1, and the calculation formula is defined as FormatReward = M / N, where M is the normal reward value. <number>The number of included digital indicators, N is the sum of the number of digital indicators that cannot be generated according to the HTML paradigm and M, that is, the number of all located numbers.

[0107] (3) Accuracy Reward (secondary reward value)

[0108] When calculating the Accuracy Reward, we set aside untraceable numerical indicators and, for M traceable numerical indicators, perform the calculation as follows: For each numerical indicator, based on the source and equation, call Python code and use the number read from the training data according to the data source, or the standard value calculated based on the number read from the training data and the equation, denoted as V; the output data number is denoted as V'. If and only if V = V', the numerical indicator is marked as accurately generated, and T is increased by 1; otherwise, it is marked as inaccurately generated. For the overall report (output data) containing N numerical indicators, its Accuracy Reward (second reward value) is defined as T / M, where T is the number of accurate numerical indicators.

[0109] (4) The training reward value reward = FormatReward + AccuracyReward is obtained in combination.

[0110] Finally, during the RL process, a comprehensive reward score is calculated for each output data output by the trained large language model, so that the model maximizes this reward score through the RL algorithm, ultimately significantly improving the accuracy of the numbers in the output data generated by the large language model.

[0111] Figure 4 2 is a schematic diagram of calculating the training reward value in the exemplary embodiment of the present invention.

[0112] refer to Figure 4 At event 01, training data IN is input into the large language model to obtain output data OUT.

[0113] At event 02, N digits of the output data OUT are extracted.

[0114] At event 03, the number M of digits that meet the format requirements in the N digits is determined.

[0115] At event 04, among the M numbers that meet the format requirements, their accuracy is verified according to the corresponding data source, and the number T of accurate numbers is obtained.

[0116] At event 05 , a first reward value Format Reward is calculated based on M / N, and a second reward value Accuracy Reward is calculated based on T / M.

[0117] At event 06 , a training reward value is obtained based on the first reward value and the second reward value.

[0118] At event 07, the large language model is trained using RL based on the training reward value.

[0119] The disclosed embodiment sets the numbers output by the large language model to carry preset digital labels and digital sources, so that each number output by the large language model can be traced (knowing where the number is referenced from or what formula is used to calculate it); secondly, by designing a verifiable "digital accuracy reward signal" after setting each digital indicator to be traceable, the LLM can optimize the digital accuracy through the RL algorithm, achieving 100% traceability of each number and a significant improvement in the accuracy of the generated numbers, thereby improving the digital generation capability of the large language model at a low cost.

[0120] Corresponding to the above method embodiments, the present disclosure also provides a large language model training device, which can be used to execute the above method embodiments.

[0121] Figure 5 It is a block diagram of a large language model training device in an exemplary embodiment of the present disclosure.

[0122] refer to Figure 5 , the large language model training device 500 may include:

[0123] An output format setting module 51 is configured to set control information for the large language model using a system prompt word of the large language model, wherein the control information is used to set a preset digital label and digital source for a number output by the large language model;

[0124] A training data acquisition module 52 is configured to acquire training data for a large language model, wherein the training data includes digital information;

[0125] An output digit extraction module 53 is configured to input the training data into the large language model, obtain output data of the large language model, and extract N digits from the output data, where N≥0;

[0126] A format evaluation module 54 is configured to obtain a first reward value according to the number M of digits carrying the preset digital label and the digital source in the N digits, 0≤M≤N;

[0127] an accuracy evaluation module 55 configured to determine, among the M numbers carrying the preset digital labels and digital sources, a standard value corresponding to the number according to the digital source, and obtain a second reward value based on a comparison result between the standard value and the number;

[0128] The feedback training module 56 is configured to perform reinforcement learning training on the large language model using a proximal strategy optimization method or a group relative strategy optimization method, wherein a training reward value in the proximal strategy optimization method or the group relative strategy optimization method is formed based on the first reward value and the second reward value.

[0129] Since the functions of the apparatus 500 have been described in detail in the corresponding method embodiments, they will not be described in detail herein.

[0130] It should be noted that although several modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.

[0131] In an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is also provided.

[0132] Those skilled in the art will appreciate that various aspects of the present invention may be implemented as systems, methods, or program products. Accordingly, various aspects of the present invention may be implemented as a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be collectively referred to herein as "circuits," "modules," or "systems."

[0133] Refer to the following Figure 6 An electronic device 600 according to this embodiment of the present invention will be described. Figure 6 The electronic device 600 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present invention.

[0134] like Figure 6 As shown, electronic device 600 is implemented as a general-purpose computing device. Components of electronic device 600 may include, but are not limited to, the aforementioned at least one processing unit 610, the aforementioned at least one storage unit 620, and a bus 630 connecting different system components (including storage unit 620 and processing unit 610).

[0135] The storage unit stores program code that can be executed by the processing unit 610, causing the processing unit 610 to perform the steps described in the "Exemplary Methods" section above according to various exemplary embodiments of the present invention. For example, the processing unit 610 can perform the methods described in the embodiments of the present disclosure.

[0136] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 6201 and / or a cache 6202 , and may further include a read-only memory unit (ROM) 6203 .

[0137] The storage unit 620 may also include a program / utility 6204 having a set (at least one) of program modules 6205, such program modules 6205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0138] Bus 630 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0139] The electronic device 600 can also communicate with one or more external devices 700 (e.g., a keyboard, a pointing device, a Bluetooth device, etc.), one or more devices that enable a user to interact with the electronic device 600, and / or any device that enables the electronic device 600 to communicate with one or more other computing devices (e.g., a router, a modem, etc.). Such communication can occur via an input / output (I / O) interface 650. Furthermore, the electronic device 600 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 660. As shown, the network adapter 660 communicates with other modules of the electronic device 600 via a bus 630. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 600, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0140] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0141] In exemplary embodiments of the present disclosure, a computer-readable storage medium is also provided, storing a program product capable of implementing the methods described above. In some possible implementations, various aspects of the present invention may also be implemented in the form of a program product comprising program code. When the program product is executed on a terminal device, the program code is configured to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the "Exemplary Methods" section above.

[0142] The program product for implementing the above-described method according to an embodiment of the present invention may be a portable compact disc read-only memory (CD-ROM) and include program code, and may be run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0143] The program product may be implemented in any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0144] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0145] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0146] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and the like, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a standalone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0147] Furthermore, the above-described figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes illustrated in the above-described figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0148] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the claims.< / number> < / number> < / number>

Claims

1. A large language model training method, characterized in that: include: Setting control information for the large language model using the system prompt word of the large language model, wherein the control information is used to set a number output by the large language model to carry a preset digital label and a digital source; Obtaining training data for a large language model, wherein the training data includes digital information; Inputting the training data into the large language model, obtaining output data of the large language model, and extracting N numbers from the output data, where N≥0; Obtaining a first reward value according to the number M of numbers carrying the preset digital label and the digital source in the N numbers, 0≤M≤N; Among the M numbers carrying the preset digital labels and digital sources, determining a standard value corresponding to the number according to the digital source, and obtaining a second reward value according to a comparison result between the standard value and the number; Reinforcement learning training is performed on the large language model using a proximal strategy optimization method or a group relative strategy optimization method, wherein a training reward value in the proximal strategy optimization method or the group relative strategy optimization method is formed according to the first reward value and the second reward value.

2. The large language model training method according to claim 1, wherein: The control information includes a format example, the format example includes an HTML paradigm including the preset digital tag, and the preset digital tag includes an HTML digital tag.

3. The large language model training method according to claim 1 or 2, characterized in that: The source of the number includes the position of the number in the training data, or the calculation process of the number and the position of the number involved in the calculation process in the training data.

4. The large language model training method according to claim 1, wherein: Forming a training reward value in the proximal strategy optimization method or the group relative strategy optimization method according to the first reward value and the second reward value includes: The training reward value is formed according to a weighted value of the first reward value and the second reward value.

5. The large language model training method according to claim 1, wherein: Forming a training reward value in the proximal strategy optimization method or the group relative strategy optimization method according to the first reward value and the second reward value includes: The training reward value is formed according to a weighted value of the first reward value, the second reward value, and a third reward value, and the third reward value includes a custom reward value.

6. The large language model training method according to claim 1, wherein: The first reward value is obtained according to M / N, and the second reward value is obtained according to T / M, where T is the number of digits that are consistent with the corresponding standard value.

7. A large language model training device, characterized in that: include: an output format setting module, configured to set control information for the large language model using a system prompt word of the large language model, wherein the control information is used to set a number output by the large language model to carry a preset digital label and a digital source; a training data acquisition module configured to acquire training data for a large language model, wherein the training data includes digital information; an output digit extraction module, configured to input the training data into the large language model, obtain output data of the large language model, and extract N digits from the output data, where N≥0; a format evaluation module configured to obtain a first reward value according to a number M of digits carrying the preset digital label and the digital source in the N digits, 0≤M≤N; an accuracy assessment module configured to determine, among the M numbers carrying the preset digital labels and digital sources, a standard value corresponding to the number according to the digital source, and obtain a second reward value based on a comparison result between the standard value and the number; The feedback training module is configured to perform reinforcement learning training on the large language model using a proximal strategy optimization method or a group relative strategy optimization method, wherein a training reward value in the proximal strategy optimization method or the group relative strategy optimization method is formed based on the first reward value and the second reward value.

8. An electronic device, characterized in that: include: Memory; as well as A processor coupled to the memory, wherein the processor is configured to execute the method according to any one of claims 1 to 6 based on instructions stored in the memory.

9. A computer-readable storage medium having a program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Method and device for reinforcement learning of large language model

    CN117808120A

  • Big language model illusion relieving scheme based on citation correction

    CN117910449A

  • Method, device and equipment for training large language model

    CN118153624A

  • Method and system for enhancing large language model generation by using network search

    CN119271893A

  • Large language model training method and device

    CN119443155A

Cited By

  • Multi-modal large model training method and related device

    CN121257750A

  • Classification model training method, information classification method and device and electronic equipment

    CN121388783A

  • Training method, platform and equipment for diffusion language model, medium and product

    CN121859920A