Model training method, target reasoning task execution method, device and apparatus
Patent Information
- Application Number
- CN202610868223.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-15
- Publication Date
- 2026-09-08
Smart Images

Figure CN122713428A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, particularly to the fields of large language models, intelligent agents, and intelligent agents, specifically to a model training method, a method for executing a target reasoning task, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology
[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies mainly include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.
[0003] In recent years, large reasoning models (LRMs) such as OpenAI-ol and DeepSeek-R1 have achieved significant performance improvements in complex reasoning tasks through extended chain-of-thought (CoT) sequences. Large reasoning models have also been widely used in application scenarios such as intelligent reasoning assistants, code generation and debugging, complex logic planning, and deep analysis of scientific literature.
[0004] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention
[0005] This disclosure provides a model training method, a method for performing a target reasoning task, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product.
[0006] According to one aspect of this disclosure, a model training method is provided, comprising: inputting a sample query into an inference model to obtain at least one first inference result output by the inference model; obtaining a target inference length corresponding to the sample query, wherein the target inference length is determined based on the shortest inference length among the inference lengths of each correct inference result generated by the inference model for the sample query, and the inference length indicates the number of output tokens in the process of the inference model generating the corresponding inference result; for each of the at least one first inference results, determining a reward value corresponding to the first inference result based on the target inference length, the first inference length, and a correctness index, wherein the first inference length is the inference length corresponding to the first inference result, and when the correctness index indicates that the first inference result is a correct inference result, the reward value when the first inference length is greater than the target inference length is less than the reward value when the first inference length is not greater than the target inference length; calculating a target loss based on the reward value corresponding to each of the at least one first inference results; and updating the parameters of the inference model based on the target loss.
[0007] According to another aspect of this disclosure, a method for performing a target reasoning task is provided, comprising: obtaining a target query; and inputting the target query into a target reasoning model to obtain a reasoning result output by the target reasoning model, wherein the target reasoning model is trained using the model training method of this disclosure.
[0008] According to another aspect of this disclosure, a model training apparatus is provided, comprising: a first acquisition unit configured to input a sample query into an inference model to obtain at least one first inference result output by the inference model; a second acquisition unit configured to acquire a target inference length corresponding to the sample query, wherein the target inference length is determined based on the shortest inference length among the inference lengths of each correct inference result generated by the inference model for the sample query, and the inference length indicates the number of output tokens in the process of the inference model generating the corresponding inference result; a first determination unit configured to determine a reward value corresponding to each of the at least one first inference result based on the target inference length, the first inference length, and a correctness index, wherein the first inference length is the inference length corresponding to the first inference result, and when the correctness index indicates that the first inference result is a correct inference result, the reward value when the first inference length is greater than the target inference length is less than the reward value when the first inference length is not greater than the target inference length; a first calculation unit configured to calculate a target loss based on the reward value corresponding to each of the at least one first inference result; and a first update unit configured to update the parameters of the inference model based on the target loss.
[0009] According to another aspect of this disclosure, an execution apparatus for a target reasoning task is provided, comprising: a first acquisition unit configured to acquire a target query; and a second acquisition unit configured to input the target query into a target reasoning model to obtain a reasoning result output by the target reasoning model, wherein the target reasoning model is trained using the model training method of this disclosure.
[0010] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the model training method of this disclosure or the execution method of the target inference task of this disclosure.
[0011] According to another aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the model training method of this disclosure or the execution method of the target inference task of this disclosure.
[0012] According to another aspect of this disclosure, a computer program product is provided, including a computer program, wherein the computer program, when executed by a processor, implements the model training method of this disclosure or the execution method of the target inference task of this disclosure.
[0013] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0014] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0015] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein may be implemented according to embodiments of the present disclosure is shown; Figure 2 A flowchart of a model training method according to an embodiment of the present disclosure is shown; Figure 3 A flowchart illustrating the process of obtaining the target inference length of a sample query according to an embodiment of the present disclosure is shown; Figure 4 A flowchart illustrating the determination of a reward value according to an embodiment of the present disclosure is shown; Figure 5A flowchart illustrating the calculation of the target loss according to an embodiment of the present disclosure is shown; Figure 6 A flowchart illustrating the calculation of the target loss according to an embodiment of the present disclosure is shown; Figure 7 A flowchart illustrating the determination of lexical loss terms according to embodiments of the present disclosure is shown; Figure 8 A schematic flowchart of a model training method according to an exemplary embodiment of the present disclosure is shown; Figure 9 A flowchart illustrating a method for performing a target reasoning task according to an embodiment of the present disclosure is shown; Figure 10 A structural block diagram of a model training apparatus according to an embodiment of the present disclosure is shown; Figure 11 A structural block diagram of an execution apparatus for a target reasoning task according to an embodiment of the present disclosure is shown; Figure 12 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0016] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0017] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.
[0018] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.
[0019] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0020] Figure 1A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1 The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.
[0021] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of model training methods or target inference tasks of this disclosure.
[0022] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105 and / or 106 under a Software as a Service (SaaS) model.
[0023] exist Figure 1 In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the various methods described herein, and is not intended to be limiting.
[0024] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to obtain sample data or issue targeted queries. The client devices can provide an interface that allows users to interact with them. The client devices can also output information to users through this interface. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.
[0025] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.
[0026] Network 110 can be any type of network well known to those skilled in the art, and can support data communication using any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0027] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.
[0028] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0029] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105 and / or 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105 and / or 106.
[0030] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.
[0031] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located away from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.
[0032] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.
[0033] Figure 1The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.
[0034] According to embodiments of this disclosure, such as Figure 2 As shown, a model training method is provided, including: step S201, inputting a sample query into an inference model to obtain at least one first inference result output by the inference model; step S202, obtaining the target inference length corresponding to the sample query, wherein the target inference length is determined based on the shortest inference length among the inference lengths of each correct inference result generated by the inference model for the sample query, and the inference length indicates the number of output tokens in the process of the inference model generating the corresponding inference result; step S203, for each first inference result in the at least one first inference result, determining the reward value corresponding to the first inference result based on the target inference length, the first inference length, and the correctness index, wherein the first inference length is the inference length corresponding to the first inference result, and when the correctness index indicates that the first inference result is a correct inference result, the reward value when the first inference length is greater than the target inference length is less than the reward value when the first inference length is not greater than the target inference length; step S204, calculating the target loss based on the reward value corresponding to each first inference result in the at least one first inference result; and step S205, updating the parameters of the inference model based on the target loss.
[0035] Therefore, by introducing a target inference length that is dynamically determined based on the shortest correct inference length, when determining the reward value, a smaller reward is applied to correct inference results that exceed the target inference length, while a larger reward is given to results that meet or fall below the target inference length. This allows the inference model to learn to reduce redundant output while maintaining task correctness, thereby achieving the effect of reducing the word output length in the inference process while ensuring the accuracy of the inference model, thus improving computational efficiency.
[0036] In some embodiments, the inference model can refer to a deep neural network model with logical reasoning or complex problem-solving capabilities. Such inference models can be large language models or multimodal large models, etc. The query input into the inference model can be a description of a reasoning requirement for a target task sent by the target object (such as a user), and the inference result output by the inference model can be the response to that reasoning requirement.
[0037] In specific application scenarios, inference models can output corresponding inference results for different types of queries. For example, in the office scenario of intelligent inference assistants, the query can be a description of specific task requirements, such as logical reasoning requirements for a complex business, in-depth data analysis tasks for financial statements, or process optimization problems involving multi-department collaboration. The inference result can be a logical chain containing detailed derivation steps, data insight conclusions, and decision-making basis. In code-assisted writing scenarios, the query can be a description of specific programming requirements, and the inference result can be a source code fragment that meets the grammatical requirements. In multimodal understanding scenarios, the query can be an image and its corresponding query text, and the inference result can be a logical explanation text generated based on image content analysis.
[0038] In some embodiments, inference length indicates the number of output tokens in the process of the inference model generating the corresponding inference result. Inference length reflects the effort and time the model spends processing the sample query. Generally, the longer the inference length, the more inference steps or the longer the thought process of the model.
[0039] In some embodiments, the target inference length corresponding to a sample query can be determined based on global historical data. Specifically, the target inference length can be the shortest inference length among all correct inference results generated by the current inference model for this specific sample query. This data range includes not only the inference length of the correct inference result generated by the inference model when the input sample query is executed, but also the inference lengths of all correct inference results generated by the model for this specific sample query in all historical inputs.
[0040] By using the shortest inference length in the global history as a benchmark, the inference model can continuously move closer to the optimal, concise generation path. During this process, the target inference length is automatically updated as the model's performance evolves, ensuring that the inference model always uses the best performance from the historical generation process as a benchmark during training, thereby achieving continuous improvement in the model's inference efficiency.
[0041] In some embodiments, the correct inference result in the first inference result can be determined by determining whether the first inference result matches the sample inference result corresponding to the sample query. In some embodiments, the correctness of the first inference result can also be verified by another high-performance inference model, which is not limited here.
[0042] In some embodiments, such as Figure 3As shown, obtaining the target inference length corresponding to the sample query may include: step S301, obtaining the historical experience inference length, wherein the historical experience inference length is the preset initial length or the shortest inference length among the inference lengths of each correct inference result generated by the inference model for the sample query before generating at least one first inference result; step S302, obtaining the shortest inference length corresponding to each correct inference result among at least one first inference result, as the current shortest inference length; and step S303, determining the smaller value between the current shortest inference length and the historical experience inference length as the benchmark shortest inference length, so as to determine the target inference length based on the benchmark shortest inference length.
[0043] Therefore, by comparing the current shortest inference length with the historical inference length during the process of obtaining the target inference length, and taking the smaller value as the benchmark shortest inference length, the target inference length used to determine the reward value can be determined. This enables dynamic convergence and continuous optimization of the target inference length, thereby continuously using historical training experience to calibrate the current target, more accurately defining the optimization boundary of each sample query, and effectively preventing learning bias caused by excessively high or inaccurate reference standards in the early stages of model training.
[0044] In some embodiments, the above model training method may further include: in response to determining that the current shortest inference length is less than the historical experience inference length, updating the historical experience inference length based on the current shortest inference length.
[0045] Therefore, by determining that the current shortest inference length is less than the historical inference length and updating the historical inference length in a timely manner, it is possible to ensure that the inference model always uses the best performance in the historical generation process as a reference benchmark, so that the target inference length can automatically evolve downward as the model's capabilities improve, thereby achieving continuous length optimization and efficiency improvement in the model training process.
[0046] Figure 4 A schematic flowchart of a model training method according to an exemplary embodiment of the present disclosure is shown.
[0047] In some embodiments, combined with Figure 4 As shown, the historical inference length can be used to record the best performance of the inference model for a sample query in past training cycles, that is, the shortest inference length in the historical inference of the model before the current inference round when the correct inference result is output, i.e. Figure 4 In .
[0048] In some embodiments, the historical experience inference length can be stored in a data structure such as a global dictionary, i.e., as an experience buffer, to enable fine-grained management of queries for each sample in the sample dataset.
[0049] In some embodiments, during the initial stage of model training, if the inference model has not yet generated any correct inference results for a particular sample query, the historical experience inference length can be initialized to a preset initial length. This preset initial length can be set to the maximum inference length allowed for the inference model to output, ensuring sufficient space for the model to generate thought chains in the early stages of training. It is understood that the aforementioned preset initial length can be determined by those skilled in the art based on their actual needs, and is not limited herein.
[0050] Through the records maintained by the experience buffer, the model training process can achieve differentiated training for different sample queries. Since the historical experience inference length corresponding to each sample query is stored and updated independently, the inference model can learn and master personalized output strategies for sample queries of different complexities. That is, it maintains a very short inference length for simple sample queries, while adjusting it appropriately for complex sample queries while ensuring accuracy. This allows the model to learn to output different output token lengths for sample queries of different complexities.
[0051] In some exemplary embodiments, the update of the historical experience reasoning length can be represented by the following formula:
[0052] in, The inference model is represented in the first... The second query for this sample The determined length of historical experience reasoning after performing the reasoning. The inference model is represented in the first... The second query for this sample The length of historical reasoning experience stored after inference is performed. This represents the current shortest inference length mentioned above. If the inference model does not produce any correct inference results in the current batch, the inner minimum is considered to be infinite, thus keeping the historical inference length unchanged.
[0053] In some embodiments, the process of determining the target inference length based on the baseline shortest inference length is flexible and can be to directly determine the baseline shortest inference length as the target inference length.
[0054] In some embodiments, determining the target inference length based on the baseline shortest inference length may include: determining the target inference length based on the baseline shortest inference length and a preset tolerance parameter, wherein the target inference length is greater than the baseline shortest inference length.
[0055] Therefore, by introducing a preset tolerance parameter and determining the target inference length based on the minimum benchmark inference length, a reasonable optimization margin is set for the inference model. This allows the inference model to have a certain logical generation space while ensuring the correctness of the task, avoiding a decrease in the accuracy of model inference due to excessive pursuit of extremely short output. This ensures the conciseness of the inference model output while further guaranteeing the reliability of the inference results.
[0056] In some exemplary embodiments, see continue to see Figure 4 The target inference length can be represented as ,in, The minimum inference length is represented by the baseline, and α is a preset tolerance parameter that satisfies α∈[0,+∞). For example, α can be set to 0.1.
[0057] In some embodiments, the reward value corresponding to the first inference result, based on the target inference length, the first inference length, and the correctness metric, can be dynamically calculated using a continuous length penalty function. For example, in response to the correctness metric indicating that the first inference result is an incorrect inference result, the reward value corresponding to the first inference result can be determined as a preset base penalty value; in response to the correctness metric indicating that the first inference result is a correct inference result, and the first inference length is not greater than the target inference length, the reward value corresponding to the first inference result can be determined as the maximum base reward value; in response to the correctness metric indicating that the first inference result is a correct inference result, and the first inference length is greater than the target inference length, a continuously decaying penalty can be applied to the maximum base reward value based on the difference between the first inference length and the target inference length to calculate the final reward value corresponding to the first inference result.
[0058] In some embodiments, such as Figure 5 As shown, determining the reward value corresponding to the first inference result based on the target inference length, the first inference length, and the correctness index may include: step S501, in response to the determination that the correctness index indicates the first inference result is an incorrect inference result, determining the reward value corresponding to the first inference result as a first value; step S502, in response to the determination that the correctness index indicates the first inference result is a correct inference result, and the first inference length is greater than the target inference length, determining the reward value corresponding to the first inference result as a second value; and step S503, in response to the determination that the correctness index indicates the first inference result is a correct inference result, and the first inference length is not greater than the target inference length, determining the reward value corresponding to the first inference result as a third value; wherein, the third value is greater than the second value, and the second value is greater than the first value.
[0059] Therefore, by refining the reward value into three levels—first, second, and third values—that decrease sequentially, a hierarchical punishment and incentive system is established. This not only effectively suppresses the generation of erroneous outputs but also provides greater incentives for correct results that meet the target inference length by comparing the inference length of the correct inference results. This strengthens the inference model's preference for concise and correct answers, ensuring the accuracy of the model's inference results while reducing the computational resources consumed by the model's inference and improving the model's inference efficiency.
[0060] In some embodiments, if the correctness indicator indicates that the first inference result is an incorrect inference result, the system determines the reward value corresponding to the first inference result as a first value (i.e., zero reward, which can be assigned a value of 0) to clearly penalize the incorrect inference path; if the correctness indicator indicates that the first inference result is a correct inference result, and the first inference length exceeds the target inference length, the system determines the reward value as a second value (i.e., a discounted reward, which can be assigned a value of 0). ),in A preset discount coefficient greater than 0 and less than 1 can be used to ensure that lengthy but correct responses still receive positive signals, preventing the inference model from losing its unique learning motivation when dealing with difficult problems because it has not yet found the shortest correct solution. If the correctness index indicates that the first inference result is a correct inference result and the length of the first inference is not greater than the target inference length, the system will determine the reward value as the third value (i.e., full reward, which can be assigned a value of 1.0) to maximize the incentive for the model to produce high-quality answers that meet efficiency requirements.
[0061] In some exemplary embodiments, see continue to see Figure 4 The above reward value can be determined by the following formula:
[0062] in, Indicates the first The reward value corresponding to the first reasoning result This represents the first result of reasoning. This indicates the correct inference result corresponding to the sample query. This represents the target inference length, where α represents the preset tolerance parameter. This indicates the discount / reward coefficient.
[0063] In some embodiments, the number of at least one first inference result can be multiple, such as Figure 6As shown, calculating the target loss based on the reward value corresponding to each of the at least one first inference results may include: step S601, obtaining the baseline reward value of at least one first inference result and the target number of correct inference results in the at least one first inference result; step S602, for each of the at least one first inference results, calculating the advantage value corresponding to the first inference result based on the reward value corresponding to the first inference result, the baseline reward value, and the target number, wherein the advantage value is positively correlated with the difference between the reward value corresponding to the first inference result and the baseline reward value, and the target number is used to attenuate the magnitude of the advantage value; and step S603, calculating the target loss based on the advantage value corresponding to each of the at least one first inference results.
[0064] Therefore, by calculating the advantage value based on the baseline reward value of multiple first inference results and the target number of correct inference results, and by introducing the target number of correct inference results in the first inference results to adjust the magnitude of the advantage value, the magnitude of gradient update can be dynamically adjusted according to the difficulty of sample query. That is, for samples with a large number of correct inference results, the fluctuation magnitude of gradient can be effectively reduced, so that in inference tasks with low accuracy (difficulty), larger gradients are prioritized to improve accuracy, while in inference tasks with high accuracy (easy), the advantage difference between long and short responses is used as the dominant signal to shorten the output length, further improving the accuracy and computational efficiency of the inference model.
[0065] In some embodiments, the aforementioned benchmark reward value may be an exponential moving average of the reward values from several inferences for a sample query throughout history. In some embodiments, the aforementioned benchmark reward value may also be a preset benchmark value.
[0066] In some embodiments, the baseline reward value may be the average reward value of at least one of the first inference results described above.
[0067] Therefore, by using the average reward value of multiple first inference results as the benchmark reward value, a more objective and stable group evaluation reference point is provided, which can effectively eliminate the influence of the randomness of a single inference on the evaluation of the advantage value. This enables the inference model to perform differentiated learning based on the overall generation quality within the group, and further improves the consistency of the strategy performance of the inference model when dealing with different samples.
[0068] In some embodiments, for each of the at least one first inference results, calculating the advantage value corresponding to the first inference result based on the reward value, the baseline reward value, and the target quantity can include: calculating the advantage value of the first inference result among the at least one first inference result based on the following formula. The first inference result, corresponding to the advantage value :
[0069] in, Indicates the first The reward value corresponding to the first reasoning result Indicates the baseline reward value. Indicates the target quantity. It is a numerical stability constant that is greater than zero.
[0070] Therefore, by using the above-mentioned advantage value calculation method, the gradient update magnitude is dynamically adjusted according to the difficulty of the sample query. This allows for the priority acquisition of larger gradients to improve accuracy in inference tasks with low accuracy (difficulty), while in inference tasks with high accuracy (easy), the advantage difference between long and short responses is used as the dominant signal to shorten the output length, further improving the accuracy and computational efficiency of the inference model.
[0071] In some exemplary embodiments, the baseline reward value It can be calculated based on the following formula:
[0072] Wherein, G is the first inference result obtained by the model inference for the sample query q. The total number.
[0073] In some exemplary embodiments, the target quantity It can be expressed based on the following formula: .
[0074] In some exemplary embodiments, For example, it can be set to 1 or other constants greater than zero; there are no restrictions here.
[0075] In the above calculation method, by introducing the target quantity as the denominator in the advantage value calculation, a difficulty-adaptive advantage estimation mechanism is achieved. Related technologies typically use standard deviation for normalization. The standard deviation reaches its maximum when the number of correct inference results approaches half the total, and then tends to zero at both ends. This non-monotonic characteristic can disrupt the inference model's perception of the sample query difficulty. However, in the embodiments of this disclosure, the target quantity is a monotonic surrogate for the accuracy rate, ensuring that the magnitude of the advantage value calculation strictly decreases as the target quantity increases. This allows the gradient update magnitude to strictly and accurately reflect the current sample query difficulty level, guaranteeing that feedback signals related to the question difficulty can be transmitted all the way to the update process of the inference model parameters.
[0076] Specifically, when the sample query is a difficult problem, the inference model generates fewer correct inference results, resulting in a smaller number of objectives. This leads to a smaller denominator in the formula and a larger overall magnitude of the calculated advantage value. In this case, a larger gradient update magnitude will focus on improving the accuracy of the inference model in generating correct inference results. Conversely, when the sample query is a simple problem, the inference model generates more correct inference results, resulting in a larger number of objectives. This leads to a larger denominator in the formula and a smaller overall magnitude of the advantage value. In this case, although the absolute magnitude of the advantage value decreases, the difference in reward value between correct inference results with an inference length no greater than the target inference length and correct inference results with an inference length greater than the target inference length dominates the advantage value difference signal. This causes the gradient update to focus on promoting a more concise output from the inference model.
[0077] Through the aforementioned computational mechanism, the embodiments of this disclosure naturally generate a learning dynamic during the model training phase that first improves accuracy and then compresses inference length. This dynamic mechanism allows the inference model to be automatically pushed toward exploring concise outputs only on problems where correctness is already reliable, while on problems where correctness is still a bottleneck, it is continuously pushed toward improving accuracy. The entire difficulty adaptation process is entirely guided by experience, without any manual coefficient switching operations.
[0078] In some embodiments, such as Figure 7 As shown, calculating the target loss based on the advantage value corresponding to each of the at least one first inference results may include: Step S701, for each output word corresponding to each of the at least one first inference results, performing the following operations: Step S7011, determining the word advantage value of the output word based on the advantage value corresponding to the first inference result; and Step S7012, obtaining the first generation probability of the output word under the inference model and the second generation probability under the reference model, so as to determine the original probability ratio corresponding to the output word based on the first generation probability and the second generation probability, wherein the reference model is the inference model in the initial state; and Step S7013, determining the word loss term corresponding to the output word based on the original probability ratio and word advantage value corresponding to the output word; and Step S702, calculating the target loss based on the word loss term of each output word corresponding to each of the at least one first inference results.
[0079] Therefore, by independently calculating the lexical advantage value and probability ratio for each output lexical in each first inference result, fine-grained control over the model's inference process is achieved. By refining the evaluation scale to the lexical level, the inference model can clearly identify which specific lexical's generation has a positive or negative impact on the overall inference effect. This allows for more precise guidance during training, significantly improving the inference model's ability to control the generation path.
[0080] In some embodiments, such as Figure 8 As shown, determining the lexical loss term corresponding to the output lexical based on the original probability ratio and lexical advantage value can include: Step S801, obtaining a lower bound pruning threshold and an upper bound pruning threshold to limit the original probability ratio, wherein the absolute value of the difference between the upper bound pruning threshold and a preset benchmark value is greater than the absolute value of the difference between the preset benchmark value and the lower bound pruning threshold; Step S802, performing asymmetric pruning on the original probability ratio using the lower bound pruning threshold and the upper bound pruning threshold to obtain a pruned probability ratio; Step S803, calculating the first product of the original probability ratio and the lexical advantage value, and the second product of the pruned probability ratio and the lexical advantage value, respectively; and Step S804, determining the smaller value between the first product and the second product as the lexical loss term corresponding to the output lexical.
[0081] Therefore, by setting an asymmetric pruning threshold based on a preset benchmark value and performing asymmetric pruning on the original probability ratios, a parameter update constraint mechanism that balances exploratory nature and stability is constructed. This mechanism can preserve the model's motivation to make large strides on the correct path while strictly limiting the step size of parameter updates through effective truncation of the probability ratios. This effectively prevents performance collapse caused by a single large update, achieving an optimized balance between training stability and convergence speed.
[0082] In some exemplary embodiments, for each output lexical corresponding to each of the at least one first inference results, the system can first determine the lexical advantage value of the output lexical based on the advantage value corresponding to the first inference result. Subsequently, the system obtains the first generation probability of the output lexical under the current inference model and the second generation probability under the reference model, and compares and divides the first generation probability and the second generation probability to determine the original probability ratio corresponding to the output lexical. Here, the reference model is the inference model in its initial state, i.e., the inference model before any sample queries are applied for model training, and its parameters remain fixed during the current training phase, serving as a stable probability benchmark for calculating the original probability ratio.
[0083] After calculating the original probability ratio, the system can obtain lower and upper bound pruning thresholds to limit the original probability ratio, thus performing asymmetric pruning. The core characteristic of this asymmetric pruning is that the absolute value of the difference between the upper bound pruning threshold and the preset baseline value is greater than the absolute value of the difference between the preset baseline value and the lower bound pruning threshold. For example, when the preset baseline value is 1, the relaxation of the upper bound is strictly greater than the contraction of the lower bound. This asymmetric threshold setting allows the inference model to make larger parameter updates when evolving in a favorable direction, while maintaining a conservative lower bound constraint in unfavorable directions. The system uses the aforementioned lower and upper bound pruning thresholds to truncate the original probability ratio, obtaining the pruned probability ratio.
[0084] Next, the system calculates the first product of the original probability ratio and the lexical advantage value, and the second product of the pruning probability ratio and the lexical advantage value. The system compares these two products and determines the smaller value as the lexical loss term corresponding to the output lexical. By selecting the smaller value, the system prevents the inference model from greedily and excessively updating its parameters due to overestimating a particular generation action. Finally, the system accumulates and normalizes the lexical loss terms of all output lexical terms in all first inference results to calculate the target loss used to finally update the inference model parameters. By maximizing this target loss and adjusting the inference model's parameters, the inference model can be trained.
[0085] In some exemplary embodiments, the calculation of the target loss described above can be expressed by the following formula:
[0086] in, Let G represent the target loss, and G represent the number of first inference results generated. This indicates the number of output tokens contained in the i-th first inference result. This indicates that the number of output tokens for all first inference results is summed over a traversal. Used to perform word-level normalization on the summation result.
[0087] Let represent the original probability ratio corresponding to the t-th output word in the i-th first inference result, which can be calculated using the following formula:
[0088] in, This represents the inference model currently in the training state. (i.e., the inference model whose parameters are being updated) generates the first generation probability of this output term when processing the sample query q. Representation of reference model The second generation probability of this output term is generated when processing sample query q.
[0089] It represents the lexical advantage value corresponding to the t-th output lexical in the i-th first inference result. This indicates using the lower bound clipping threshold. and upper bound clipping threshold The clipping probability ratio obtained after asymmetric clipping of the original probability ratio, where the preset baseline value is 1, and and These represent the deviations from the lower and upper bounds, respectively. < The min function selects the smaller value between the uncropped first product and the asymmetric pruning second product as the lexical loss term for the corresponding output lexical. This indicates the goal of obtaining the expected value.
[0090] In some exemplary embodiments, taking a code-assisted writing scenario as an example, the inference model can be a large language model pre-trained specifically for programming languages and code logic. Sample queries can be specific descriptions of programming requirements, such as natural language functional requirements input by developers, partial code snippets, or function definition requirements. Correspondingly, the inference results can be source code snippets that meet grammatical requirements, as well as thought-chain content output by the inference model before generating code, used to organize the logical architecture.
[0091] The training method for the aforementioned large language model used for code-assisted programming may include: First, inputting a sample programming requirement description as a sample query into the large language model to obtain at least one first inference result from the large language model, i.e., at least one candidate source code fragment and its associated thought chain. Then, the system obtains the target inference length corresponding to the programming requirement description. For each candidate source code fragment, the system determines the reward value corresponding to that candidate source code fragment based on the target inference length, the first inference length, and the correctness metric. In this process, the correctness metric can be determined based on the results of automated unit testing.
[0092] If the unit test results indicate that the candidate source code snippet contains syntax errors or fails to implement the expected functionality, meaning the correctness metric indicates it is an incorrect inference result, the system will assign the first reward value. If the unit test results indicate that the candidate source code snippet is a correct inference result that correctly implements the programming requirements, the system will further compare the first inference length with the target inference length. If the first inference length of the candidate source code snippet is greater than the target inference length, the system will assign the second reward value; if the first inference length is not greater than the target inference length, the system will assign the third reward value.
[0093] Finally, based on the reward values corresponding to each candidate source code fragment output by the above inference model, the system calculates the target loss using the method described above, and then trains the above large language model based on the target loss to obtain a target inference model that can accurately and efficiently generate source code fragments that meet user needs.
[0094] In some embodiments, such as Figure 9 As shown, a method for executing a target reasoning task is also provided, including: step S901, obtaining a target query; and step S902, inputting the target query into a target reasoning model to obtain the reasoning result output by the target reasoning model, wherein the target reasoning model is trained using the above-mentioned model training method.
[0095] In some exemplary embodiments, the target inference task described above may be, for example, a code-assisted writing task, and the target inference model may be obtained through the large language model training method for code-assisted programming described above. Users can input their programming requirements into the target inference model to obtain the target source code fragment output by the model.
[0096] In some embodiments, such as Figure 10 As shown, a model training apparatus 1000 is provided. The apparatus 1000 includes: a first acquisition unit 1010 configured to input a sample query into an inference model to obtain at least one first inference result output by the inference model; a second acquisition unit 1020 configured to acquire a target inference length corresponding to the sample query, wherein the target inference length is determined based on the shortest inference length among the inference lengths of each correct inference result generated by the inference model for the sample query, and the inference length indicates the number of output tokens in the process of the inference model generating the corresponding inference result; and a first determination unit 1030 configured to determine each first inference result in the at least one first inference result. The inference result, based on the target inference length, the first inference length, and the correctness index, determines the reward value corresponding to the first inference result, wherein the first inference length is the inference length corresponding to the first inference result. When the correctness index indicates that the first inference result is a correct inference result, the reward value when the first inference length is greater than the target inference length is less than the reward value when the first inference length is not greater than the target inference length; the first calculation unit 1040 is configured to calculate the target loss based on the reward value corresponding to each of the at least one first inference result; and the first update unit 1050 is configured to update the parameters of the inference model based on the target loss.
[0097] The operations performed by units 1010 to 1050 in the model training device 1000 and the technical effects they can achieve are similar to steps S201 to S205 in the model training method, and will not be described in detail here.
[0098] In some embodiments, the second acquisition unit may be further configured to: acquire the historical experience inference length, wherein the historical experience inference length is a preset initial length or the shortest inference length among the inference lengths of each correct inference result generated by the inference model for the sample query before generating at least one first inference result; acquire the shortest inference length corresponding to each correct inference result in at least one first inference result as the current shortest inference length; and determine the smaller value between the current shortest inference length and the historical experience inference length as the benchmark shortest inference length, so as to determine the target inference length based on the benchmark shortest inference length.
[0099] In some embodiments, determining the target inference length based on the baseline shortest inference length may include: determining the target inference length based on the baseline shortest inference length and a preset tolerance parameter, wherein the target inference length is greater than the baseline shortest inference length.
[0100] In some embodiments, the above-described model training apparatus may further include: a second update unit configured to update the historical inference length based on the current shortest inference length in response to the current shortest inference length being less than the historical inference length.
[0101] In some embodiments, the first determining unit may be further configured to: determine the reward value corresponding to the first inference result as a first value in response to a determination correctness indicator indicating that the first inference result is an incorrect inference result; determine the reward value corresponding to the first inference result as a second value in response to a determination correctness indicator indicating that the first inference result is a correct inference result and the first inference length is greater than the target inference length; and determine the reward value corresponding to the first inference result as a third value in response to a determination correctness indicator indicating that the first inference result is a correct inference result and the first inference length is not greater than the target inference length; wherein the third value is greater than the second value, and the second value is greater than the first value.
[0102] In some embodiments, the number of at least one first inference result can be multiple, and the first calculation unit can be further configured to: obtain a baseline reward value for at least one first inference result and a target number of correct inference results among the at least one first inference result; for each of the at least one first inference result, calculate an advantage value corresponding to the first inference result based on the reward value corresponding to the first inference result, the baseline reward value, and the target number, wherein the advantage value is positively correlated with the difference between the reward value corresponding to the first inference result and the baseline reward value, and the target number is used to attenuate the magnitude of the advantage value; and calculate a target loss based on the advantage value corresponding to each of the at least one first inference result.
[0103] In some embodiments, the baseline reward value is the average reward value of at least one first inference result.
[0104] In some embodiments, for each of the at least one first inference results, calculating the advantage value corresponding to the first inference result based on the reward value, the baseline reward value, and the target quantity can include: calculating the advantage value of the first inference result among the at least one first inference result based on the following formula. The first inference result, corresponding to the advantage value :
[0105] in, Indicates the first The reward value corresponding to the first reasoning result Indicates the baseline reward value. Indicates the target quantity. It is a numerical stability constant that is greater than zero.
[0106] In some embodiments, calculating the target loss based on the advantage value corresponding to each of the at least one first inference results may include: for each output lexical corresponding to each of the at least one first inference results, performing the following operations: determining the lexical advantage value of the output lexical based on the advantage value corresponding to the first inference result; obtaining a first generation probability of the output lexical under the inference model and a second generation probability under the reference model, so as to determine the original probability ratio corresponding to the output lexical based on the first generation probability and the second generation probability, wherein the reference model is the inference model in the initial state; determining the lexical loss term corresponding to the output lexical based on the original probability ratio and the lexical advantage value; and calculating the target loss based on the lexical loss term corresponding to each output lexical corresponding to each of the at least one first inference results.
[0107] In some embodiments, determining the lexical loss term corresponding to the output lexical based on the original probability ratio and lexical advantage value may include: obtaining a lower bound pruning threshold and an upper bound pruning threshold for limiting the original probability ratio, wherein the absolute value of the difference between the upper bound pruning threshold and a preset benchmark value is greater than the absolute value of the difference between the preset benchmark value and the lower bound pruning threshold; performing asymmetric pruning on the original probability ratio using the lower bound pruning threshold and the upper bound pruning threshold to obtain a pruned probability ratio; calculating a first product of the original probability ratio and the lexical advantage value, and a second product of the pruned probability ratio and the lexical advantage value, respectively; and determining the smaller value between the first product and the second product as the lexical loss term corresponding to the output lexical.
[0108] In some embodiments, such as Figure 11 As shown, an execution device 1100 for a target reasoning task is provided. The device 1100 includes: a first acquisition unit 1110 configured to acquire a target query; and a second acquisition unit 1120 configured to input the target query into a target reasoning model to obtain the reasoning result output by the target reasoning model, wherein the target reasoning model is trained using the model training method described above.
[0109] The operations performed by units 1110 and 1120 in the execution device 1100 for the aforementioned target reasoning task, as well as the technical effects they can achieve, are similar to steps S901 and S902 in the execution method for the aforementioned target reasoning task, and will not be described in detail here.
[0110] According to embodiments of this disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.
[0111] refer to Figure 12 The present invention describes a structural block diagram of an electronic device 1200 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0112] like Figure 12As shown, the electronic device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1202 or a computer program loaded from a storage unit 1208 into a random access memory (RAM) 1203. The RAM 1203 may also store various programs and data required for the operation of the electronic device 1200. The computing unit 1201, ROM 1202, and RAM 1203 are interconnected via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.
[0113] Multiple components in electronic device 1200 are connected to I / O interface 1205, including: input unit 1206, output unit 1207, storage unit 1208, and communication unit 1209. Input unit 1206 can be any type of device capable of inputting information to electronic device 1200. Input unit 1206 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of the electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 1207 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 1208 may include, but is not limited to, a hard disk and an optical disk. The communication unit 1209 allows the electronic device 1200 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers and / or chipsets, such as Bluetooth devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices and / or the like.
[0114] The computing unit 1201 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1201 performs the various methods and processes described above, such as the above-described model training method or the execution method of the target inference task. For example, in some embodiments, the above-described model training method or the execution method of the target inference task can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 1208. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 1200 via ROM 1202 and / or communication unit 1209. When the computer program is loaded into RAM 1203 and executed by the computing unit 1201, one or more steps of the above-described model training method or the execution method of the target inference task can be performed. Alternatively, in other embodiments, the computing unit 1201 may be configured by any other suitable means (e.g., by means of firmware) to perform the above-described model training method or the execution method of the target inference task.
[0115] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0116] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0117] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0118] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0119] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0120] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0121] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0122] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.
Claims
1. A model training method, the method comprising: Input a sample query into the inference model to obtain at least one first inference result output by the inference model; Obtain the target inference length corresponding to the sample query, wherein the target inference length is determined based on the shortest inference length among the inference lengths of each correct inference result generated by the inference model for the sample query, and the inference length indicates the number of output tokens in the process of the inference model generating the corresponding inference result; For each of the at least one first reasoning results, a reward value corresponding to the first reasoning result is determined based on the target reasoning length, the first reasoning length, and the correctness index. The first reasoning length is the reasoning length corresponding to the first reasoning result. When the correctness index indicates that the first reasoning result is a correct reasoning result, the reward value when the first reasoning length is greater than the target reasoning length is less than the reward value when the first reasoning length is not greater than the target reasoning length. Based on the reward value corresponding to each of the at least one first inference results, calculate the target loss; and The parameters of the inference model are updated based on the target loss.
2. The method according to claim 1, wherein, The step of obtaining the target inference length corresponding to the sample query includes: Obtain the historical experience inference length, wherein the historical experience inference length is a preset initial length or the shortest inference length among the inference lengths of each correct inference result generated by the inference model for the sample query before generating the at least one first inference result; Obtain the shortest inference length corresponding to each correct inference result among the at least one first inference result, and use it as the current shortest inference length; and The smaller of the current shortest inference length and the historical experience inference length is determined as the baseline shortest inference length, and the target inference length is determined based on the baseline shortest inference length.
3. The method according to claim 2, wherein, Determining the target inference length based on the benchmark shortest inference length includes: Based on the baseline shortest inference length and the preset tolerance parameter, the target inference length is determined, wherein the target inference length is greater than the baseline shortest inference length.
4. The method according to claim 2 or 3, further comprising: In response to determining that the current shortest inference length is less than the historical experience inference length, the historical experience inference length is updated based on the current shortest inference length.
5. The method according to any one of claims 1 to 4, wherein, The step of determining the reward value corresponding to the first reasoning result based on the target reasoning length, the first reasoning length, and the correctness index includes: In response to determining that the correctness indicator indicates that the first reasoning result is an incorrect reasoning result, the reward value corresponding to the first reasoning result is determined as the first value; In response to determining that the correctness indicator indicates the first inference result is a correct inference result, and that the first inference length is greater than the target inference length, the reward value corresponding to the first inference result is determined as a second value; and In response to determining that the correctness indicator indicates the first inference result is a correct inference result, and that the first inference length is not greater than the target inference length, the reward value corresponding to the first inference result is determined as the third value; wherein, The third value is greater than the second value, and the second value is greater than the first value.
6. The method according to any one of claims 1 to 5, wherein, The number of the at least one first inference result is multiple, and the calculation of the target loss based on the reward value corresponding to each of the at least one first inference result includes: Obtain the baseline reward value of the at least one first reasoning result and the target number of correct reasoning results among the at least one first reasoning result; For each of the at least one first inference results, based on the reward value corresponding to the first inference result, the baseline reward value, and the target quantity, an advantage value corresponding to the first inference result is calculated, wherein the advantage value is positively correlated with the difference between the reward value corresponding to the first inference result and the baseline reward value, and the target quantity is used to attenuate the magnitude of the advantage value; and The target loss is calculated based on the advantage value corresponding to each of the at least one first inference results.
7. The method according to claim 6, wherein, The baseline reward value is the average reward value of the at least one first inference result.
8. The method according to claim 6 or 7, wherein, The step of calculating the advantage value corresponding to each of the at least one first inference result, based on the reward value corresponding to the first inference result, the baseline reward value, and the target quantity, includes: The first inference result is calculated based on the following formula. The advantage value corresponding to the first reasoning result : in, Indicates the first The reward value corresponding to the first reasoning result This represents the baseline reward value. This indicates the target quantity. It is a numerical stability constant that is greater than zero.
9. The method according to any one of claims 6 to 8, wherein, The calculation of the target loss based on the advantage value corresponding to each of the at least one first inference result includes: For each output word corresponding to each of the at least one first inference result, the following operations are performed: Based on the advantage value corresponding to the first inference result, determine the lexical advantage value of the output lexical; and Obtain the first generation probability of the output word under the inference model and the second generation probability under the reference model, and determine the original probability ratio corresponding to the output word based on the first generation probability and the second generation probability, wherein the reference model is the inference model in its initial state; and Based on the original probability ratio and lexical advantage value corresponding to the output lexical, determine the lexical loss term corresponding to the output lexical; and The target loss is calculated based on the lexical loss term of each output lexical corresponding to each of the at least one first inference results.
10. The method according to claim 9, wherein, The step of determining the lexical loss term corresponding to the output lexical based on the original probability ratio and lexical advantage value includes: Obtain a lower bound pruning threshold and an upper bound pruning threshold for limiting the original probability ratio, wherein the absolute value of the difference between the upper bound pruning threshold and a preset benchmark value is greater than the absolute value of the difference between the preset benchmark value and the lower bound pruning threshold. The original probability ratio is asymmetrically clipped using the lower bound clipping threshold and the upper bound clipping threshold to obtain the clipping probability ratio. Calculate the first product of the original probability ratio and the lexical advantage value, and the second product of the trimming probability ratio and the lexical advantage value; and The smaller value between the first product and the second product is determined as the lexical loss term corresponding to the output lexical.
11. A method for performing a target reasoning task, the method comprising: Retrieve the target query; as well as The target query is input into the target inference model to obtain the inference result output by the target inference model, wherein the target inference model is trained using the model training method as described in any one of claims 1 to 10.
12. A model training apparatus, the apparatus comprising: The first acquisition unit is configured to input a sample query into the inference model to obtain at least one first inference result output by the inference model; The second acquisition unit is configured to acquire the target inference length corresponding to the sample query, wherein the target inference length is determined based on the shortest inference length among the inference lengths of each correct inference result generated by the inference model for the sample query, and the inference length indicates the number of output tokens in the process of the inference model generating the corresponding inference result; The first determining unit is configured to determine a reward value corresponding to each of the at least one first inference results, based on the target inference length, the first inference length, and the correctness index, wherein the first inference length is the inference length corresponding to the first inference result, and when the correctness index indicates that the first inference result is a correct inference result, the reward value when the first inference length is greater than the target inference length is less than the reward value when the first inference length is not greater than the target inference length. A first calculation unit is configured to calculate a target loss based on the reward value corresponding to each of the at least one first inference results; and The first update unit is configured to update the parameters of the inference model based on the target loss.
13. An execution apparatus for a target reasoning task, the apparatus comprising: The first acquisition unit is configured to acquire the target query; as well as The second acquisition unit is configured to input the target query into the target inference model to obtain the inference result output by the target inference model, wherein the target inference model is trained using the model training method as described in any one of claims 1 to 10.
14. An electronic device, comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-11.
15. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-11.
16. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the method of any one of claims 1-11.