Model training method and device, nonvolatile storage medium and electronic equipment

By converting the tasks to be learned into graph structures and using reinforcement learning to train large models, the problem of low accuracy in execution of large language models in new tasks is solved, and more efficient and flexible task processing is achieved.

CN120338097APending Publication Date: 2025-07-18CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510361664.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Large language models cannot independently learn execution strategies when facing new tasks, resulting in low execution accuracy and inflexible application.

Method used

Convert the task to be learned into a graph structure, train the big model through graph tasks, clarify the sub-tasks and dependencies, and use reinforcement learning methods to optimize the execution strategy.

Benefits of technology

Improves the execution efficiency and accuracy of large models when facing new tasks and enhances application flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338097A_ABST
    Figure CN120338097A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method and device, a nonvolatile storage medium and electronic equipment. The method comprises the steps that a to-be-learned task in a natural language format and a task target corresponding to the to-be-learned task are obtained, and the task target is a result expected to be output by a large model when the large model is adopted to achieve the to-be-learned task; determining a graph task corresponding to the to-be-learned task, the graph task being a graph structure form of the to-be-learned task, each node in the graph task representing a sub-task in the to-be-learned task, and an edge in the graph task representing a dependency relationship of the plurality of sub-tasks; the graph task is adopted to train the large model, a target model is obtained, and the target model is the large model which learns the execution strategy of the to-be-learned task. According to the method and the device, the technical problem that the accuracy is low when the large language model executes the new task due to the fact that the large language model cannot autonomously learn the execution strategy of the new task when facing the new task in the related technology is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology. Specifically, it relates to a method and device for model training, a non-volatile storage medium, and an electronic device. Background Art

[0002] With the development of deep learning technology, large language models based on the Transformer architecture have achieved remarkable results in tasks such as natural language processing and multi-modal understanding. However, the above large language models rely on static training data. When dealing with new tasks not involved in the training data, they are unable to make autonomous decisions and lack the ability to autonomously layout and plan the multi-stage execution of new tasks, resulting in a low accuracy rate of the execution results output when facing new tasks.

[0003] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention

[0004] Embodiments of this application provide a method and device for model training, a non-volatile storage medium, and an electronic device to at least solve the technical problem of low accuracy rate when a large language model in related technologies executes new tasks because it is unable to autonomously learn the execution strategy of new tasks.

[0005] According to one aspect of the embodiments of this application, a method for model training is provided, including: obtaining a to-be-learned task in natural language format and a task objective corresponding to the to-be-learned task, where the task objective is the result expected to be output by a large model when implementing the to-be-learned task; determining a graph task corresponding to the to-be-learned task, where the graph task is a graph structure form of the to-be-learned task, each node in the graph task represents a subtask in the to-be-learned task, and the edges in the graph task represent the dependency relationships of multiple subtasks; training the large model with the graph task to obtain a target model, where the target model is a large model that has learned the execution strategy of the to-be-learned task.

[0006] Optionally, determining the graph task corresponding to the to-be-learned task includes: parsing the to-be-learned task to obtain a parsing result, where the parsing result includes: an object field, an operation field, and a constraint condition field. The information recorded in the object field is used to indicate the operation object of the to-be-learned task, the information recorded in the operation field is used to indicate the action to be performed by the large model, and the information recorded in the constraint condition field is used to indicate the triggering condition for the large model to perform the action; determining subtasks and dependency relationships according to the parsing result; determining the subtasks as nodes, and determining the dependency relationships as edges, establishing a graph structure corresponding to the to-be-learned task according to the nodes and edges, and taking the task corresponding to the graph structure as the graph task.

[0007] Optionally, determine subtasks and dependencies based on the parsing result. Among them, determining subtasks includes: determining the action corresponding to each operation object according to the object field and the operation field; determining subtasks according to each operation object and its corresponding action.

[0008] Optionally, determine subtasks and dependencies based on the parsing result. Among them, determining dependencies includes: for each subtask, determining the value of the tool call field corresponding to the subtask, where the tool call field is included in the operation field, and the information recorded in the tool call field is used to instruct the large model to call external resources, and the external resources are used to implement functions that the large model cannot implement; when the value of the tool call field is a valid value, determine that there is a dependency between the first type of node and the second type of node, where the first type of node is a subtask whose value of the included tool call field is a valid value, and the second type of node is a subtask that includes an operation to call external resources; when the value of the tool call field is an invalid value, determine that there is no dependency between the second type of node and the third type of node, where the third type of node is a subtask whose value of the included tool call field is an invalid value.

[0009] Optionally, the large model is trained by the following method: determining the actions performed by the large model in the dataset, where the dataset stores multiple groups of historical interaction data, the multiple groups of historical interaction data correspond to multiple historical moments, and each group of historical interaction data includes the historical state of the large model at the historical moment, the historical actions performed by the large model at the historical moment, the reward obtained by the large model after performing the historical actions, and the next state of the large model after obtaining the reward, and the reward is used to indicate the contribution degree of the historical actions performed by the large model to the completion of the historical task goal; determining the function value of the objective function according to the actions performed by the large model and the current state, where the current state is used to indicate the progress of the large model in completing the graph task, and the objective function is used to evaluate the contribution degree of the actions performed by the large model to the completion of the task goal; when the number of consecutive times that the change amount of the function value is less than the preset change amount is greater than or equal to the preset number of times, determine that the training is completed, where the change amount is the difference between the previous function value and the next function value, and the previous function value and the next function value are the function values of the objective function determined in two adjacent iterative training processes.

[0010] Optionally, the multiple groups of historical interaction data include: real interaction data generated when the large model performs historical tasks, and test interaction data generated when the large model performs actions in the test set, where the test set contains multiple preset actions.

[0011] Optionally, determining the actions performed by the large model in the dataset includes: determining the group of historical interaction data corresponding to the largest reward as the target dataset; determining the actions included in the target dataset as the actions performed by the large model.

[0012] According to another aspect of the embodiments of the present application, there is also provided an apparatus for model training, including: an acquisition module, configured to acquire a to-be-learned task in natural language format and a task objective corresponding to the to-be-learned task, where the task objective is the result expected to be output by a large model when implementing the to-be-learned task using the large model; a determination module, configured to determine a graph task corresponding to the to-be-learned task, where the graph task is a graph structure form of the to-be-learned task, each node in the graph task represents a subtask in the to-be-learned task, and the edges in the graph task represent the dependency relationships of multiple subtasks; a training module, configured to train the large model using the graph task to obtain a target model, where the target model is a large model that has learned the execution strategy of the to-be-learned task.

[0013] According to another aspect of the embodiments of the present application, there is also provided a non-volatile storage medium storing a computer program, where the computer program is executed by a device where the non-volatile storage medium is located to perform the above-mentioned method for model training.

[0014] According to another aspect of the embodiments of the present application, there is also provided an electronic device including a memory and a processor, where the memory stores a computer program, and the processor is configured to execute the above-mentioned method for model training through the computer program.

[0015] According to another aspect of the embodiments of the present application, there is also provided a computer program product including computer instructions, where the computer instructions, when executed by a processor, implement the steps of the above-mentioned method for model training.

[0016] In the embodiments of the present application, by acquiring a to-be-learned task in natural language format and a task objective corresponding to the to-be-learned task, where the task objective is the result expected to be output by a large model when implementing the to-be-learned task using the large model; determining a graph task corresponding to the to-be-learned task, where the graph task is a graph structure form of the to-be-learned task, each node in the graph task represents a subtask in the to-be-learned task, and the edges in the graph task represent the dependency relationships of multiple subtasks; and training the large model using the graph task to obtain a target model, where the target model is a large model that has learned the execution strategy of the to-be-learned task, the large model is trained during the process of the large model implementing the to-be-learned task, and the understanding degree of the large model for the to-be-learned task is improved during the training process, achieving the purpose of enabling the large model to learn the execution strategy of the to-be-learned task, thereby realizing the technical effect of improving the execution efficiency of the large model when facing new tasks and enhancing the application flexibility of the large model, and further solving the technical problem of low accuracy when the large language model in the related technology executes new tasks because it cannot autonomously learn the execution strategy of new tasks. Description of the Drawings

[0017] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0018] Figure 1 is a hardware structural block diagram of a computer terminal for a method of implementing model training according to an embodiment of the present application;

[0019] Figure 2 is a flowchart of steps of a method of model training according to an embodiment of the present application;

[0020] Figure 3 is a structural diagram of a device for model training according to an embodiment of the present application. Detailed implementation manners

[0021] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0022] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above accompanying drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0023] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are explained as follows:

[0024] Large model: A deep learning model with a large number of parameters and a complex structure, such as a natural language processing model (Bidirectional Encoder Representations from Transformers, BERT).

[0025] Machine building: A building owned by a business operator or a data center service provider for installing and operating communication equipment, servers, storage devices, etc.

[0026] Machine building level: A level determined for a machine building based on multiple aspects such as reliability, redundancy, maintenance, security, etc., divided into four levels from 1 to 4, where level 4 is the highest level.

[0027] Station: A node or site in a network, usually containing equipment for data transmission, exchange, and processing.

[0028] In the related art, large models rely on static training data and can only learn the execution strategies and technical fields covered by the training data. When dealing with new tasks not involved in the training data, large models do not know the execution strategies of the new tasks. Therefore, large models have problems such as low execution efficiency and low accuracy of execution results when facing new tasks, and there are also problems with being unable to be flexibly applied to different tasks. To solve this problem, relevant solutions are provided in the embodiments of the present application, which are described in detail below.

[0029] According to the embodiments of the present application, a method embodiment for model training is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0030] The method embodiments provided by the embodiments of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal for implementing a method for model training is shown. As Figure 1 shown, the computer terminal 10 may include one or more (shown as 102a, 102b,..., 102n in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.

[0031] It should be noted that one or more of the above-mentioned processors 102 and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of the other components in the computer terminal 10. As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).

[0032] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the model training method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned model training method. The memory 104 can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 can further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.

[0033] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network can include the wireless network provided by the communication provider of the computer terminal 10. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0034] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables a user to interact with the user interface of the computer terminal 10.

[0035] The embodiments of the present application provide a model training method that can run in the above operating environment. Figure 2 is a step flowchart of the model training method provided according to the embodiments of the present application, as Figure 2 shown, the method includes the following steps:

[0036] Step S202: Obtain the to-be-learned task in natural language format and the task objective corresponding to the to-be-learned task. Here, the task objective is the result expected to be output by the large model when implementing the to-be-learned task using the large model.

[0037] The model training method provided by the embodiments of this application directly uses the to-be-learned task to train the large model: Let the large model execute the to-be-learned task, and after each execution result, give feedback information to the large model so that it can learn the optimal execution strategy of the to-be-learned task during the execution of the to-be-learned task. Using the method provided by the embodiments of this application, the large model can be used to execute various tasks, improving the application flexibility of the large model. In step S202, obtain the to-be-learned task in natural language format and the task objective in natural language format transmitted by the user terminal device. Here, the task objective is the result expected to be output by the large model when executing the to-be-learned task. For example, the to-be-learned task in natural language format is: "Help me query how many stations each building level in each district of City A has, conduct grouped statistics, and represent it using a bar chart"; then the task objective corresponding to the to-be-learned task is: statistical result and bar chart.

[0038] Step S204: Determine the graph task corresponding to the to-be-learned task. Here, the graph task is the graph structure form of the to-be-learned task, and each node in the graph task represents a subtask in the to-be-learned task, and the edges in the graph task represent the dependency relationships between multiple subtasks.

[0039] In step S204, convert the to-be-learned task in natural language format obtained in step S202 into a graph structure form to obtain the graph task corresponding to the to-be-learned task. Here, the graph structure form means forming a graph using the form of nodes and edges. Each node represents a subtask in the to-be-learned task, and each edge is used to connect two subtasks with a dependency relationship. That is to say, all the subtasks in the graph task form a to-be-learned task under the concatenation of dependency relationships. The above-mentioned edge representing the dependency relationship is a vector edge, and this vector edge can at least indicate the execution order of the subtasks. For example, when the to-be-learned task is: "Help me query how many stations each building level in each district of City A has, conduct grouped statistics, and represent it using a bar chart", the subtasks of the to-be-learned task include: query data, grouped statistics, and generate a bar chart. Among them, querying data includes: querying the building levels in each district of City A and querying the number of stations for each building level.

[0040] Optionally, determining the graph task corresponding to the task to be learned includes: parsing the task to be learned to obtain a parsing result, where the parsing result includes: an object field, an operation field, and a constraint condition field. The information recorded in the object field is used to indicate the operation object of the task to be learned, the information recorded in the operation field is used to indicate the action to be performed by the large model, and the information recorded in the constraint condition field is used to indicate the triggering condition for the large model to perform the action; determining subtasks and dependencies according to the parsing result; determining the subtasks as nodes, and determining the dependencies as edges, establishing a graph structure corresponding to the task to be learned according to the nodes and edges, and using the task corresponding to the graph structure as the graph task.

[0041] In the embodiments of the present application, the to-be-learned task in natural language format received in step S202 can be converted into a graph task in graph structure form through the method provided in this embodiment. First, the to-be-learned task in natural language format is parsed to obtain an object field, an operation field, and a constraint condition field. Among them, the information recorded in the object field is the operation object in the to-be-learned task. For example, when the to-be-learned task is "Help me query how many stations there are in each building level in each district of City A, conduct grouped statistics, and represent it using a bar chart" as described above, the information recorded in the object field is the building; the information recorded in the operation field is the action performed by the large model. For example, when the to-be-learned task is "Help me query how many stations there are in each building level in each district of City A, conduct grouped statistics, and represent it using a bar chart" as described above, the information recorded in the operation field includes the following actions: query, statistics, and drawing; the information recorded in the constraint condition field is used to indicate the trigger condition for the large model to perform the actions recorded in the operation field. For example, when the to-be-learned task is "Help me query how many stations there are in each building level in each district of City A, conduct grouped statistics, and represent it using a bar chart" as described above, the information recorded in the constraint condition is: "Each district of City A". That is, only when the location of the station belongs to a certain district of City A, the sub-task corresponding to the query action of querying the building level, the sub-task under the statistics action of counting the number of stations under each building level, and the sub-task under the drawing action of drawing a bar chart are executed. In this embodiment, in order to convert the to-be-learned task in natural language format into a graph task in graph structure form, after obtaining the parsing result of the to-be-learned task in natural language format (including: object field, operation field, and constraint condition field), continue to determine the sub-tasks and the dependency relationships between the sub-tasks according to the parsing result. For example, when the to-be-learned task is "Help me query how many stations there are in each building level in each district of City A, conduct grouped statistics, and represent it using a bar chart" as described above, the sub-tasks include: "Query the building information in City A (i.e., the above query data)", "Group and count the data by building level (i.e., the above grouped statistics)", "Display the statistical result using a bar chart (i.e., the above generate a bar chart)". The dependency relationships between the sub-tasks include the following: The drawing task (i.e., the above display the statistical result using a bar chart) depends on the result of the statistical task (i.e., the above group and count the data by building level), and the statistical task depends on the completion of the query task (i.e., the above query the building information in City A). Finally, based on the sub-tasks and dependency relationships determined according to the above parsing result, a graph task is constructed so that the model can understand the task process.

[0042] According to some optional embodiments of the present application, sub-tasks and dependency relationships are determined according to the parsing result. Among them, determining sub-tasks includes: determining the action corresponding to each operation object according to the object field and the operation field; determining sub-tasks according to each operation object and its corresponding action.

[0043] As mentioned in the previous embodiment, in the method provided in the embodiment of the present application, the to-be-learned task in natural language format is converted into a graph task in graph structure form. After obtaining the parsing result of the to-be-learned task in natural language format (including: object field, operation field, and constraint condition field), multiple subtasks decomposed from the to-be-learned task are determined. In this embodiment, determining the subtasks mainly depends on the object field and the operation field in the parsing result. Specifically, multiple operation objects are determined according to the object field, and the action corresponding to each operation object is determined according to the operation field. Each operation object and its corresponding operation action are determined as a subtask. For example, if the association between the "query" action and the "building information" object is recognized, then "query building information" is a subtask, corresponding to a node. Additionally, in this embodiment, in the case of multiple operation objects, all subtasks corresponding to each operation object can also be determined as a subtask. Through the above method, specific subtasks are determined according to each operation object and its corresponding execution action, which helps the large model clarify each specific task it needs to complete, such as "query the grade information of all buildings in City A".

[0044] According to some other optional embodiments of the present application, subtasks and dependency relationships are determined according to the parsing result. Among them, determining the dependency relationship includes: for each subtask, determining the value of the tool call field corresponding to the subtask, where the tool call field is included in the operation field, and the information recorded in the tool call field is used to instruct the large model to call external resources, and the external resources are used to implement functions that the large model cannot implement; in the case where the value of the tool call field is a valid value, it is determined that there is a dependency relationship between the first type of node and the second type of node, where the first type of node is the subtask whose value of the included tool call field is a valid value, and the second type of node is the subtask including the operation of calling external resources; in the case where the value of the tool call field is an invalid value, it is determined that there is no dependency relationship between the second type of node and the third type of node, where the third type of node is the subtask whose value of the included tool call field is an invalid value.

[0045] As mentioned in the above embodiments, in this embodiment, the to-be-learned task in natural language format is converted into a graph task in graph structure form. After obtaining the parsing results of the to-be-learned task in natural language format (including: object field, operation field, and constraint condition field), the dependency relationship between the sub-tasks decomposed from the to-be-learned task is determined. Specifically, in this embodiment, the dependency relationship between the sub-tasks can be determined according to the tool call field included in the operation field. Specifically, analyze the tool call field in each sub-task to determine whether external resources need to be called. The external resources can be in various forms such as software and hardware devices, and the external resources can implement functions that the large model itself cannot implement. For example, in this embodiment, when the to-be-learned task is "Help me query how many stations there are in each district of City A for the building level, perform grouped statistics, and represent it with a bar chart", the external resources include: a database storing the building information of City A, a drawing tool, etc. If the value of the tool call field is a valid value (such as 1), it indicates that the sub-task needs to call a specific tool or resource. For example, data query needs to call a database. At this time, create an edge representing the association relationship between the node corresponding to the sub-task of "querying the building information of City A (i.e., the above query data)" (i.e., the first type of node) and the node corresponding to the sub-task of "calling the database" (i.e., the second type of node). If the value of the tool call field is an invalid value (such as 0), it indicates that the sub-task does not need to call a specific tool or resource. Then, it is determined that there is no dependency relationship between the node corresponding to the sub-task that does not need to call external resources (i.e., the third type of node) and the node including the operation of calling external resources (i.e., the second type of node), and there is no need to create an edge representing the dependency relationship between the second type of node and the third type of node. In addition, in this embodiment, the dependency relationship can also be determined according to the constraint conditions. The two sub-tasks associated with the constraint conditions are determined as sub-tasks with a dependency relationship, and an edge for representing the dependency relationship is created between the two nodes corresponding to these two sub-tasks.

[0046] Step S206: Train the large model using the graph task to obtain a target model, where the target model is a large model that has learned the execution strategy of the to-be-learned task.

[0047] In step S206, the graph task obtained in step S204 is used to train the large model. The large model learns how to optimize the decision-making process through multiple attempts in the environment to achieve the above to-be-learned task and output the task objective. That is, during the training process, multiple sub-tasks in the graph task are executed according to the dependency relationship in the graph task. After each execution is completed, the strategy for the large model to execute the to-be-learned task is adjusted according to the result of the large model's execution until the result output by the large model is the same as the task objective in step S202.

[0048] In step S206, the large model can be loaded into the memory. For example, the original data of the large model can be loaded from the non-volatile memory to the volatile memory so that the processor can run the large model. The original data of the large model refers to the unprocessed data, which usually includes the parameters and structure data of the large model. The structure data can be the calculation relationship based on the parameters, such as the forward propagation calculation relationship between intermediate layers and between neurons. Specifically, the structure data can include the code related to the structure of the large model, such as the code used to perform the calculations related to the intermediate layers and between neurons.

[0049] In one implementation, an area for loading the large model can be partitioned in the memory, which can include a structure data storage area and a parameter storage area. The structure data storage area is used to store the structure-related code, and the parameters it references can point to the addresses of specific parameters in the parameter storage area through pointers. During the training process of the large model, the parameters may need to be updated frequently, so only the parameter values in the parameter storage area need to be updated.

[0050] Optionally, the large model is trained by the following method: determining the actions performed by the large model in the dataset, where the dataset stores multiple sets of historical interaction data corresponding to multiple historical moments. Each set of historical interaction data includes the historical state of the large model at the historical moment, the historical actions performed by the large model at the historical moment, the rewards obtained by the large model after performing the historical actions, and the next state of the large model after obtaining the rewards. The rewards are used to indicate the contribution of the historical actions performed by the large model to the completion of the historical task goal; determining the function value of the objective function according to the actions performed by the large model and the current state, where the current state is used to indicate the progress of the large model in completing the graph task, and the objective function is used to evaluate the contribution of the actions performed by the large model to the completion of the task goal; determining that the training is completed when the number of consecutive times that the change amount of the function value is less than the preset change amount is greater than or equal to the preset number, where the change amount is the difference between the previous function value and the next function value, and the previous function value and the next function value are the function values of the objective function determined in two adjacent iterative training processes.

[0051] The method provided by the embodiments of this application uses a training method of reinforcement learning to train a large model. During the training process, the large model executes a graph task by selecting actions to be executed from the historical interaction dataset: the large model selects actions to be executed from the historical interaction dataset and, when implementing the graph task, executes the selected actions. The above-mentioned dataset stores multiple sets of historical interaction data generated when the large model executes actions at multiple historical moments (moments before the current moment); among them, each data point at a historical moment (i.e., each set of historical interaction data) represents the behavior and results of the large model in a specific task environment, including: the state of the large model at a historical moment, the executed action, the obtained reward, and the next state after the action is executed; in this embodiment, the reward obtained by the large model refers to the change in the progress of completing the historical task objective (i.e., the contribution degree) after the large model executes the corresponding action at this historical moment. This change is a vector value. When the contribution degree is greater than 0, it indicates that the progress of completing the historical task objective is increasing. When the contribution degree is less than 0, it indicates that the progress of completing the historical task objective is regressing. The historical task objective is the expected result of the historical task executed by the large model at the historical moment. In this embodiment, when using the reinforcement learning method for training, the impact of the large model's execution of this action on completing the task to be learned is evaluated according to a predefined evaluation function (i.e., the objective function); during the training process, the specific actions executed by the large model when achieving each subtask are determined according to the change in the objective function. When the change in the objective function is less than the preset change for N consecutive times (N is a preset number), it is determined that the large model has learned the execution strategy of the task to be learned, and the training of the large model is completed. The change in the above-mentioned objective function refers to the difference between the function values corresponding to two adjacent iterative trainings of the objective function (i.e., the previous function value and the next function value). The function value corresponding to each iterative training can be determined using the following formula representing the objective function: Q(S t ,a t )←Q(S t ,a t )+α[r t +γ·max a′ Q(S t+1 ,a′)-Q(S t ,a t )], where Q(S t ,a t ) is the function value corresponding to state S t , S t represents the state of the large model at time (t) (that is, the progress of the large model in completing the graph task), a t represents the action executed by the large model determined from the dataset, α represents the learning rate, which is a preset value, r t is the reward obtained by the large model for executing action a tThe reward obtained later, γ is a value belonging to the interval [0, 1], S t+1 is the state where the large model executes action a t (i.e., the next state), and a′ is any action stored in the dataset. According to the evaluation function provided in the embodiments of the present application, it is possible to evaluate whether the current action helps to achieve the task goal more efficiently, so that the large model adopts the optimal execution strategy to implement the task to be learned.

[0052] According to some optional embodiments of the present application, the multi-group historical interaction data includes: real interaction data generated when the large model executes historical tasks, and test interaction data generated when the large model executes actions in the test set, where the test set contains multiple preset actions.

[0053] In the embodiments of the present application, in order to more comprehensively enhance the autonomy and task execution ability of the large model, a training strategy that combines real interaction data and test interaction data is adopted, that is, the multi-group historical interaction data stored in the dataset includes: interaction data generated when the large model executes the task to be learned in step S202 at a historical moment (i.e., real interaction data generated when executing historical tasks) and test interaction data generated when the large model executes actions in the test set. Among them, the real interaction data includes: model states (including the state before and after executing the action), input information in natural language format, the actions executed by the model, and environmental feedback (rewards); the test set only contains data representing actions (i.e., preset actions). During the training process, the large model executes each action in the test set and records the states before and after each action execution and the obtained rewards.

[0054] According to some other optional embodiments of the present application, determining the action executed by the large model in the dataset includes: determining a group of historical interaction data corresponding to the largest reward value as the target dataset; and determining the action included in the target dataset as the action executed by the large model.

[0055] In some embodiments, when the large model determines the action to be executed in the dataset, the reward obtained after executing the action is used as a reference standard. Since each action in the dataset is associated with a reward, and this reward is used to indicate the contribution of the large model's execution of this action to the completion of the task to be learned, then, in order to improve the efficiency of executing the task to be learned, in this embodiment, the reward values of each action in the historical interaction data (i.e., the contribution corresponding to the reward) are analyzed, and a group of interaction data with the largest reward value is determined as the target dataset. Extract the actions executed by the model from the target dataset as the learning target during training to promote the model to learn the most efficient execution strategy.

[0056] Through the above steps, by transforming natural language tasks into graph-structured tasks, the internal logic of the tasks can be understood and processed more intuitively, especially the dependency relationships among subtasks in the tasks. This helps the large model learn the execution strategies of the tasks more effectively during the training process. By parsing natural language tasks and extracting key fields, the subtasks and dependency relationships can be defined more precisely, thereby improving the efficiency and accuracy of model training. In addition, by introducing the invocation of external resources during the training process, the execution ability of the model can be enhanced, enabling it to handle more complex and extensive task types. It can achieve improving the execution efficiency of the large model when processing new tasks, improving the accuracy of the output results of the large model when executing new tasks, and improving the application flexibility of the large model.

[0057] Figure 3 is a structural diagram of a device for model training provided according to an embodiment of the present application, as Figure 3 shown, the device for model training includes: an acquisition module 30, configured to acquire a to-be-learned task in natural language format and a task objective corresponding to the to-be-learned task, where the task objective is the result expected to be output by the large model when implementing the to-be-learned task using the large model; a determination module 32, configured to determine a graph task corresponding to the to-be-learned task, where the graph task is a graph structure form of the to-be-learned task, each node in the graph task represents a subtask in the to-be-learned task, and the edges in the graph task represent the dependency relationships of multiple subtasks; a training module 34, configured to train the large model using the graph task to obtain a target model, where the target model is a large model that has learned the execution strategy of the to-be-learned task.

[0058] When training a large model using the training device of the above model, the acquisition module 30 receives the to-be-learned task in natural language format and the task objective in natural language format input by the user terminal. Among them, the task objective is the result expected to be output by the large model when the large model executes the to-be-learned task. When training the large model to learn the strategy of the to-be-learned task, the acquisition module 30 sends the received to-be-learned task and its corresponding task objective to the determination module 32. The determination module 32 converts the to-be-learned task in natural language format into a graph structure form to obtain the graph task corresponding to the to-be-learned task. Specifically, based on the parsed natural language task, a graph containing nodes and edges is created, where each node represents a subtask and the edge represents the dependency relationship between subtasks, so that the large model can understand the multiple subtasks obtained by decomposing the to-be-learned task and the dependency relationship between the multiple subtasks, ensuring that the large model executes the to-be-learned task according to the correct logic. Further, the determination module 32 sends the graph task to the training module 34, and the training module 34 trains the large model using the graph task and the method of reinforcement learning, so that it can quickly learn the execution strategy of the new task (i.e., the to-be-learned task). When training the large model using the reinforcement learning strategy, the large model executes the actions (subtasks) in the graph task, observes the environmental feedback (reward), and adjusts the executed actions according to the feedback; in multiple rounds of iteration, the model continuously executes the graph task, learns how to optimize the execution order and manner of subtasks, and how to handle the dependency relationship between subtasks, and finally obtains the target model that can efficiently execute the to-be-learned task.

[0059] It should be noted that Figure 3 For the preferred implementation manners of the illustrated embodiments, reference may be made to Figure 2 the relevant descriptions of the illustrated embodiments, which will not be elaborated here.

[0060] The embodiments of the present application further provide a non-volatile storage medium. A computer program is stored in the non-volatile storage medium. Among them, on the device where the non-volatile storage medium is located, the above model training method is executed by running the computer program.

[0061] The above non-volatile storage medium is used to store a program for executing the following functions: acquiring a to-be-learned task in natural language format and a task objective corresponding to the to-be-learned task, where the task objective is the result expected to be output by the large model when the large model implements the to-be-learned task; determining a graph task corresponding to the to-be-learned task, where the graph task is the graph structure form of the to-be-learned task, each node in the graph task represents a subtask in the to-be-learned task, and the edge in the graph task represents the dependency relationship of multiple subtasks; training the large model using the graph task to obtain a target model, where the target model is a large model that has learned the execution strategy of the to-be-learned task.

[0062] The embodiments of the present application further provide an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to execute the above method for model training through the computer program.

[0063] The processor in the above electronic device is used to run a program that performs the following functions: obtaining a to-be-learned task in natural language format and a task objective corresponding to the to-be-learned task, where the task objective is the result expected to be output by a large model when implementing the to-be-learned task using the large model; determining a graph task corresponding to the to-be-learned task, where the graph task is the graph structure form of the to-be-learned task, each node in the graph task represents a subtask in the to-be-learned task, and the edges in the graph task represent the dependency relationships of multiple subtasks; training the large model using the graph task to obtain a target model, where the target model is a large model that has learned the execution strategy of the to-be-learned task.

[0064] The embodiments of the present application further provide a computer program product, including computer instructions, which implement the steps of the above method for model training when executed by a processor.

[0065] It should be noted that each module in the above device for model training can be a program module (for example, a set of program instructions that implement a specific function), or a hardware module. For the latter, it can be presented in the following forms, but not limited to this: the manifestation form of each of the above modules is a processor, or the functions of each of the above modules are implemented by a processor.

[0066] The serial numbers of the above embodiments of the present application are only for description and do not represent the superiority or inferiority of the embodiments.

[0067] In the above embodiments of the present application, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0068] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.

[0069] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0070] In addition, each functional unit in various embodiments of the present application may be integrated into a processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0071] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the related technology, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs and other various media that can store program codes.

[0072] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for model training, characterized in that, Including: Obtain a to-be-learned task in natural language format and a task objective corresponding to the to-be-learned task, where the task objective is the result expected to be output by the large model when implementing the to-be-learned task using the large model; Determine a graph task corresponding to the to-be-learned task, where the graph task is the graph structure form of the to-be-learned task, each node in the graph task represents a subtask in the to-be-learned task, and the edges in the graph task represent the dependency relationships of multiple subtasks; Train the large model using the graph task to obtain a target model, where the target model is the large model that has learned the execution strategy of the to-be-learned task.

2. The method according to claim 1, wherein Determining the graph task corresponding to the to-be-learned task includes: Parse the to-be-learned task to obtain a parsing result, where the parsing result includes: an object field, an operation field, and a constraint condition field. The information recorded in the object field is used to indicate the operation object of the to-be-learned task, the information recorded in the operation field is used to indicate the action to be performed by the large model, and the information recorded in the constraint condition field is used to indicate the trigger condition for the large model to perform the action; Determine the subtasks and the dependency relationships according to the parsing result; Determine the subtasks as the nodes, determine the dependency relationships as the edges, establish a graph structure corresponding to the to-be-learned task according to the nodes and the edges, and use the task corresponding to the graph structure as the graph task.

3. The method according to claim 2, wherein Determine the subtasks and the dependency relationships according to the parsing result, where determining the subtasks includes: Determine the action corresponding to each operation object according to the object field and the operation field; Determine the subtasks according to each operation object and its corresponding action.

4. The method according to claim 2, wherein Determine the subtasks and the dependency relationships according to the parsing result, where determining the dependency relationships includes: For each subtask, determine the value of the tool call field corresponding to the subtask, where the tool call field is included in the operation field, and the information recorded in the tool call field is used to indicate that the large model calls an external resource, and the external resource is used to implement a function that the large model cannot implement; When the value of the tool call field is a valid value, determine that there is a dependency relationship between the first type of node and the second type of node, where the first type of node is a subtask whose value of the included tool call field is the valid value, and the second type of node is a subtask including an operation of calling the external resource; When the value of the tool call field is an invalid value, determine that there is no dependency relationship between the second type of node and the third type of node, where the third type of node is a subtask whose value of the included tool call field is the invalid value.

5. The method according to claim 1, wherein The large model is trained by the following method: Determine the actions performed by the large model in the dataset, where the dataset stores multiple sets of historical interaction data. The multiple sets of historical interaction data correspond to multiple historical moments. Each set of historical interaction data includes the historical state of the large model at the historical moment, the historical actions performed by the large model at the historical moment, the rewards obtained by the large model after performing the historical actions, and the next state of the large model after obtaining the rewards. The rewards are used to indicate the contribution degree of the historical actions performed by the large model to the completion of the historical task goal; Determine the function value of the objective function according to the actions performed by the large model and the current state, where the current state is used to indicate the progress of the large model in completing the graph task, and the objective function is used to evaluate the contribution degree of the actions performed by the large model to the completion of the task goal; When the number of consecutive times that the change amount of the function value is less than the preset change amount is greater than or equal to the preset number of times, determine that the training is completed, where the change amount is the difference between the previous function value and the next function value, and the previous function value and the next function value are the function values of the objective function determined in two adjacent iterative training processes.

6. The method according to claim 5, characterized in that The multiple sets of historical interaction data include: the real interaction data generated when the large model performs historical tasks, and the test interaction data generated when the large model performs the actions in the test set, where the test set contains multiple preset actions.

7. The method according to claim 5, wherein Determining the actions performed by the large model in the dataset includes: determining a set of historical interaction data corresponding to the largest reward value as the target dataset; determining the actions included in the target dataset as the actions performed by the large model.

8. An apparatus for model training, characterized in that, Include: An acquisition module for acquiring a to-be-learned task in natural language format and a task goal corresponding to the to-be-learned task, where the task goal is the result expected to be output by the large model when implementing the to-be-learned task using the large model; A determination module for determining the graph task corresponding to the to-be-learned task, where the graph task is the graph structure form of the to-be-learned task, each node in the graph task represents a subtask in the to-be-learned task, and the edges in the graph task represent the dependency relationships of multiple subtasks; A training module for training the large model using the graph task to obtain a target model, where the target model is the large model that has learned the execution strategy of the to-be-learned task.

9. A non-volatile storage medium, characterized in that, A computer program is stored in the non-volatile storage medium, where the method for training the model according to any one of claims 1 to 7 is executed by the device where the non-volatile storage medium is located by running the computer program.

10. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method for training the model according to any one of claims 1 to 7 through the computer program.

11. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, the steps of the method for training the model according to any one of claims 1 to 7 are implemented.