Model output control method based on reinforcement learning
By presetting the model output data and using the output control model and reinforcement learning model for differential judgment and training, the problem of unreliable model output is solved, and highly customized model output control is realized, which improves the reliability and efficiency of model output.
Patent Information
- Application Number
- CN202510615568.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-12
AI Technical Summary
The existing model output control methods are difficult to achieve customization of model output, resulting in unreliable model output.
By presetting the output data of the model, collecting real-time output data and making differential judgments, using the output control model and reinforcement learning model for training and adjustment, the model output is highly customized.
It improves the reliability of model output control, avoids harmful guidance, saves training time, improves computing resource utilization and model training efficiency, and enhances the adaptability and accuracy of the model.
Smart Images

Figure CN120471132A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of model output control, and in particular to a model output control method based on reinforcement learning. Background Art
[0002] By controlling the model output through reinforcement learning technology, the legality and compliance of the model output content can be effectively achieved. However, the existing relevant technical solutions do not control the model output through pre-set models, making it difficult to achieve the customization of the model output. Summary of the Invention
[0003] This application provides a model output control method based on reinforcement learning, which aims to improve existing related technical problems.
[0004] An embodiment of the present application provides a model output control method based on reinforcement learning, which may include the following steps:
[0005] Presetting the output data of the model and collecting the real-time output data of the model, the output control model determines whether there is a difference between the output data of the model and the output data of the preset model based on the collected real-time output data of the model, and when there is no difference between the real-time output data of the model and the output data of the preset model, putting the model into use;
[0006] When the output control model determines, based on the real-time output data of the model collected, that the output data of the model differs from the output data of a preset model, learning the output data of the preset model through a reinforcement learning model, and training the model based on the output data of the preset model;
[0007] Collecting real-time output data of the trained model, and determining whether there is a difference between the real-time output data of the trained model and the output data of a preset model through an output control model;
[0008] When there is no difference between the real-time output data of the trained model and the output data of the preset model, the trained model is put into use.
[0009] In the above technical solution, by pre-setting the output data of the model, the output control model controlling the model output, and the reinforcement learning model strengthening the model output, the model output is highly customizable, thereby improving the reliability of the model output control and effectively avoiding harmful guidance of the model output.
[0010] In a preferred example, the solution of the first aspect of the present application can be further configured as follows:
[0011] The model output control method based on reinforcement learning may further include the following steps:
[0012] When the output control model determines that there is a difference between the output data of the model and the output data of the preset model based on the real-time output data of the collected model, the model parameters of the model are adjusted according to the difference value between the real-time output data of the collected model and the output data of the preset model.
[0013] In the above technical solution, by appropriately adjusting the model parameters of the model, the training time of the model is saved and the training efficiency of the model training is improved.
[0014] In a preferred example, the solution of the first aspect of the present application can be further configured as follows:
[0015] The model output control method based on reinforcement learning may further include the following steps:
[0016] If the difference between the real-time output data of the collected model and the output data of the preset model exceeds a preset threshold, the model is trained in distributed parallel based on the output data of the preset model.
[0017] In the above technical solution, when the difference value exceeds the preset threshold, the model is trained in distributed parallel, which improves the utilization of computing resources and meets the computing power requirements of the model in the low computing power case of multi-card and multi-node clustering.
[0018] In a preferred example, the solution of the first aspect of the present application can be further configured as follows:
[0019] The model output control method based on reinforcement learning may further include the following steps:
[0020] The preset threshold is adaptively generated by reinforcing the learning model's learning of the output data of the preset model and the collected real-time output data of the model.
[0021] Through the above technical solutions, the adaptability of large-scale model deployment of reinforcement learning models is improved.
[0022] In a preferred example, the solution of the first aspect of the present application can be further configured as follows:
[0023] The model output control method based on reinforcement learning may further include the following steps:
[0024] The real-time output data of the model put into application is continuously detected to see whether abnormal output data occurs. When abnormal output data is detected in the real-time output data of the model put into application, the model is trained based on the output data of the preset model.
[0025] Through the above technical solution, the output control of the model is strengthened.
[0026] In a preferred example, the solution of the first aspect of the present application can be further configured as follows:
[0027] In the step of presetting the output data of the model and collecting the real-time output data of the model, the output control model determines whether there is a difference between the output data of the model and the output data of the pre-set model based on the collected real-time output data of the model, and when there is no difference between the real-time output data of the model and the output data of the pre-set model, the model is put into application. The method for determining whether there is a difference between the real-time output data of the collected model and the output data of the pre-set model includes the following expression:
[0028]
[0029] Where ε is a preset constant, C n Indicates the nth real-time output data of the collected model, μ n It represents the entropy value of the nth real-time output data of the collected model, N represents the total number of real-time output data of the collected model, 1≤n≤N, n is a positive integer, Y m represents the mth output data of the preset model, |·| represents the absolute value, τ m It represents the entropy value of the mth output data of the preset model, M represents the total number of output data of the preset model, 1≤m≤M, and m is a positive integer.
[0030] Through the above technical solution, it is possible to accurately determine whether there is a difference between the real-time output data of the collected model and the output data of the pre-set model. When the above formula is satisfied, it is determined that there is a difference between the real-time output data of the collected model and the output data of the pre-set model.
[0031] In a preferred example, the solution of the first aspect of the present application can be further configured as follows:
[0032] The model output control method based on reinforcement learning may further include the following steps:
[0033] If the output data of the model needs to be changed, the output control model is trained based on the changed output data, and the output data of the model is controlled by the trained output control model.
[0034] In the above technical solution, when the user's demand for the output data of the model changes, the output data of the control model can be adjusted in time, thereby being able to flexibly meet the user's demand.
[0035] In a preferred example, the solution of the first aspect of the present application can be further configured as follows:
[0036] The model output control method based on reinforcement learning may further include the following steps:
[0037] The reinforcement learning model performs deep reinforcement learning on the real-time output data of the collected model and the output data of the preset model, outputs control data of the model output, and inputs the control data into the output control model.
[0038] In the above technical solution, the accuracy of the output control model for model output control is improved through deep reinforcement learning of the reinforcement learning model.
[0039] In a preferred example, the solution of the first aspect of the present application can be further configured as follows:
[0040] The model output control method based on reinforcement learning may further include the following steps:
[0041] The control instruction data output by the output control model and the model output control data output by the reinforcement learning model are integrated to generate final model output control data, and the output of the model is controlled by the final model output control data.
[0042] Through the above technical solution, the output control of the model is further strengthened.
[0043] In a preferred example, the solution of the first aspect of the present application can be further configured as follows:
[0044] The model output control method based on reinforcement learning may further include the following steps:
[0045] The reinforcement learning model performs deep reinforcement learning on the detected abnormal output data and the output data of the trained model.
[0046] Through the above technical solutions, the reinforcement learning model can solve the challenges brought by the complexity of data and improve the generalization ability of the reinforcement learning model.
[0047] Based on the above method embodiment, the present application provides a corresponding terminal embodiment;
[0048] The present application provides a terminal, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements a model output control method based on reinforcement learning as described in any embodiment of the present application.
[0049] Based on the above method embodiment, the present application provides a storage medium embodiment;
[0050] The present application provides a storage medium, including a processor, a memory, and a computer program stored in the above-mentioned memory and configured to be executed by the above-mentioned processor. When the above-mentioned processor executes the above-mentioned computer program, a model output control method based on reinforcement learning described in any embodiment of the present application is implemented.
[0051] This application has at least the following beneficial effects:
[0052] The present application provides a reinforcement learning-based model output control method, which realizes high customizability of model output, thereby improving the reliability of model output control and effectively avoiding harmful guidance of model output. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 This is a flow chart of a model output control method based on reinforcement learning in one embodiment of the present application.
[0054] Figure 2 This is a structural block diagram of a dynamic construction system for a digital power grid knowledge base according to an embodiment of the present application. DETAILED DESCRIPTION
[0055] The following will be combined with the accompanying drawings to clearly, completely, and comprehensively describe the technical solutions in this application. Obviously, the embodiments described are only some of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0056] like Figure 1 As shown, an embodiment of the present application provides a model output control method based on reinforcement learning, which may specifically include the following steps:
[0057] Step S1: presetting the output data of the model and collecting the real-time output data of the model; the output control model determines whether there is a difference between the output data of the model and the output data of the pre-set model based on the collected real-time output data of the model; when there is no difference between the real-time output data of the model and the output data of the pre-set model, the model is put into use;
[0058] Step S2: when the output control model determines, based on the real-time output data of the collected model, that the output data of the model differs from the output data of the preset model, learning the output data of the preset model through the reinforcement learning model, and training the model based on the output data of the preset model;
[0059] Step S3: collecting the real-time output data of the trained model, and determining whether there is a difference between the real-time output data of the trained model and the output data of the preset model through the output control model;
[0060] Step S4: When there is no difference between the real-time output data of the trained model and the output data of the preset model, the trained model is put into use.
[0061] In the above embodiment, by pre-setting the output data of the model, the output control model's control of the model output, and the reinforcement learning model's reinforcement learning of the model output, the model output is highly customizable, thereby improving the reliability of the model output control and effectively avoiding harmful guidance of the model output.
[0062] In a preferred embodiment, in order to save model training time and improve model training efficiency by appropriately adjusting the model parameters of the model, the model output control method based on reinforcement learning may further include the following steps:
[0063] When the output control model determines that there is a difference between the output data of the model and the output data of the preset model based on the real-time output data of the collected model, the model parameters of the model are adjusted according to the difference value between the real-time output data of the collected model and the output data of the preset model.
[0064] In a preferred embodiment, in order to implement distributed parallel training of the model when the difference value exceeds a preset threshold, thereby improving the utilization of computing resources and meeting the computing power requirements of the model in the case of low computing power in a multi-card and multi-node cluster, the model output control method based on reinforcement learning may further include the following steps:
[0065] If the difference between the real-time output data of the collected model and the output data of the preset model exceeds a preset threshold, the model is trained in distributed parallel based on the output data of the preset model.
[0066] In a preferred embodiment, in order to improve the adaptability of large-scale model deployment of the reinforcement learning model, the model output control method based on reinforcement learning may further include the following steps:
[0067] The preset threshold is adaptively generated by reinforcing the learning model's learning of the output data of the preset model and the collected real-time output data of the model.
[0068] In a preferred embodiment, in order to strengthen the output control of the model, the model output control method based on reinforcement learning may further include the following steps:
[0069] The real-time output data of the model put into application is continuously detected to see whether abnormal output data occurs. When abnormal output data is detected in the real-time output data of the model put into application, the model is trained based on the output data of the preset model.
[0070] In a preferred embodiment, when the following formula is satisfied, it is determined that there is a difference between the real-time output data of the collected model and the output data of the preset model. In order to accurately determine whether there is a difference between the real-time output data of the collected model and the output data of the preset model, the output data of the preset model is collected and the real-time output data of the model is collected. The output control model determines whether there is a difference between the output data of the model and the output data of the preset model based on the real-time output data of the collected model. When there is no difference between the real-time output data of the model and the output data of the preset model, the model is put into application. The method for determining whether there is a difference between the real-time output data of the collected model and the output data of the preset model includes the following expression:
[0071]
[0072] Where ε is a preset constant, C n Indicates the nth real-time output data of the collected model, μ n It represents the entropy value of the nth real-time output data of the collected model, N represents the total number of real-time output data of the collected model, 1≤n≤N, n is a positive integer, Y m represents the mth output data of the preset model, |·| represents the absolute value, τ m represents the mth output data entropy value of the preset model, M represents the total number of output data of the preset model, 1≤m≤M, m is a positive integer, sin(·) represents the sine function, represents the limit as N approaches ∞, represents the limit as M approaches ∞.
[0073] In a preferred embodiment, in order to timely adjust the output data of the control model when the user's demand for the output data of the model changes, thereby flexibly meeting the user's demand, the model output control method based on reinforcement learning may further include the following steps:
[0074] If the output data of the model needs to be changed, the output control model is trained based on the changed output data, and the output data of the model is controlled by the trained output control model.
[0075] In a preferred embodiment, in order to improve the accuracy of the output control model for model output control through deep reinforcement learning of the reinforcement learning model, the model output control method based on reinforcement learning may further include the following steps:
[0076] The reinforcement learning model performs deep reinforcement learning on the real-time output data of the collected model and the output data of the preset model, outputs control data of the model output, and inputs the control data into the output control model.
[0077] In a preferred embodiment, in order to further strengthen the output control of the model, the model output control method based on reinforcement learning may further include the following steps:
[0078] The control instruction data output by the output control model and the model output control data output by the reinforcement learning model are integrated to generate final model output control data, and the output of the model is controlled by the final model output control data.
[0079] In a preferred embodiment, in order to enable the reinforcement learning model to address the challenges brought by data complexity and improve the generalization ability of the reinforcement learning model, the model output control method based on reinforcement learning may further include the following steps:
[0080] The reinforcement learning model performs deep reinforcement learning on the detected abnormal output data and the output data of the trained model.
[0081] An embodiment of the present application provides a dynamic construction system for a digital power grid knowledge base, such as Figure 2 Specifically, it may include:
[0082] A model data difference judgment module is used to pre-set the output data of the model and collect the real-time output data of the model. The output control model judges whether there is a difference between the output data of the model and the output data of the pre-set model based on the collected real-time output data of the model. When there is no difference between the real-time output data of the model and the output data of the pre-set model, the model is put into use;
[0083] a model training module for learning the output data of the preset model through a reinforcement learning model when the output data of the model is determined to be different from the output data of a preset model based on the real-time output data of the model collected by the output control model, and training the model based on the output data of the preset model;
[0084] A post-training model data difference judgment module collects the real-time output data of the trained model and judges whether there is a difference between the real-time output data of the trained model and the output data of the preset model through the output control model;
[0085] The model application module is used to apply the trained model when there is no difference between the real-time output data of the trained model and the output data of the preset model.
[0086] It should be noted that the system embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network elements. Some or all of the modules may be selected to achieve the objectives of the present embodiment according to actual needs. Furthermore, in the system embodiment drawings provided herein, the connection relationship between modules indicates that they have a communication connection, which may be implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement the present invention without inventive effort. The above schematic diagram is merely an example of a system for dynamically constructing a digital power grid knowledge base and does not constitute a limitation on the system for dynamically constructing a digital power grid knowledge base. The present invention may include more or fewer components than shown, or combine certain components, or have different components.
[0087] Based on the above method embodiment, the present application provides a corresponding terminal embodiment.
[0088] Another embodiment of the present application provides a terminal, including a processor, a memory, and a computer program stored in the above memory and configured to be executed by the above processor. When the above processor executes the above computer program, a model output control method based on reinforcement learning as described in any embodiment of the present application is implemented.
[0089] For example, in this embodiment, the computer program may be divided into one or more modules, which are stored in the memory and executed by the processor to implement the present application. The one or more module elements may be a series of computer program instruction segments capable of performing specific functions, which are used to describe the execution process of the computer program in the device.
[0090] The terminals can be computing devices such as desktop computers, notebook computers, PDAs, and cloud servers. The devices can include, but are not limited to, processors and memories.
[0091] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor. The processor is the control center of the device, connecting various parts of the device using various interfaces and lines.
[0092] The above-mentioned memory can be used to store the above-mentioned computer programs and / or modules. The above-mentioned processor realizes various functions of the above-mentioned device by running or executing the computer programs and / or modules stored in the above-mentioned memory, and calling the data stored in the memory. The above-mentioned memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; in addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0093] Based on the above method embodiment, the present application provides a corresponding storage medium embodiment.
[0094] Another embodiment of the present application provides a storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute a model output control method based on reinforcement learning as described in any embodiment of the present application.
[0095] In this embodiment, the storage medium is a computer-readable storage medium, and the computer program includes computer program code, which may be in source code form, object code form, an executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium.
[0096] In the above-mentioned embodiment of the present application, the enterprise's intranet and intranet are integrated, so that internal staff participating in the enterprise's intranet project can obtain external network information related to the enterprise's intranet project by simply logging into the enterprise's intranet, and can obtain relevant information about the enterprise's intranet project very conveniently; by setting up external network information access accounts for internal staff participating in the enterprise's intranet project to access the project-related external network information, and setting different external network information access permissions for different external network information access accounts, internal staff participating in the enterprise's intranet project can conveniently and accurately obtain relevant external network information of the enterprise's intranet project in which they participate.
[0097] The above is a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications are also considered to be within the scope of protection of the present application.
Claims
1. A model output control method based on reinforcement learning, characterized in that: The following steps are involved: Presetting the output data of the model and collecting the real-time output data of the model, the output control model determines whether there is a difference between the output data of the model and the output data of the preset model based on the collected real-time output data of the model, and when there is no difference between the real-time output data of the model and the output data of the preset model, putting the model into use; When the output control model determines, based on the real-time output data of the model collected, that the output data of the model differs from the output data of a preset model, learning the output data of the preset model through a reinforcement learning model, and training the model based on the output data of the preset model; Collecting real-time output data of the trained model, and determining whether there is a difference between the real-time output data of the trained model and the output data of a preset model through an output control model; When there is no difference between the real-time output data of the trained model and the output data of the preset model, the trained model is put into use.
2. The model output control method based on reinforcement learning according to claim 1, characterized in that: It also includes the following steps: When the output control model determines that there is a difference between the output data of the model and the output data of the preset model based on the real-time output data of the collected model, the model parameters of the model are adjusted according to the difference value between the real-time output data of the collected model and the output data of the preset model.
3. The model output control method based on reinforcement learning according to claim 2, characterized in that: It also includes the following steps: If the difference between the real-time output data of the collected model and the output data of the preset model exceeds a preset threshold, the model is trained in distributed parallel based on the output data of the preset model.
4. The model output control method based on reinforcement learning according to claim 3, characterized in that: It also includes the following steps: The preset threshold is adaptively generated by reinforcing the learning model's learning of the output data of the preset model and the collected real-time output data of the model.
5. The model output control method based on reinforcement learning according to claim 1, characterized in that: It also includes the following steps: The real-time output data of the model put into application is continuously detected to see whether abnormal output data occurs. When abnormal output data is detected in the real-time output data of the model put into application, the model is trained based on the output data of the preset model.
6. The model output control method based on reinforcement learning according to claim 1, characterized in that: In the step of presetting the output data of the model and collecting the real-time output data of the model, the output control model determines whether there is a difference between the output data of the model and the output data of the pre-set model based on the collected real-time output data of the model, and when there is no difference between the real-time output data of the model and the output data of the pre-set model, the model is put into application. The method for determining whether there is a difference between the real-time output data of the collected model and the output data of the pre-set model includes the following expression: Where ε is a preset constant, C n Indicates the nth real-time output data of the collected model, μ n It represents the entropy value of the nth real-time output data of the collected model, N represents the total number of real-time output data of the collected model, 1≤n≤N, n is a positive integer, Y m represents the mth output data of the preset model, |·| represents the absolute value, τ m It represents the entropy value of the mth output data of the preset model, M represents the total number of output data of the preset model, 1≤m≤M, and m is a positive integer.
7. The model output control method based on reinforcement learning according to claim 1, characterized in that: It also includes the following steps: If the output data of the model needs to be changed, the output control model is trained based on the changed output data, and the output data of the model is controlled by the trained output control model.
8. The model output control method based on reinforcement learning according to claim 1, characterized in that: It also includes the following steps: The reinforcement learning model performs deep reinforcement learning on the real-time output data of the collected model and the output data of the preset model, outputs control data of the model output, and inputs the control data into the output control model.
9. The model output control method based on reinforcement learning according to claim 8, characterized in that: It also includes the following steps: The control instruction data output by the output control model and the model output control data output by the reinforcement learning model are integrated to generate final model output control data, and the output of the model is controlled by the final model output control data.
10. The model output control method based on reinforcement learning according to claim 5, characterized in that: It also includes the following steps: The reinforcement learning model performs deep reinforcement learning on the detected abnormal output data and the output data of the trained model.