A method, device and equipment for automatic evolution of a model based on reinforcement learning

By using random sampling, inference and variation to generate amplified sample sets in the reinforcement learning method, the model is retrained, and the problem that reinforcement learning cannot achieve automatic evolution of the model is solved, the automatic evolution and optimization of the model is realized, and the manual workload is reduced.

CN119623568BActive Publication Date: 2025-07-25NO 15 INST OF CHINA ELECTRONICS TECH GRP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510147901.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-07-25
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

In the prior art, reinforcement learning methods only involve the construction process before the model is launched, and cannot realize the retraining process after the model is launched, and rely entirely on external input reward and punishment information, so the automatic evolution of the model cannot be achieved.

Method used

The random sampling method is used to extract data from the original data set, generate a new data set, inference is performed based on the first local model, and the inference result set is formed, and the inference result set is mutated to form an amplified sample set. The first local model is retrained by amplified sample set and the training sample set to generate a second local model. If the accuracy of the second local model is not lower than the accuracy of the first local model, evolution will be stopped, and the first local model will be used as the online model.

Benefits of technology

The automatic evolution ability of the model is realized without manual participation and external input, which significantly reduces the workload of model developers, can form high-quality amplified samples, and optimizes the model ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119623568B_ABST
    Figure CN119623568B_ABST
Patent Text Reader

Abstract

This application relates to the field of computer technology. Embodiments of this specification disclose a method, apparatus, and device for automatically evolving a model based on reinforcement learning, including: extracting data from an original data set by means of random sampling to generate a new data set; performing inference on the new data set based on a first local model to form an inference result set; mutating the inference result set to form an augmented sample set; retraining the first local model with the augmented sample set and a training sample set to generate a second local model; if the accuracy of the second local model ≤ the accuracy of the first local model, stop the evolution, and use the first local model as the online model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular, to a method, apparatus, and device for automatically evolving a model based on reinforcement learning. Background Art

[0002] During the model training / retraining process, the quality of training samples is directly related to the inference effect of the model. The training samples used in model training / retraining mainly rely on the method of manual labeling. This method requires a large amount of manual work and the labeling levels are also inconsistent, making it difficult to meet the model training / retraining scenarios under large-scale samples. Therefore, how to achieve the model evolution of the model and make the model "get better and better" is an important problem that needs to be solved.

[0003] In the current technology, reinforcement learning is used for model construction. Reinforcement learning is an online learning method that adopts the idea of "trial and error". Different from supervised learning and unsupervised learning, it does not require any data to be given in advance, but obtains corresponding learning information by receiving the feedback of the environment on actions, and adjusts and optimizes the relevant parameters of the model accordingly to achieve the continuous improvement of the model. This method only involves the construction process before the model goes online and does not involve the retraining process after the model goes online. Moreover, this method completely depends on the externally input reward and punishment information to complete the training of the model itself and cannot achieve the automatic evolution of the model.

[0004] Based on this, a new method for automatically evolving a model based on reinforcement learning is needed. Summary of the Invention

[0005] Embodiments of this specification provide a method, apparatus, and device for automatically evolving a model based on reinforcement learning to solve the following technical problems: In the current technology, reinforcement learning is used for model construction. Reinforcement learning is an online learning method that adopts the idea of "trial and error". Different from supervised learning and unsupervised learning, it does not require any data to be given in advance, but obtains corresponding learning information by receiving the feedback of the environment on actions, and adjusts and optimizes the relevant parameters of the model accordingly to achieve the continuous improvement of the model. This method only involves the construction process before the model goes online and does not involve the retraining process after the model goes online. Moreover, this method completely depends on the externally input reward and punishment information to complete the training of the model itself and cannot achieve the automatic evolution of the model.

[0006] To solve the above technical problems, the embodiments of this specification are implemented as follows:

[0007] Embodiments of this specification provide a method for automatically evolving a model based on reinforcement learning, including:

[0008] Randomly sample data from the original data set to generate a new data set;

[0009] Perform inference on the new data set based on the first local model to form an inference result set;

[0010] Mutate the inference result set to form an augmented sample set;

[0011] Retrain the first local model with the augmented sample set and the training sample set to generate a second local model;

[0012] If the accuracy of the second local model ≤ the accuracy of the first local model, stop evolution, and use the first local model as the online model.

[0013] This embodiment of the specification also provides a model automatic evolution device based on reinforcement learning, including:

[0014] A random generation module that uses random sampling to extract data from the original data set to generate a new data set;

[0015] An inference module that performs inference on the new data set based on the first local model to form an inference result set;

[0016] A mutation module that mutates the inference result set to form an augmented sample set;

[0017] A retraining module that retrains the first local model with the augmented sample set and the training sample set to generate a second local model;

[0018] An online module that, if the accuracy of the second local model ≤ the accuracy of the first local model, stops evolution, and uses the first local model as the online model.

[0019] This embodiment of the specification also provides an electronic device, including:

[0020] At least one processor; and,

[0021] A memory communicatively connected to the at least one processor; wherein,

[0022] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to:

[0023] Use random sampling to extract data from the original data set to generate a new data set;

[0024] Perform inference on the new data set based on the first local model to form an inference result set;

[0025] Mutate the inference result set to form an augmented sample set;

[0026] Retrain the first local model with the amplified sample set and the training sample set to generate a second local model;

[0027] If the accuracy rate of the second local model ≤ the accuracy rate of the first local model, stop the evolution, and use the first local model as the online model.

[0028] The model automatic evolution method based on reinforcement learning provided by the embodiments of this specification uses a random sampling method to extract data from the original data set to generate a new data set; based on the first local model, perform inference on the new data set to form an inference result set; mutate the inference result set to form an amplified sample set; retrain the first local model with the amplified sample set and the training sample set to generate a second local model; if the accuracy rate of the second local model ≤ the accuracy rate of the first local model, stop the evolution, and use the first local model as the online model. It can form high-quality amplified samples through the automatic extraction, automatic mutation, and automatic evaluation of the model inference results, thereby realizing the automatic evolution ability of the model. The entire process requires no manual participation and other external inputs, significantly reducing the workload of model developers and can be effectively applied to various data analysis business fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0030] Figure 1 It is a schematic diagram of the system architecture of a model automatic evolution method based on reinforcement learning provided by the embodiments of this specification;

[0031] Figure 2 It is a schematic flowchart of a model automatic evolution method based on reinforcement learning provided by the embodiments of this specification;

[0032] Figure 3 It is a schematic flowchart of the generation process of the amplified sample set provided by the embodiments of this specification;

[0033] Figure 4 A framework diagram of a model automatic evolution method based on reinforcement learning provided by the embodiments of this specification;

[0034] Figure 5 A schematic diagram of a model automatic evolution device based on reinforcement learning provided by the embodiments of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0036] Reinforcement learning is an online learning method that adopts the idea of "trial and error". Different from supervised learning and unsupervised learning, it does not require any pre-given data. Instead, it obtains corresponding learning information by receiving the feedback of the environment on actions, and adjusts and optimizes the relevant parameters of the model accordingly to continuously improve the model. According to the given conditions, reinforcement learning can be divided into model-based reinforcement learning and model-free reinforcement learning.

[0037] Therefore, the embodiments of this specification provide a new method for automatic model evolution based on reinforcement learning, which uses the method of "self-extracting inference results, self-mutating inference results, and self-evaluating inference results" to realize the ability to discover augmented samples for model inference results, so as to support subsequent model retraining. It only depends on the existing evaluation samples and inference results inside, without other external inputs, and continuously optimizes the model capabilities.

[0038] Figure 1 It is a schematic diagram of the system architecture of a method for automatic model evolution based on reinforcement learning provided by the embodiments of this specification. As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0039] The terminal devices 101, 102, 103 interact with the server 105 through the network 104 to receive or send messages, etc. Various client applications may be installed on the terminal devices 101, 102, 103. For example, special programs such as a method for automatic model evolution based on reinforcement learning.

[0040] The terminal devices 101, 102, and 103 can be hardware or software. When the terminal devices 101, 102, and 103 are hardware, they can be various dedicated or general-purpose electronic devices, including but not limited to smartphones, tablets, laptop computers, and desktop computers, etc. When the terminal devices 101, 102, and 103 are software, they can be installed in the above-listed electronic devices. It can be implemented as multiple software or software modules (such as multiple software or software modules for providing distributed services), or it can be implemented as a single software or software module.

[0041] The server 105 can be a server that provides various services. For example, it can be a backend server that provides services for the client applications installed on the terminal devices 101, 102, and 103. For example, the server can perform automatic evolution of the model based on reinforcement learning so as to display the model automatic evolution results on the terminal devices 101, 102, and 103. The server can also perform automatic evolution of the model based on reinforcement learning so as to display the model automatic evolution results on the terminal devices 101, 102, and 103.

[0042] The server 105 can be hardware or software. When the server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or it can be implemented as a single server. When the server 105 is software, it can be implemented as multiple software or software modules (such as multiple software or software modules for providing distributed services), or it can be implemented as a single software or software module.

[0043] Figure 2 It is a schematic flowchart of a method for automatic evolution of a model based on reinforcement learning provided by the embodiments of this specification. From a program perspective, the execution subject of the process can be a program running on an application server or an application terminal. It can be understood that this method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities. As Figure 2 shown, the model automatic evolution method includes:

[0044] Step S201: Extract data from the original data set by means of random sampling to generate a new data set.

[0045] In the embodiments of this specification, the original data set D is an unprocessed data set, which includes data that can be used for model training and also includes data that is useless for model training. Therefore, it needs to go through steps such as cleaning and preprocessing to remove noise, fill in missing values, convert formats, or extract new features to be suitable for model training. The original data set D is numerical data and / or data that can be converted into numerical data. Specifically, the original data set D is in the key-value format, that is, key-value pairs, where key represents the index value and value represents the value.

[0046] From the original dataset D, randomly extract data with the first preset ratio x% to form a new dataset D1. The selection of x depends on the specific business scenario. In one embodiment, x is defaulted to 10 and its value is adjustable, and no specific limitation is made here. Similarly, the new dataset D1 is numerical data and / or data that can be converted into numerical data. In the illustrated embodiment, the new dataset D1 is used for the first local model M1.

[0047] Step S203: The new dataset is inferred by the first local model to form an inference result set.

[0048] Continuing with the previous example, after preprocessing and other steps on the original dataset D, a training dataset S is formed. The training dataset S is a data set for training a model. Similarly, the training dataset S is numerical data and / or data that can be converted into numerical data.

[0049] In this embodiment, the training dataset S is used to train the local model M to form the first local model M1. Of course, the first local model M1 can also be a model that has been trained. That is to say, the first local model is a model obtained by training or a model that has completed training.

[0050] In the illustrated embodiment of this specification, the first local model can process numerical data and / or data that can be converted into numerical data. It should be particularly noted that in the illustrated embodiment of this specification, the data involved are all numerical data and / or data that can be converted into numerical data, and will not be elaborated hereinafter.

[0051] It can be understood that the method provided in the illustrated embodiment of this specification is used to process numerical data and / or data that can be converted into numerical data, and thus to retrain the model corresponding to the numerical data and / or data that can be converted into numerical data.

[0052] The method provided in the illustrated embodiment of this specification, specifically in terms of data types, can be image data, and / or text data, and / or video data, and / or audio data, etc.

[0053] The first local model is a model that has completed training or a model obtained by training, that is, the first local model is an on-line model. Since for a model that has completed training, during the training process, there may be certain differences between the data used as training samples and the real data of the model user. To ensure the accuracy of the data recognition result, therefore, after the first local model is deployed and goes on-line, that is, during the usage stage after model deployment, model adjustment is required, that is, model auto-evolution. It can be seen that the model auto-evolution method provided in this embodiment is not applied before the model is deployed and goes on-line, but is applied during the usage stage after model deployment.

[0054] Continuing with the previous example, the first local model M1 is used to infer the new dataset D1, forming an inference result set D1' with inference results.

[0055] Step S205: Infer the new dataset based on the first local model to form an inference result set.

[0056] Since the accuracy rate of the data quality of the inference result set D1' is not high and model evolution cannot be performed well, based on this, the inference result set D1' is mutated to form an inference result set D1''.

[0057] In the embodiments of this specification, mutating the inference result set to form an amplified sample set specifically includes:

[0058] Using the inference result set as the initial amplified sample set;

[0059] Randomly selecting data from the initial amplified sample set as the new inference result set;

[0060] Mutating the new inference result set to form a mutated inference result set;

[0061] If the accuracy rate of the mutated inference result set > the accuracy rate of the new inference result set, then update the amplified sample set with the mutated inference result set to obtain an updated amplified sample set;

[0062] Using the updated amplified sample set as the initial amplified sample set, performing cyclic operations and mutations to form the amplified sample set.

[0063] In the embodiments of this specification, mutating the new inference result set to form a mutated inference result set specifically includes:

[0064] From the initial amplified sample set, determining the target records in the new inference result set, using the value with the highest weight sum of the target records as the first candidate value, and the value with the second highest weight sum of the target records as the second candidate value;

[0065] According to the first preset probability, replacing each record in the new inference result set with the first candidate value, and according to the second preset probability, replacing each record in the new inference result set with the second candidate value to form the mutated inference result set, where the sum of the first preset probability and the second preset probability is 1.

[0066] In the embodiments of this specification, determining the target records in the new inference result set from the initial amplified sample set specifically includes:

[0067] From the initialized amplification sample set, select several records that are closest to each record in the new inference result set as the target records in the new inference result set.

[0068] In the embodiments of this specification, use the updated amplification sample set as the initialized amplification sample set, perform cyclic operations, mutate, and form the amplification sample set. The default value of the number of times n for performing cyclic operations is 10. The number of times n for cyclic operations can be adjusted according to specific business scenarios and is not limited herein.

[0069] The initialized amplification sample set refers to the initial state of the amplification sample set. Continuing with the previous example, use the inference result set D1’ as the initialized amplification sample set D01”. Randomly select data with a second preset ratio y% from the initialized amplification sample set as the new inference result set E. The value range of y is 1 to 10, and the selection of y depends on the specific business scenario. In one embodiment, y defaults to 5 and its value is adjustable and is not specifically limited herein.

[0070] For each record d in the new inference result set E, obtain z records that are closest to the distribution distance of record d from the initialized amplification sample set D01” as the target records in the new inference result set. The weight value corresponding to each record is calculated as 1 / (z + the ranking of the distribution distance between the record and record d from small to large). The sum of the weight values of each record is used as the weight sum of the target record. Select the value with the highest weight sum among the z records as the first candidate value v1, select the value with the second highest weight sum as the candidate value v2, replace the value of record d with the first candidate value v1 with a probability of the first preset probability w, and replace the value of record d with the second candidate value v2 with a probability of the second preset probability (1 - w) to form the mutated inference result set E’.

[0071] In this embodiment, the value range of z is 1 to 5, the default is 5, and the value of z can be determined according to the business scenario and its value can be adjusted. Each record d in the inference result set E is <key-value>Form. The value range of the first preset threshold w is 0.8 to 1.0, and the default value is 0.9.

[0072] Continuing with the previous example, the evaluation sample set C is formed by selecting and annotating part of the data from the original data set D. Therefore, the original data set D covers the evaluation sample set C.

[0073] The accuracy rate q1 of the mutated inference result set E' is evaluated based on the evaluation sample set C. Specifically, the accuracy rate q1 of the mutated inference result set E' = the number of samples with the same value from the mutated inference result set E' and the evaluation sample set C / count (the key repetition set of the mutated inference result set E' and the evaluation sample set C).

[0074] Similarly, the accuracy rate q2 of the new inference result set E is evaluated based on the evaluation sample set C. Specifically, the accuracy rate q2 of the new inference result set E = the number of samples with the same value from the new inference result set E and the evaluation sample set C / count (the key repetition set of the new inference result set E and the evaluation sample set C).

[0075] That is: if the accuracy rate q1 of the mutated inference result set E' > the accuracy rate q2 of the new inference result set E, then update the mutated inference result set to the amplified sample set.

[0076] In the embodiments of this specification, if the accuracy rate of the mutated inference result set > the accuracy rate of the new inference result set, then update the mutated inference result set to the amplified sample set to obtain an updated amplified sample set, which further includes:

[0077] If the accuracy rate of the updated amplified sample set ≤ the accuracy rate of the inference result set, stop mutation, and replace the updated amplified sample set with the inference result set as the amplified sample set;

[0078] If the accuracy rate of the updated amplified sample set > the accuracy rate of the inference result set, and the difference between the accuracy rate of the updated amplified sample set and the accuracy rate of the inference result set < the second preset threshold, stop mutation, and use the updated amplified sample set as the amplified sample set;

[0079] If the accuracy rate of the updated amplified sample set > the accuracy rate of the inference result set, and the difference between the accuracy rate of the updated amplified sample set and the accuracy rate of the inference result set ≥ the second preset threshold, then use the updated amplified sample set as the initial amplified sample set.

[0080] Continuing with the previous example, the accuracy rate q3 of the updated amplified sample set D1” is evaluated based on the evaluation sample set C. Specifically, the accuracy rate q3 of the updated amplified sample set D1” = the number of samples with the same value from the updated amplified sample set D1” and the evaluation sample set C / count (the set of keys that are duplicates between the updated amplified sample set D1” and the evaluation sample set C).

[0081] Similarly, the accuracy rate q4 of the inference result set D1’ is evaluated based on the evaluation sample set C. Specifically, the accuracy rate q4 of the inference result set D1’ = the number of samples with the same value from the inference result set D1’ and the evaluation sample set C / count (the set of keys that are duplicates between the inference result set D1’ and the evaluation sample set C).

[0082] That is: the accuracy rate q3 of the updated amplified sample set D1” ≤ the accuracy rate q4 of the inference result set D1’, stop the mutation, and replace the updated amplified sample set with the inference result set as the amplified sample set.

[0083] If the accuracy rate q3 of the updated amplified sample set D1” > the accuracy rate q4 of the inference result set D1’, and q3 - q4 < the second preset threshold p2, then stop the mutation, and use the updated amplified sample set as the amplified sample set;

[0084] If the accuracy rate q3 of the updated amplified sample set D1” > the accuracy rate q4 of the inference result set D1’, and q3 - q4 ≥ the second preset threshold p2, then stop the mutation, and repeat the following steps: randomly select data from the initialized amplified sample set as the new inference result set.

[0085] The selection of the second preset threshold p2 should not be greater than the second preset ratio, that is, the second preset threshold ≤ the second preset ratio. In this embodiment, the default value of the second preset threshold p2 is 0.05%, and the second preset threshold p2 can be adjusted according to the specific business scenario, and no specific limitation is made here.

[0086] To further understand the generation of the amplified sample set, the generation process of the amplified sample set is described below. Figure 3 It is a schematic diagram of the generation process of the amplified sample set provided by the embodiment of this specification. As Figure 3 As shown, use the inference result set as the initial augmented sample set; adopt a random sampling method to extract data from the initial augmented sample set to generate a new inference result set; based on the new inference result set, perform mutation to form a mutated inference result set; if the accuracy rate q1 of the mutated inference result set > the accuracy rate q2 of the new inference result set, then use the mutated inference result set to update the augmented sample set; perform loop operations to form an updated augmented sample set; if the accuracy rate q3 of the updated augmented sample set ≤ the accuracy rate q4 of the inference result set, then stop mutation, and replace the inference result set with the updated augmented sample set as the augmented sample set; if the accuracy rate q3 of the updated augmented sample set D1” > the accuracy rate q4 of the inference result set D1’, and q3 - q4 ≥ the second preset threshold p2, then stop mutation, and repeat the following steps: randomly select data from the initial augmented sample set as the new inference result set, and thus perform the subsequent operations of "randomly select data from the initial augmented sample set as the new inference result set".

[0087] Step S207: Retrain the first local model with the augmented sample set and the training sample set to generate a second local model.

[0088] When retraining the first local model, if the keys of the augmented sample set and the training sample set are the same but the values are different, that is, there is a value conflict under the same key, then use the value of the training sample set as the result value.

[0089] Step S209: If the accuracy rate of the second local model ≤ the accuracy rate of the first local model, then stop evolution, and use the first local model as the online model.

[0090] The determination of the accuracy rate of the first local model and the accuracy rate of the second local model are both obtained based on the evaluation of the evaluation sample set C. Among them, the accuracy rate q5 of the first local model = (the number of results that are the same as the evaluation result set obtained by predicting the samples of the evaluation sample set using the first local model) / the number of the evaluation sample set; the accuracy rate q6 of the second local model = (the number of results that are the same as the evaluation result set obtained by predicting the samples of the evaluation sample set using the second local model) / the number of the evaluation sample set. If q6 ≤ q5, then stop evolution, and use the first local model as the online model.

[0091] In the embodiments of this specification, the method further includes:

[0092] If the accuracy rate of the second local model > the accuracy rate of the first local model, and (the accuracy rate of the second local model - the accuracy rate of the first local model) < the first preset threshold, then stop evolution, and use the second local model as the online model;

[0093] If the accuracy rate of the second local model > the accuracy rate of the first local model, and (the accuracy rate of the second local model - the accuracy rate of the first local model) ≥ the first preset threshold, then add the augmented sample set to the training sample set as the updated training sample set, and replace the first local model with the second local model, and repeat the following steps: randomly sample data from the original dataset to generate a new dataset.

[0094] Continuing with the previous example, if q6 > q5 and q6 - q5 < the first preset threshold p1, then stop the evolution and put the second local model M2 online; if q6 > q5 and q6 - q5 ≥ the first preset threshold p1, then add the augmented sample set to the training sample set as the updated training sample set, and replace the first local model with the second local model, and repeat the following steps: randomly sample data from the original dataset to generate a new dataset, so as to perform the subsequent operations of "randomly sample data from the original dataset to generate a new dataset".

[0095] The determination of the first preset threshold depends on the specific business scenario, and the default value of the first preset threshold p1 is 0.1%.

[0096] The model automatic evolution method based on reinforcement learning provided in the embodiments of this specification is used to perform model automatic evolution after the model is deployed and put online.

[0097] To further understand the model automatic evolution method based on reinforcement learning provided in the embodiments of this specification, the embodiments of this specification further provide a framework diagram of the model automatic evolution method based on reinforcement learning. As Figure 4 As shown in the figure, the local model M is trained using the training sample set S to obtain the first local model M1. Data D is randomly sampled from the original dataset to generate a new dataset D1. The first local model M1 is used to perform inference on the new dataset D1 to form an inference result set D1' with inference results. The inference result set D1' is mutated to form an augmented sample set D1". The first local model M1 is retrained using the training sample set S + the augmented sample set D1" to obtain the second local model M2, and the accuracy q5 of the first local model M1 and the accuracy q6 of the second local model M2 are respectively evaluated using the evaluation sample set C. If q6 ≤ q5, the evolution is stopped, and the first local model M1 is put online, with the first local model M1 serving as the online model. If q6 > q5 and q6 - q5 < the first preset threshold, the evolution is stopped, and the second local model M2 is put online, with the second local model M2 replacing the first local model M1 as the online model. If q6 > q5 and q6 - q5 ≥ the first preset threshold, the augmented sample set is added to the training sample set as the updated training sample set, and the second local model replaces the first local model. The following steps are repeated: Data is randomly sampled from the original dataset to generate a new dataset.

[0098] The model automatic evolution method based on reinforcement learning provided in the embodiments of this specification randomly samples data from the original dataset to generate a new dataset, performs inference on the new dataset based on the first local model to form an inference result set, mutates the inference result set to form an augmented sample set, retrains the first local model with the augmented sample set and the training sample set to generate a second local model. If the accuracy of the second local model ≤ the accuracy of the first local model, the evolution is stopped, and the first local model serves as the online model. It can automatically extract, mutate, and evaluate the model inference results to form high-quality augmented samples, thereby realizing the automatic evolution ability of the model. The entire process requires no manual participation and other external inputs, significantly reducing the workload of model developers and can be effectively applied to various data analysis business fields.

[0099] The above content details a model automatic evolution method based on reinforcement learning. Correspondingly, this specification also provides a model automatic evolution device based on reinforcement learning, as Figure 5 shown. Figure 5 It is a schematic diagram of a model automatic evolution device based on reinforcement learning provided in the embodiments of this specification. The model automatic evolution device includes:

[0100] A random generation module 501 that randomly samples data from the original dataset to generate a new dataset;

[0101] The inference module 503 performs inference on the new data set based on the first local model to form an inference result set;

[0102] The mutation module 505 mutates the inference result set to form an amplified sample set;

[0103] The retraining module 507 retrains the first local model with the amplified sample set and the training sample set to generate a second local model;

[0104] The online module 509 stops the evolution if the accuracy rate of the second local model ≤ the accuracy rate of the first local model, and the first local model is used as the online model.

[0105] In the embodiments of this specification, the model automatic evolution device further includes:

[0106] The update module 511 stops the evolution if the accuracy rate of the second local model > the accuracy rate of the first local model, and (the accuracy rate of the second local model - the accuracy rate of the first local model) < the first preset threshold, and the second local model is used as the online model;

[0107] If the accuracy rate of the second local model > the accuracy rate of the first local model, and (the accuracy rate of the second local model - the accuracy rate of the first local model) ≥ the first preset threshold, then the inference result set is added to the training sample set as an updated training sample set, the second local model replaces the first local model, and the following steps are repeated: data is randomly sampled from the original data set to generate a new data set.

[0108] The embodiments of this specification also provide an electronic device, including:

[0109] At least one processor; and,

[0110] A memory communicatively connected to the at least one processor; wherein,

[0111] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can:

[0112] Randomly sample data from the original data set to generate a new data set;

[0113] Perform inference on the new data set based on the first local model to form an inference result set;

[0114] Mutate the inference result set to form an amplified sample set;

[0115] Retrain the first local model with the amplified sample set and the training sample set to generate a second local model;

[0116] If the accuracy rate of the second local model ≤ the accuracy rate of the first local model, stop the evolution, and use the first local model as the online model.

[0117] The specific embodiments of this specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0118] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the apparatus, electronic device, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiments.

[0119] The apparatus, electronic device, and non-volatile computer storage medium provided in the embodiments of this specification correspond to the method. Therefore, the apparatus, electronic device, and non-volatile computer storage medium also have beneficial technical effects similar to the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the corresponding apparatus, electronic device, and non-volatile computer storage medium will not be repeated here.

[0120] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (e.g., a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a hardware description language (HDL), and there is not just one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0121] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.

[0122] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0123] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing one or more embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0124] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, the embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0125] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or in one block or multiple blocks.

[0126] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or in one block or multiple blocks.

[0127] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or in one block or multiple blocks.

[0128] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0129] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.

[0130] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0131] It should also be noted that the term "includes", "comprising" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the existence of other identical elements in the process, method, commodity or device including the element.

[0132] The specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media including storage devices.

[0133] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0134] The above description is only for the embodiments of this specification and is not intended to limit this application. For those skilled in the art, various changes and modifications can be made to this application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the scope of the claims of this application.

Claims

1. A method for automatic evolution of a model based on reinforcement learning, characterized in that The described model automatic evolution method is applied to the usage stage after model deployment, and includes: Adopting a random sampling method to extract data from the original dataset to generate a new dataset; Inferring the new dataset based on the first local model to form an inference result set, where the first local model is the online model and is used to process image data, or text data, or video data, or audio data; Mutating the inference result set to form an augmented sample set, specifically including: using the inference result set as the initial augmented sample set; randomly selecting data from the initial augmented sample set as the new inference result set; mutating the new inference result set to form a mutated inference result set; if the accuracy of the mutated inference result set > the accuracy of the new inference result set, then update the augmented sample set with the mutated inference result set to obtain an updated augmented sample set; using the updated augmented sample set as the initial augmented sample set, and performing cyclic operations to mutate and form the augmented sample set; Retraining the first local model with the augmented sample set and the training sample set to generate a second local model; If the accuracy of the second local model ≤ the accuracy of the first local model, stop evolution, and use the first local model as the online model.

2. The model automatic evolution method according to claim 1, wherein The method further includes: If the accuracy of the second local model > the accuracy of the first local model, and (the accuracy of the second local model - the accuracy of the first local model) < the first preset threshold, stop evolution, and use the second local model as the online model; If the accuracy of the second local model > the accuracy of the first local model, and (the accuracy of the second local model - the accuracy of the first local model) ≥ the first preset threshold, then add the inference result set to the training sample set as the updated training sample set, replace the first local model with the second local model, and repeat the steps: adopting a random sampling method to extract data from the original dataset to generate a new dataset.

3. The model automatic evolution method according to claim 1, wherein The mutating the new inference result set to form a mutated inference result set specifically includes: Determining the target records in the new inference result set from the initial augmented sample set, using the value with the highest weight sum of the target records as the first candidate value, and the value with the second highest weight sum of the target records as the second candidate value; Replacing each record in the new inference result set with the first candidate value according to the first preset probability, and replacing each record in the new inference result set with the second candidate value according to the second preset probability to form the mutated inference result set, where the sum of the first preset probability and the second preset probability is 1.

4. The model automatic evolution method according to claim 3, characterized in that, The determining the target records in the new inference result set from the initial augmented sample set specifically includes: Selecting several records closest to each record in the new inference result set from the initial augmented sample set as the target records in the new inference result set.

5. The model automatic evolution method according to claim 1, wherein If the accuracy rate of the mutated inference result set > the accuracy rate of the new inference result set, updating the amplified sample set with the mutated inference result set to obtain an updated amplified sample set further includes: If the accuracy rate of the updated amplified sample set ≤ the accuracy rate of the inference result set, stop mutation, and replace the updated amplified sample set with the inference result set as the amplified sample set; If the accuracy rate of the updated amplified sample set > the accuracy rate of the inference result set, and the difference between the accuracy rate of the updated amplified sample set and the accuracy rate of the inference result set < the second preset threshold, stop mutation, and use the updated amplified sample set as the amplified sample set; If the accuracy rate of the updated amplified sample set > the accuracy rate of the inference result set, and the difference between the accuracy rate of the updated amplified sample set and the accuracy rate of the inference result set ≥ the second preset threshold, repeat the step: randomly select data from the initialized amplified sample set as the new inference result set.

6. The model automatic evolution method according to claim 1, characterized in that The first local model is used to process numerical data or data that can be converted into numerical data.

7. The model automatic evolution method according to claim 1, characterized in that The model automatic evolution method is used for automatic model evolution after the model is deployed online.

8. An automatic model evolution device based on reinforcement learning, characterized in that, The model automatic evolution device is applied to the usage stage after the model is deployed, and includes: A random generation module that uses random sampling to extract data from the original data set to generate a new data set; An inference module that infers the new data set based on the first local model to form an inference result set. The first local model is an online model and is used to process image data, or text data, or video data, or audio data; A mutation module that mutates the inference result set to form an amplified sample set, specifically including: using the inference result set as the initialized amplified sample set; randomly selecting data from the initialized amplified sample set as the new inference result set; mutating the new inference result set to form a mutated inference result set; if the accuracy rate of the mutated inference result set > the accuracy rate of the new inference result set, updating the amplified sample set with the mutated inference result set to obtain an updated amplified sample set; using the updated amplified sample set as the initialized amplified sample set, and performing cyclic operations for mutation to form the amplified sample set; A retraining module that retrains the first local model with the amplified sample set and the training sample set to generate a second local model; An online module that, if the accuracy rate of the second local model ≤ the accuracy rate of the first local model, stops evolution, and uses the first local model as the online model.

9. An electronic device, including: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor can: Use random sampling to extract data from the original data set to generate a new data set; Infer the new data set based on the first local model to form an inference result set. The first local model is an online model and is used to process image data, or text data, or video data, or audio data; Mutate the inference result set to form an augmented sample set, specifically including: using the inference result set as the initial augmented sample set; randomly selecting data from the initial augmented sample set as the new inference result set; mutating the new inference result set to form a mutated inference result set; if the accuracy of the mutated inference result set > the accuracy of the new inference result set, then update the augmented sample set with the mutated inference result set to obtain an updated augmented sample set; use the updated augmented sample set as the initial augmented sample set, and perform cyclic operations for mutation to form the augmented sample set; Retrain the first local model with the augmented sample set and the training sample set to generate a second local model; If the accuracy of the second local model ≤ the accuracy of the first local model, stop the evolution, and use the first local model as the online model.

Citation Information

Patent Citations

  • Data identification method, automatic continuous learning model, device and equipment

    CN115859122A

  • Method and device for constructing endogenous security artificial intelligence system based on dynamic heterogeneous redundancy

    CN116595511A