Value function continuous learning method and system based on online data
Through the online data-driven continuous learning method of value functions, the generative model and neural network extract fluid states and value scalars are solved, and the problem of difficult to achieve universal human value function modeling in the existing technology is realized, and the continuous learning and updating of value functions is achieved, and complex and multi-dimensional human value modeling is supported.
Patent Information
- Application Number
- CN202510475337.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing value function learning methods are difficult to achieve universal human value function modeling, and lack clear numerical value functions, making it difficult to support complex and multi-dimensional human value modeling.
Through the continuous learning method of value functions based on online data, data samples are generated using the generative model, and the flow state and value scalar are extracted through the neural network model, and the value function is continuously trained based on this information.
It realizes continuous learning and update of value functions, ensures effective update of artificial intelligence systems and improves the adaptability of value functions, and supports complex and multi-dimensional human value modeling.
Smart Images

Figure CN119989068A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and system for continuous learning of value functions based on online data. Background Art
[0002] In the fields of psychology and cognitive science, a relatively mature theoretical system has been formed for the definition and classification of human values. For example, Maslow's Hierarchy of Needs divides human needs into five levels, arguing that needs are progressive: from physiological needs to safety needs, social needs, respect needs, and finally to the highest self-actualization needs. This theory shows that human values are gradually satisfied in different levels of needs. The Schwartz Theory of Basic Values proposes ten basic values from the perspective of social psychology, covering multiple dimensions from power, achievement, enjoyment to fraternity, universalism, etc., emphasizing the universality of values and cross-cultural applicability. The above theories provide a theoretical basis for understanding the motivation of human behavior, but their frameworks are essentially qualitative descriptions and lack clear numerical modeling methods.
[0003] In machine learning and artificial intelligence research, the calculation of value functions is often associated with reinforcement learning. In reinforcement learning, value functions are used to estimate the future rewards of an agent in a given state. However, existing value function learning mainly focuses on the optimization of specific tasks, such as expressing multi-objective values under different tasks through generalized value functions (GVFs). Although these methods enhance task flexibility, they are usually only effective in specific scenarios and are difficult to generalize to universal human value function modeling.
[0004] On the other hand, in the field of artificial intelligence in recent years, value alignment has become a hot topic, with the goal of aligning the decision-making process of artificial intelligence systems with human values. The main methods for implementing value alignment include human-in-the-loop reinforcement learning and preference learning. Human-in-the-loop reinforcement learning adjusts the behavior of AI through direct human feedback, while preference learning gradually internalizes human values by observing human choices. These methods rely on example data in an attempt to align AI output with human expectations. However, current value alignment research is more about adaptive adjustments for specific tasks or scenarios, lacks a clear numerical value function, and is difficult to support complex, multi-dimensional human value modeling. How to achieve continuous learning of value functions mainly depends on the collection and update of data, that is, how to effectively collect corresponding data is an important core of continuous learning of value functions. Summary of the invention
[0005] One of the purposes of the present invention is to provide a method and system for continuous learning of value functions based on online data, so as to realize effective and continuous collection of online data, ensure continuous learning and updating of value functions, and ensure effective updating of artificial intelligence systems.
[0006] An embodiment of the present invention provides a method for continuous learning of a value function based on online data, comprising: Based on the current learning value, the expected value is inferred and the corresponding flow state is further determined, and the first data sample corresponding to the flow state is generated using the generation model; Publish tasks and task scenarios, receive other users' operations on tasks to generate corresponding scenario data; process the scenario data to obtain a second data sample; The first data sample and the second data sample are added to the original training data sample, and after the training data sample is processed, the value function is retrained.
[0007] Preferably, the steps for training the value function are as follows: Extracting flow patterns from data samples; Discover value scalars from the flow state; Train a pre-configured model based on a value scalar.
[0008] Preferably, the rules for publishing tasks and task scenarios are as follows: Filter the information of the training data corresponding to the current value function to determine the data group to be replaced; Analyze each data of the data group to be replaced to determine the description parameter set of the target object, the target task and the task scenario corresponding to the target task; Based on the description parameter set, other users are screened to determine the target object; Send the target tasks and task scenarios to the target objects; When the target object does not respond within the preset time threshold, the screening of other users is repeated.
[0009] Preferably, analyzing each data of the data group to be replaced to determine a description parameter set of a target object, a target task, and a task scenario corresponding to the target task includes: After extracting features from the data, the data is input into a pre-trained and converged neural network model to obtain the call factor; Based on the calling factor, the description parameter set of the target object, the target task and the task scenario corresponding to the target task are called from the preset target object and task analysis library.
[0010] Preferably, based on the description parameter set, other users are screened to determine the target object, including: Matching the description parameter set with the classification parameter sets of other users to obtain a matching result; Evaluate the matching result using a pre-configured first scoring library to determine a first scoring value; Evaluate the activity data of other users using a pre-configured second scoring library to determine a second scoring value; Sorting other users in descending order based on the weighted sum of the first rating value and the second rating value; The other users ranked first are targeted.
[0011] The present invention also provides a continuous learning system for a value function based on online data, comprising: an inference generation module, a publishing and collection module, and a retraining module; wherein the inference generation module infers and generates an expected value based on the current learning value and further determines the corresponding flow state, and uses a generation model to generate a first data sample corresponding to the flow state; the publishing and collection module publishes tasks and task scenarios, receives other users' operations on tasks and generates corresponding scenario data; processes the scenario data to obtain a second data sample; and the retraining module adds the first data sample and the second data sample to the original training data sample, and after processing the training data sample, retrains the value function.
[0012] Preferably, the steps of training the value function by the retraining module are as follows: Extracting flow patterns from data samples; Discover value scalars from the flow state; Train a pre-configured model based on a value scalar.
[0013] Preferably, the rules for publishing tasks and task scenarios are as follows: Filter the information of the training data corresponding to the current value function to determine the data group to be replaced; Analyze each data of the data group to be replaced to determine the description parameter set of the target object, the target task and the task scenario corresponding to the target task; Based on the description parameter set, other users are screened to determine the target object; Send the target tasks and task scenarios to the target objects; When the target object does not respond within the preset time threshold, the screening of other users is repeated.
[0014] Preferably, analyzing each data of the data group to be replaced to determine a description parameter set of a target object, a target task, and a task scenario corresponding to the target task includes: After extracting features from the data, the data is input into a pre-trained and converged neural network model to obtain the call factor; Based on the calling factor, the description parameter set of the target object, the target task and the task scenario corresponding to the target task are called from the preset target object and task analysis library.
[0015] Preferably, based on the description parameter set, other users are screened to determine the target object, including: Matching the description parameter set with the classification parameter sets of other users to obtain a matching result; Evaluate the matching result using a pre-configured first scoring library to determine a first scoring value; Evaluate the activity data of other users using a pre-configured second scoring library to determine a second scoring value; Sorting other users in descending order based on the weighted sum of the first rating value and the second rating value; The other users ranked first are targeted.
[0016] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings.
[0017] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1A schematic diagram of a method for continuous learning of a value function based on online data in an embodiment of the present invention; Figure 2 Schematic diagram of a value function continuous learning system based on online data in an embodiment of the present invention. DETAILED DESCRIPTION
[0019] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0020] The embodiment of the present invention provides a method for continuous learning of value function based on online data, such as Figure 1 As shown, including: Step 1: Based on the current learning value, infer the expected value and further determine the corresponding flow state, and use the generation model to generate the first data sample corresponding to the flow state; Step 2: Publish the task and task scenario, receive other users' operations on the task and generate corresponding scenario data; process the scenario data to obtain a second data sample; Step 3: Add the first data sample and the second data sample to the original training data sample, process the training data sample, and then retrain the value function.
[0021] The working principle and beneficial effects of the above technical solution are: The present invention is based on the continuous learning of the value function based on online data. First, the value function needs to be understood; its specific generation process is as follows: Sample Data: This is the raw data collected from the environment or task, which may be images, audio, video, text or structured data, etc. Fluent / Latent Variable: These are the potential dynamic features or hidden patterns behind the data, which can reflect the changes of the data in different states. Latent variables represent a certain potential attribute or state of a data sample. Value Scalar: This is a numerical value that represents the value of a data sample, usually obtained through the output of the model, representing the quality of the target state or behavior corresponding to the sample. The steps of generating latent variables and sample data are as follows: For sequence data (such as time series, text, etc.), we use RNN or Transformer, where the latent variables are used as the hidden layer input of the network to capture the dependencies and state transitions in the time series. For image data, CNN or UNet is used, where the latent variables are used as additional feature encodings to help extract key features in the image and generate data samples. For graph data (such as social networks or traffic networks), GNN (graph neural network) can be used to process the relationship between nodes in the data through the graph structure, and then infer hidden variables and sample data. Through these different architectures, hidden variables that conform to different modal data can be generated to help establish accurate value functions.
[0022] The specific steps of the online data acquisition process are as follows: 1. Spontaneous exploration by the agent (based on value inference): The agent infers the desired value goal based on the currently learned value system. To achieve this, the agent first identifies possible behavioral goals in the current state and evaluates the potential value of these goals. Through the currently learned value function, the agent is able to infer the potential flow state (hidden variable) from the expected value. For example, in an autonomous driving application, the expected value may be "minimize the risk of collision", and the agent infers the corresponding flow state (e.g., road safety, traffic density, etc.) based on the current road and traffic conditions. Using a generative model (such as VAE or GAN), the agent generates new data samples based on these flow states, such as simulated traffic scenarios or new driving decisions.
[0023] 2. Cloud platform task interface: The cloud platform provides an interface for task scenarios. Users interact with the system through the web interface to complete specific tasks. The system will generate corresponding scenario data (such as user operation logs, time series of completed tasks, etc.) based on user operations.
[0024] These data can be post-processed to obtain higher-level information. For example, raw data (such as operation sequences) can be converted into structured graphs (such as task flow charts or behavioral decision diagrams) through a given interpretation algorithm.
[0025] Through these two methods, the system can continuously generate new data samples and provide rich raw data for the continuous learning of the value function.
[0026] For all kinds of collected data, after processing and labeling, they will form a new data set together with the original data. Using the new data, we must first infer the potential latent variables, namely the flow state, through the sample data and value, and then use the inferred flow state to generate sample data and value scalars. Using the learning theory principles of the latent variable generation model, the newly generated data and value should be as consistent as possible with the original data and value. Based on this, the flow state variables will emerge spontaneously, and how the flow state latent variables generate sample data and value.
[0027] After given sample data, the fluid latent variables can also be inferred, and then the value of the current sample data can be calculated based on the generative model from latent variables to value. This constitutes the calculation from sample data to value, which is dual to the inference of latent variables based on given values in the spontaneous exploration of the intelligent agent to generate data. In fact, any modality can be added to the value system, such as text, audio, video, and structure diagram. These modalities are all generated by fluid latent variables. In practical applications, the latent variables can be inferred based on any given modal data and other modalities and values of the current modal data can be generated, which greatly increases the scalability of the value system. Adding modality (text, audio, video, etc.): For given modal data (such as images, text or audio), firstly process the data through existing neural network models (such as CNN, RNN, etc.) to extract the key features in the data. Then, through the latent variable inference model (such as variational autoencoder, generative adversarial network, etc.), we can infer the latent variables behind these data (such as object features in images, emotional features in text). Then, using the generative model from latent variables to value (such as MLP or other neural networks), the value corresponding to each modal data is calculated. For example, for a piece of text, its emotional value (positive, negative, neutral, etc.) can be inferred based on its latent variables, and its value to the user experience can be further calculated. This method can unify data of different modalities (such as text, images, videos, etc.) into the same value system for calculation and optimization, while greatly improving the scalability of the system and supporting more data modalities and complex application scenarios.
[0028] In the process of continuous data collection and learning, the neural network interpretability theory, such as sparse autoencoders, can be used to detect the fluid latent variables and obtain the specific meaning of each basis vector in the latent variable space, which represents a certain attribute of the sample data. Based on the correspondence between this attribute and the value function, more interpretable and controllable value alignment, task planning and data generation can be performed.
[0029] In order to achieve retraining of the value function; in one embodiment, the steps of training the value function are as follows: Extracting flow patterns from data samples; Discover the value scalar from the flow state; Train a pre-configured model based on a value scalar.
[0030] In order to accurately collect data from the platform, in one embodiment, the rules for publishing tasks and task scenarios are as follows: The information of the training data corresponding to the current value function is screened to determine the data group to be replaced; the screening is mainly based on the authenticity evaluation of the generation time and generation path of each data in the training data and the difference between the data and other data, and the screening is performed based on the evaluation results, and each screened data is placed in the data group to be replaced; Analyze each data of the data group to be replaced to determine the description parameter set of the target object, the target task and the task scenario corresponding to the target task; Based on the description parameter set, other users are screened to determine the target object; Send the target tasks and task scenarios to the target objects; When the target object does not respond within the preset time threshold, the screening of other users is repeated.
[0031] The data of the data group to be replaced are analyzed to determine the description parameter set of the target object, the target task and the task scenario corresponding to the target task, including: After extracting features from the data, the data is input into a pre-trained and converged neural network model to obtain the call factor; Based on the calling factor, the description parameter set of the target object, the target task and the task scenario corresponding to the target task are called from the preset target object and task analysis library.
[0032] Among them, based on the description parameter set, other users are screened to determine the target object, including: The description parameter set is matched with the classification parameter set of other users to obtain a matching result; the classification parameter set is constructed by analyzing the user information of other users, and the classification parameter set includes: parameters corresponding to information such as user age, gender, occupation, and location; the description parameter set is basically the same as the classification parameter set, and each data represents the parameters corresponding to information such as age, gender, occupation, and location of the target object; matching can be performed by calculating the similarity of two data sets, that is, the representation of the matching result can be the similarity between the two data sets; the similarity can be calculated by the cosine similarity calculation method; The matching result is evaluated using a pre-configured first scoring library to determine a first scoring value; the first scoring library is analyzed and configured in advance by professionals, and the first scoring value is associated with the matching result in a one-to-one correspondence in the first scoring library; The active data of other users are evaluated with a pre-configured second scoring library to determine a second scoring value; the specific evaluation may be performed by extracting features from the active data to extract active feature parameters; then the second scoring value corresponding to the active feature parameters is retrieved from the second scoring library through the active feature parameters; wherein the active feature parameters include: a parameter indicating the online time of a person, a parameter indicating the effective online time of a person (online with operations), a parameter indicating the number of times a person has participated in a solicitation activity, a parameter indicating the number of times feedback from a person who has participated in the solicitation activity is adopted, a parameter indicating the time of the most recent participation, etc.; Sort other users in descending order based on the weighted sum of the first score value and the second score value; the weighted coefficients of the first score value and the second score value are both pre-configured; The other users ranked first are targeted.
[0033] When reselecting the target object, the target object that has not been fed back needs to be deleted from the sorting; In addition, the filtering rules for filtering the information of the training data corresponding to the current value function are as follows: Extract features from the information of the training data to obtain data feature parameters; data feature parameters include: quantitative parameters representing the generation time, generation path, and differences between other data; the differences between data can be described by the similarity between data; Then, the data characteristic parameters are used to determine the corresponding evaluation value from the pre-configured screening library; The data whose evaluation value is less than or equal to the preset evaluation threshold value is used as the data to be replaced.
[0034] In order to ensure that the collected data is representative, when selecting the target object, multiple selections can be made in a top-down order, that is, N items are selected, where N is greater than or equal to 2; then the similarities of the data returned between the target objects are calculated, and the feedback data of the target object with the largest sum of similarities is used as the data to replace the data to be replaced; in order to achieve more accurate data replacement, the difference between the data to be replaced and the data used for replacement can also be considered, that is, when there are multiple feedback data, not only the difference between the feedback data but also the difference with the data to be replaced should be considered, and a first coefficient corresponding to the difference between the feedback data and a second coefficient corresponding to the difference with the data to be replaced can be configured; the target objects are sorted in descending order of the sum of the similarities between the feedback data and the first coefficient, and the similarity with the data to be replaced and the second coefficient, and the data corresponding to the first place in the order is used to replace the data to be replaced.
[0035] The present invention also provides a value function continuous learning system based on online data, such as Figure 2 As shown, it includes: an inference generation module 1, a publishing and collection module 2 and a retraining module 3; wherein, the inference generation module 1 infers and generates the expected value based on the current learning value and further determines the corresponding flow state, and uses the generation model to generate the first data sample corresponding to the flow state; the publishing and collection module 2 publishes tasks and task scenarios, receives other users' operations on tasks to generate corresponding scenario data; processes the scenario data to obtain a second data sample; the retraining module 3 adds the first data sample and the second data sample to the original training data sample, and after processing the training data sample, retrains the value function.
[0036] Among them, the steps of training the value function in the retraining module 3 are as follows: Extracting flow patterns from data samples; Discover the value scalar from the flow state; Train a pre-configured model based on a value scalar.
[0037] The rules for publishing tasks and task scenarios are as follows: Filter the information of the training data corresponding to the current value function to determine the data group to be replaced; Analyze each data of the data group to be replaced to determine the description parameter set of the target object, the target task and the task scenario corresponding to the target task; Based on the description parameter set, other users are screened to determine the target object; Send the target tasks and task scenarios to the target objects; When the target object does not respond within the preset time threshold, the screening of other users is repeated.
[0038] The data of the data group to be replaced are analyzed to determine the description parameter set of the target object, the target task and the task scenario corresponding to the target task, including: After extracting features from the data, the data is input into a pre-trained and converged neural network model to obtain the call factor; Based on the calling factor, the description parameter set of the target object, the target task and the task scenario corresponding to the target task are called from the preset target object and task analysis library.
[0039] Among them, based on the description parameter set, other users are screened to determine the target object, including: Matching the description parameter set with the classification parameter sets of other users to obtain a matching result; Evaluate the matching result using a pre-configured first scoring library to determine a first scoring value; Evaluate the activity data of other users using a pre-configured second scoring library to determine a second scoring value; Sorting other users in descending order based on the weighted sum of the first rating value and the second rating value; The other users ranked first are targeted.
[0040] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A method for continuous learning of value functions based on online data, characterized in that: include: Based on the current learning value, the expected value is inferred and the corresponding flow state is further determined, and the first data sample corresponding to the flow state is generated using the generation model; Publish tasks and task scenarios, receive other users' operations on tasks and generate corresponding scenario data; Processing the scene data to obtain a second data sample; The first data sample and the second data sample are added to the original training data sample, and after the training data sample is processed, the value function is retrained.
2. The method for continuous learning of value functions based on online data according to claim 1, characterized in that: The steps to train the value function are as follows: Extracting flow patterns from data samples; Discover the value scalar from the flow state; Train a pre-configured model based on a value scalar.
3. The method for continuous learning of value functions based on online data according to claim 1, characterized in that: The rules for publishing tasks and task scenarios are as follows: Filter the information of the training data corresponding to the current value function to determine the data group to be replaced; Analyze each data of the data group to be replaced to determine the description parameter set of the target object, the target task and the task scenario corresponding to the target task; Based on the description parameter set, other users are screened to determine the target object; Send the target tasks and task scenarios to the target objects; When the target object does not respond within the preset time threshold, the screening of other users is repeated.
4. The method for continuous learning of value functions based on online data as claimed in claim 3, characterized in that: Analyze each data of the data group to be replaced to determine the description parameter set of the target object, the target task, and the task scenario corresponding to the target task, including: After extracting features from the data, the data is input into a pre-trained and converged neural network model to obtain the call factor; Based on the calling factor, the description parameter set of the target object, the target task and the task scenario corresponding to the target task are called from the preset target object and task analysis library.
5. The method for continuous learning of value functions based on online data as claimed in claim 3, characterized in that: Based on the description parameter set, other users are screened to determine the target objects, including: Matching the description parameter set with the classification parameter sets of other users to obtain a matching result; Evaluate the matching result using a pre-configured first scoring library to determine a first scoring value; Evaluate the activity data of other users using a pre-configured second scoring library to determine a second scoring value; Sorting other users in descending order based on the weighted sum of the first rating value and the second rating value; The other users ranked first are targeted.
6. A value function continuous learning system based on online data, characterized in that: include: Inference generation module, publishing and collection module and retraining module; wherein, the inference generation module infers and generates expected value based on the current learning value and further determines the corresponding flow state, and uses the generation model to generate the first data sample corresponding to the flow state; the publishing and collection module publishes tasks and task scenarios, receives other users' operations on tasks to generate corresponding scenario data; processes the scenario data to obtain a second data sample; the retraining module adds the first data sample and the second data sample to the original training data sample, and after processing the training data sample, retrains the value function.
7. The value function continuous learning system based on online data as claimed in claim 6, characterized in that: The steps for the retraining module to train the value function are as follows: Extracting flow patterns from data samples; Discover the value scalar from the flow state; Train a pre-configured model based on a value scalar.
8. The value function continuous learning system based on online data as claimed in claim 6, characterized in that: The rules for publishing tasks and task scenarios are as follows: Filter the information of the training data corresponding to the current value function to determine the data group to be replaced; Analyze each data of the data group to be replaced to determine the description parameter set of the target object, the target task and the task scenario corresponding to the target task; Based on the description parameter set, other users are screened to determine the target object; Send the target tasks and task scenarios to the target objects; When the target object does not respond within the preset time threshold, the screening of other users is repeated.
9. The value function continuous learning system based on online data as claimed in claim 8, characterized in that: Analyze each data of the data group to be replaced to determine the description parameter set of the target object, the target task, and the task scenario corresponding to the target task, including: After extracting features from the data, the data is input into a pre-trained and converged neural network model to obtain the call factor; Based on the calling factor, the description parameter set of the target object, the target task and the task scenario corresponding to the target task are called from the preset target object and task analysis library.
10. The value function continuous learning system based on online data as claimed in claim 8, characterized in that: Based on the description parameter set, other users are screened to determine the target objects, including: Matching the description parameter set with the classification parameter sets of other users to obtain a matching result; Evaluate the matching result using a pre-configured first scoring library to determine a first scoring value; Evaluate the activity data of other users using a pre-configured second scoring library to determine a second scoring value; Sorting other users in descending order based on the weighted sum of the first rating value and the second rating value; The other users ranked first are targeted.
Citation Information
Patent Citations
Model deployment method and device, equipment and storage medium
CN117010474A
Automatic driving decision-making method and device for reinforcement learning guided by human driving data, and medium
CN118261233A
End-to-end automatic driving control system and device based on human preference reinforcement learning
CN119018181A
Strategy exploration model training method and device, computer equipment and storage medium
CN119204152A
Deep reinforcement learning-based captioning with embedding reward
US10467274B1