An online data-based value function continual learning method and system
By collecting data online and using cloud platform task interfaces, flow state latent variables and scene data are generated, which solves the shortcomings of human value function modeling in existing technologies, realizes continuous learning of value functions and scalability of multimodal data, and improves the adaptability and accuracy of artificial intelligence systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE
- Filing Date
- 2025-04-16
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies struggle to achieve effective, numerical modeling and continuous learning of human value functions, especially in complex, multi-dimensional human value modeling, where there is a lack of continuous data updates and adaptive adjustments.
By employing a value function continuous learning method based on online data, we utilize generative models to generate flow-state latent variables and data samples, combine cloud platform task interfaces to collect scenario data, filter and evaluate target objects, and continuously train the value function.
It enables continuous updating and adaptive adjustment of the value function, supports scalable learning of multimodal data, and enhances the value alignment capability of artificial intelligence systems.
Smart Images

Figure CN119989068B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and system for continuous learning of value functions based on online data. Background Technology
[0002] In the fields of psychology and cognitive science, the definition and classification of human values have formed relatively mature theoretical systems. For example, Maslow's Hierarchy of Needs divides human needs into five levels, arguing that needs progress progressively: from physiological needs to safety needs, social needs, esteem needs, and finally, the highest need for self-actualization. This theory shows that human values have a hierarchical nature in terms of the gradual satisfaction of different needs. Schwartz's Theory of Basic Values, on the other hand, proposes ten basic values from a social psychological perspective, covering multiple dimensions from power, achievement, and pleasure to philanthropy and universalism, emphasizing the universality and cross-cultural applicability of values. These theories provide a theoretical basis for understanding human behavioral motivations, but their frameworks are essentially qualitative descriptions, lacking explicit numerical modeling methods.
[0003] In machine learning and artificial intelligence research, the computation of value functions is often associated with reinforcement learning. In reinforcement learning, the value function is used to estimate the future reward of an agent in a given state. However, existing value function learning mainly focuses on the optimization of specific tasks, such as expressing multi-objective values under different tasks through generalized value functions (GVFs). While these methods enhance task flexibility, they are usually only effective in specific scenarios and are difficult to generalize to the modeling of universal human value functions.
[0004] On the other hand, value alignment has become a hot topic in the field of artificial intelligence in recent years, aiming to align the decision-making process of AI systems with human values. The main methods for achieving value alignment include human-in-the-loop reinforcement learning and preference learning. Human-in-the-loop reinforcement learning adjusts AI behavior through direct human feedback, while preference learning gradually internalizes human values by observing human choices. These methods rely on example data to attempt to align AI output with human expectations. However, current value alignment research focuses more on adaptive adjustments for specific tasks or scenarios, lacking a clear numerical value function, making it difficult to support complex, multi-dimensional human value modeling. The key to achieving continuous learning of the value function lies in data collection and updating; effectively collecting relevant data is a crucial core aspect of continuous value function learning. Summary of the Invention
[0005] One of the objectives of this invention is to provide a method and system for continuous learning of value functions based on online data, which enables effective and continuous collection of online data, ensuring the continuous learning and updating of value functions, thereby guaranteeing the effective updating of artificial intelligence systems.
[0006] This invention provides a method for continuous learning of value functions based on online data, comprising:
[0007] Based on the value of the current learning, the expected value is inferred and the corresponding flow state is further determined. The generative model is then used to generate the first data sample corresponding to the flow state.
[0008] Publish tasks and task scenarios, receive other users' operations on the tasks to generate corresponding scenario data; process the scenario data to obtain a second data sample;
[0009] The first and second data samples are added to the original training data samples, and after processing the training data samples, the value function is retrained.
[0010] Preferably, the steps for training the value function are as follows:
[0011] Extract flow patterns from data samples;
[0012] Extracting value scalars from the flow state;
[0013] The pre-configured model is trained based on the value scalar.
[0014] Preferably, the rules for publishing tasks and task scenarios are as follows:
[0015] Filter the information of the training data corresponding to the current value function to determine the data group to be replaced;
[0016] Analyze each piece of data in the data group to be replaced to determine the set of descriptive parameters of the target object, the target task, and the task scenario corresponding to the target task.
[0017] Based on the set of description parameters, other users are filtered to determine the target object;
[0018] Send the target task and task scenario to the target object;
[0019] If the target user does not respond within the preset time threshold, the process of filtering other users will be repeated.
[0020] Preferably, each piece of data in the data group to be replaced is analyzed to determine the descriptive parameter set of the target object, the target task, and the task scenario corresponding to the target task, including:
[0021] After feature extraction from the data, the input is fed into a pre-trained and converged neural network model to obtain the call factor;
[0022] Based on the retrieval factors, the description parameter set of the target object, the target task, and the corresponding task scenario are retrieved from the preset target object and task analysis library.
[0023] Preferably, based on the description parameter set, other users are filtered to determine the target object, including:
[0024] The description parameter set is matched with the classification parameter sets of other users to obtain the matching results;
[0025] The matching results are evaluated using a pre-configured first scoring library to determine a first scoring value;
[0026] The activity data of other users are evaluated using a pre-configured second rating library to determine a second rating value;
[0027] Other users are sorted based on the weighted sum of the first and second ratings, from largest to smallest.
[0028] The target is the other user ranked first.
[0029] This invention also provides a value function continuous learning system based on online data, comprising: an inference generation module, a publishing and acquisition module, and a retraining module; wherein, the inference generation module infers and generates a desired value based on the currently learned value and further determines the corresponding flow state, and uses a generative model to generate a first data sample corresponding to the flow state; the publishing and acquisition module publishes tasks and task scenarios, receives other users' operations on the tasks to generate corresponding scenario data; processes the scenario data to obtain a second data sample; the retraining module adds the first and second data samples to the original training data samples, processes the training data samples, and retrains the value function.
[0030] Preferably, the retraining module trains the value function using the following steps:
[0031] Extract flow patterns from data samples;
[0032] Extracting value scalars from the flow state;
[0033] The pre-configured model is trained based on the value scalar.
[0034] Preferably, the rules for publishing tasks and task scenarios are as follows:
[0035] Filter the information of the training data corresponding to the current value function to determine the data group to be replaced;
[0036] Analyze each piece of data in the data group to be replaced to determine the set of descriptive parameters of the target object, the target task, and the task scenario corresponding to the target task.
[0037] Based on the set of description parameters, other users are filtered to determine the target object;
[0038] Send the target task and task scenario to the target object;
[0039] If the target user does not respond within the preset time threshold, the process of filtering other users will be repeated.
[0040] Preferably, each piece of data in the data group to be replaced is analyzed to determine the descriptive parameter set of the target object, the target task, and the task scenario corresponding to the target task, including:
[0041] After feature extraction from the data, the input is fed into a pre-trained and converged neural network model to obtain the call factor;
[0042] Based on the retrieval factors, the description parameter set of the target object, the target task, and the corresponding task scenario are retrieved from the preset target object and task analysis library.
[0043] Preferably, based on the description parameter set, other users are filtered to determine the target object, including:
[0044] The description parameter set is matched with the classification parameter sets of other users to obtain the matching results;
[0045] The matching results are evaluated using a pre-configured first scoring library to determine a first scoring value;
[0046] The activity data of other users are evaluated using a pre-configured second rating library to determine a second rating value;
[0047] Other users are sorted based on the weighted sum of the first and second ratings, from largest to smallest.
[0048] The target is the other user ranked first.
[0049] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings.
[0050] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0051] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0052] Figure 1 This is a schematic diagram of a value function continuous learning method based on online data in an embodiment of the present invention;
[0053] Figure 2 This is a schematic diagram of a value function continuous learning system based on online data in an embodiment of the present invention. Detailed Implementation
[0054] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0055] This invention provides a method for continuous learning of value functions based on online data, such as... Figure 1 As shown, it includes:
[0056] Step 1: Based on the value of the current learning, infer the expected value and further determine the corresponding flow state. Use the generative model to generate the first data sample corresponding to the flow state.
[0057] Step 2: Publish the task and task scenario, receive other users' operations on the task to generate corresponding scenario data; process the scenario data to obtain a second data sample;
[0058] Step 3: Add the first and second data samples to the original training data samples, process the training data samples, and then retrain the value function.
[0059] The working principle and beneficial effects of the above technical solution are as follows:
[0060] This invention is based on continuous learning of the value function using online data. First, it's necessary to understand the value function; its specific generation process is as follows: Data Sample: This is raw data collected from the environment or task, which may be images, audio, video, text, or structured data. Fluent / Latent Variable: These are the underlying dynamic features or hidden patterns behind the data, reflecting changes in the data under different states. Latent variables represent a certain latent attribute or state of a data sample. Value Scalar: This is a numerical value representing the value of a data sample, usually obtained through the model's output, representing the quality of the target state or behavior corresponding to that sample. The steps for generating latent variables and sample data are as follows: For sequence data (e.g., time series, text), we use RNN or Transformer, where latent variables serve as input to the network's hidden layers, capturing dependencies and state transitions in the time series. For image data, we use CNN or UNet, where latent variables serve as additional feature encoding, helping to extract key features from the image and generate data samples. For graph data (such as social networks or transportation networks), GNNs (Graph Neural Networks) can be used to process the relationships between nodes in the data through the graph structure, thereby inferring latent variables and sample data. These different architectures can generate latent variables that fit different modalities of data, helping to establish accurate value functions.
[0061] The specific steps of online data acquisition are as follows:
[0062] I. Spontaneous Exploration by Intelligent Agents (Based on Value Inference): The intelligent agent infers the desired value objective based on the currently learned value system. To achieve this, the agent first identifies possible behavioral objectives in the current state and evaluates the potential value of these objectives. Using the currently learned value function, the agent can infer the underlying flow states (latent variables) from the desired value. For example, in an autonomous driving application, the desired value might be "minimizing collision risk," and the agent infers the corresponding flow states (e.g., road safety, traffic density, etc.) based on the current road and traffic conditions. Using generative models (such as VAEs or GANs), the agent generates new data samples based on these flow states, such as simulated traffic scenarios or new driving decisions.
[0063] II. Cloud Platform Task Interface:
[0064] The cloud platform provides interfaces for task scenarios, allowing users to interact with the system through a web-based interface to complete specific tasks. The system will then generate corresponding scenario data based on the user's actions (such as user operation logs, task completion time sequences, etc.).
[0065] These data can be post-processed to obtain higher-level information. For example, raw data (such as operation sequences) can be transformed into structured graphs (such as task flowcharts or behavioral decision graphs) using a given interpretation algorithm.
[0066] Through these two methods, the system can continuously generate new data samples, providing rich raw data for the continuous learning of the value function.
[0067] After processing and labeling, the collected data, together with the original data, will form a new dataset. Using the new data, we first need to infer the latent variables, i.e., the flow state, from the sample data and values. Then, we use the inferred flow state to generate sample data and value scalars. Using the learning theory principle of latent variable generation model, the newly generated data and values should be as consistent as possible with the original data and values. Based on this, flow state variables and how flow state latent variables generate sample data and values will spontaneously emerge.
[0068] Given sample data, we can infer the latent variables of the flow state, and then calculate the value of the current sample data based on the latent variable-to-value generative model. This constitutes the calculation of value from sample data, which is dual to the generation of data based on latent variables in the spontaneous exploration of intelligent agents, where data is generated by inferring latent variables based on given values. In fact, any modality can be added to the value system, such as text, audio, video, and structural graphs. These modalities are all generated by the latent variables of the flow state. In practical applications, latent variables can be inferred based on any given modal data, and other modalities and values of the current modal data can be generated, which greatly increases the scalability of the value system. Adding modalities (text, audio, video, etc.): For given modal data (such as images, text, or audio), the data is first processed using existing neural network models (such as CNNs, RNNs, etc.) to extract key features. Then, using latent variable inference models (such as variational autoencoders, generative adversarial networks, etc.), we can infer the latent variables behind these data (such as object features in images, sentiment features in text). Then, using the latent variable-to-value generative model (such as MLPs or other neural networks), the value corresponding to each modal data is calculated. For example, for a piece of text, its sentiment value (positive, negative, neutral, etc.) can be inferred from its latent variables, and its value to user experience can be further calculated. This method can unify data from different modalities (such as text, images, videos, etc.) into a single value system for calculation and optimization, while greatly improving the system's scalability and supporting more data modalities and complex application scenarios.
[0069] During the continuous data acquisition and learning process, the interpretability theory of neural networks, such as sparse autoencoders, can be used to detect latent variables in the flow state and obtain the specific meaning of each basis vector in the latent variable space. This represents a certain attribute of the sample data. Based on the correspondence between this attribute and the value function, more interpretable and controllable value alignment, task planning, and data generation can be carried out.
[0070] To achieve retraining of the value function, in one embodiment, the steps for training the value function are as follows:
[0071] Extract flow patterns from data samples;
[0072] Extracting value scalars from the flow state;
[0073] The pre-configured model is trained based on the value scalar.
[0074] To ensure accurate data collection from the platform, in one embodiment, the rules for publishing tasks and task scenarios are as follows:
[0075] The information of the training data corresponding to the current value function is filtered to determine the data group to be replaced. The filtering mainly evaluates the authenticity of each data in the training data based on the generation time, generation method and differences with other data. The filtered data are then placed into the data group to be replaced.
[0076] Analyze each piece of data in the data group to be replaced to determine the set of descriptive parameters of the target object, the target task, and the task scenario corresponding to the target task.
[0077] Based on the set of description parameters, other users are filtered to determine the target object;
[0078] Send the target task and task scenario to the target object;
[0079] If the target user does not respond within the preset time threshold, the process of filtering other users will be repeated.
[0080] This involves analyzing each piece of data in the data group to be replaced to determine the description parameter set of the target object, the target task, and the corresponding task scenario, including:
[0081] After feature extraction from the data, the input is fed into a pre-trained and converged neural network model to obtain the call factor;
[0082] Based on the retrieval factors, the description parameter set of the target object, the target task, and the corresponding task scenario are retrieved from the preset target object and task analysis library.
[0083] This includes filtering other users based on the description parameter set to determine the target object, including:
[0084] The description parameter set is matched with the classification parameter sets of other users to obtain the matching results. The classification parameter set is constructed by analyzing the user information of other users and includes parameters representing information such as user age, gender, occupation, and location. The description parameter set is basically the same as the classification parameter set, with each data point representing parameters corresponding to information such as the target object's age, gender, occupation, and location. Matching can be done by calculating the similarity between the two datasets, that is, the matching result can be represented as the similarity between the two datasets. The similarity can be calculated using the cosine similarity method.
[0085] The matching results are evaluated using a pre-configured first scoring library to determine the first scoring value. The first scoring library is pre-analyzed and configured by professionals, and the first scoring value is associated with the matching result in a one-to-one correspondence.
[0086] The activity data of other users is evaluated using a pre-configured second scoring library to determine a second scoring value. Specifically, the evaluation can involve feature extraction of the activity data to extract activity feature parameters. Then, the second scoring value corresponding to the activity feature parameters is retrieved from the second scoring library. These activity feature parameters include: parameters representing the user's online time, parameters representing the user's effective online time (online with active engagement), parameters representing the number of times the user participated in the solicitation activity, parameters representing the number of times the user's feedback was adopted after participating in the solicitation activity, and parameters representing the time of the most recent participation, etc.
[0087] Other users are sorted from largest to smallest based on the weighted sum of the first and second ratings; the weighting coefficients for the first and second ratings are pre-configured.
[0088] The target is the other user ranked first.
[0089] When reselecting a target object, it is necessary to remove the target objects that have not provided feedback from the sorting.
[0090] Furthermore, the filtering rules for selecting information from the training data corresponding to the current value function are as follows:
[0091] Feature extraction is performed on the training data to obtain data feature parameters. These parameters include quantitative parameters representing the generation time, generation method, and differences from other data. Differences between data can be described by the similarity between data.
[0092] Then, the corresponding evaluation value is determined from a pre-configured screening library using data feature parameters;
[0093] Data whose evaluation value is less than or equal to the preset evaluation threshold will be used as the data to be replaced.
[0094] To ensure the representativeness of the collected data, multiple selections can be made when choosing target objects, based on a top-down sorting method, i.e., selecting N objects, where N is greater than or equal to 2. Then, the similarity between the returned data of each target object is calculated, and the feedback data of the target object with the highest sum of similarity is used as the data to replace the data to be replaced. To achieve more accurate data replacement, the differences between the data to be replaced and the data used for replacement can also be considered. That is, when there are multiple feedback data, not only the differences between the feedback data should be considered, but also the differences with the data to be replaced. A first coefficient corresponding to the difference between the feedback data and a second coefficient corresponding to the difference with the data to be replaced can be configured. The sum of the similarity between the feedback data and the product of the first coefficient, and the sum of the similarity with the data to be replaced and the second coefficient, are used to sort the target objects in descending order, and the data corresponding to the first position in the sorted list replaces the data to be replaced.
[0095] This invention also provides a value function continuous learning system based on online data, such as... Figure 2 As shown, it includes: an inference generation module 1, a publishing and acquisition module 2, and a retraining module 3; wherein, the inference generation module 1 infers and generates the expected value based on the value of the current learning and further determines the corresponding flow state, and uses the generative model to generate the first data sample corresponding to the flow state; the publishing and acquisition module 2 publishes the task and task scenario, receives other users' operations on the task to generate corresponding scenario data; processes the scenario data to obtain the second data sample; the retraining module 3 adds the first data sample and the second data sample to the original training data sample, and after processing the training data sample, retrains the value function.
[0096] The retraining module 3 trains the value function using the following steps:
[0097] Extract flow patterns from data samples;
[0098] Extracting value scalars from the flow state;
[0099] The pre-configured model is trained based on the value scalar.
[0100] The rules for publishing tasks and task scenarios are as follows:
[0101] Filter the information of the training data corresponding to the current value function to determine the data group to be replaced;
[0102] Analyze each piece of data in the data group to be replaced to determine the set of descriptive parameters of the target object, the target task, and the task scenario corresponding to the target task.
[0103] Based on the set of description parameters, other users are filtered to determine the target object;
[0104] Send the target task and task scenario to the target object;
[0105] If the target user does not respond within the preset time threshold, the process of filtering other users will be repeated.
[0106] This involves analyzing each piece of data in the data group to be replaced to determine the description parameter set of the target object, the target task, and the corresponding task scenario, including:
[0107] After feature extraction from the data, the input is fed into a pre-trained and converged neural network model to obtain the call factor;
[0108] Based on the retrieval factors, the description parameter set of the target object, the target task, and the corresponding task scenario are retrieved from the preset target object and task analysis library.
[0109] This includes filtering other users based on the description parameter set to determine the target object, including:
[0110] The description parameter set is matched with the classification parameter sets of other users to obtain the matching results;
[0111] The matching results are evaluated using a pre-configured first scoring library to determine a first scoring value;
[0112] The activity data of other users are evaluated using a pre-configured second rating library to determine a second rating value;
[0113] Other users are sorted based on the weighted sum of the first and second ratings, from largest to smallest.
[0114] The target is the other user ranked first.
[0115] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for continuous learning of value functions based on online data, characterized in that, include: Based on the value of the current learning, the expected value is inferred and the corresponding flow state is further determined. The generative model is then used to generate the first data sample corresponding to the flow state. Publish tasks and task scenarios, and receive other users' actions on the tasks to generate corresponding scenario data; The scene data is processed to obtain a second data sample; The first and second data samples are added to the original training data samples, and after processing the training data samples, the value function is retrained. The first data sample, the second data sample, and the original training data sample each include one or more combinations of images, audio, video, and text. The rules for publishing tasks and task scenarios are as follows: Filter the information of the training data corresponding to the current value function to determine the data group to be replaced; Analyze each piece of data in the data group to be replaced to determine the set of descriptive parameters of the target object, the target task, and the task scenario corresponding to the target task. Based on the set of description parameters, other users are filtered to determine the target object; Send the target task and task scenario to the target object; If the target user does not respond within the preset time threshold, the process of filtering other users will be repeated. The filtering rules for selecting information from the training data corresponding to the current value function are as follows: Feature extraction is performed on the training data to obtain data feature parameters; the data feature parameters include: quantitative parameters representing the generation time, generation method, and differences from other data; The corresponding evaluation value is determined from a pre-configured screening library using data feature parameters; Data whose evaluation value is less than or equal to a preset evaluation threshold will be used as data to be replaced; This includes filtering other users based on the description parameter set to determine the target object, including: The description parameter set is matched with the classification parameter sets of other users to obtain the matching results; The matching results are evaluated using a pre-configured first scoring library to determine a first scoring value; The activity data of other users are evaluated using a pre-configured second rating library to determine a second rating value; Other users are sorted based on the weighted sum of the first and second ratings, from largest to smallest. Select N other users as target objects based on the sorting order from top to bottom; Configure a first coefficient corresponding to the difference between the feedback data and a second coefficient corresponding to the difference between the feedback data and the data to be replaced; sort the target objects in descending order of the sum of the similarity between the feedback data and the product of the first coefficient, and the sum of the similarity between the feedback data and the data to be replaced and the product of the second coefficient; and replace the data to be replaced with the data corresponding to the first sorted data.
2. The value function continuous learning method based on online data as described in claim 1, characterized in that, The steps for training the value function are as follows: Extract flow patterns from data samples; Extracting value scalars from the flow state; The pre-configured model is trained based on the value scalar.
3. The value function continuous learning method based on online data as described in claim 1, characterized in that, Analyze each data point in the data group to be replaced to determine the descriptive parameter set of the target object, the target task, and the corresponding task scenario, including: After feature extraction from the data, the input is fed into a pre-trained and converged neural network model to obtain the call factor; Based on the retrieval factors, the system retrieves the description parameter set of the target object, the target task, and the corresponding task scenario from the preset target object and task analysis library.
4. A value function continuous learning system based on online data, characterized in that, include: The system comprises an inference generation module, a publishing and acquisition module, and a retraining module. The inference generation module infers the expected value based on the currently learned value and further determines the corresponding flow state. Using a generative model, it generates the first data sample corresponding to the flow state. The publishing and acquisition module publishes tasks and task scenarios, receives other users' operations on the tasks, and generates corresponding scenario data. It then processes the scenario data to obtain the second data sample. The retraining module adds the first and second data samples to the original training data samples, processes the training data samples, and retrains the value function. The first data sample, the second data sample, and the original training data sample each include one or more combinations of images, audio, video, and text. The rules for publishing tasks and task scenarios are as follows: Filter the information of the training data corresponding to the current value function to determine the data group to be replaced; Analyze each piece of data in the data group to be replaced to determine the set of descriptive parameters of the target object, the target task, and the task scenario corresponding to the target task. Based on the set of description parameters, other users are filtered to determine the target object; Send the target task and task scenario to the target object; If the target user does not respond within the preset time threshold, the process of filtering other users will be repeated. The filtering rules for selecting information from the training data corresponding to the current value function are as follows: Feature extraction is performed on the training data to obtain data feature parameters; the data feature parameters include: quantitative parameters representing the generation time, generation method, and differences from other data; The corresponding evaluation value is determined from a pre-configured screening library using data feature parameters; Data whose evaluation value is less than or equal to a preset evaluation threshold will be used as data to be replaced; This includes filtering other users based on the description parameter set to determine the target object, including: The description parameter set is matched with the classification parameter sets of other users to obtain the matching results; The matching results are evaluated using a pre-configured first scoring library to determine a first scoring value; The activity data of other users are evaluated using a pre-configured second rating library to determine a second rating value; Other users are sorted based on the weighted sum of the first and second ratings, from largest to smallest. After sorting, select the top N other users as the target objects; Configure a first coefficient corresponding to the difference between the feedback data and a second coefficient corresponding to the difference between the feedback data and the data to be replaced; sort the target objects in descending order of the sum of the similarity between the feedback data and the product of the first coefficient, and the sum of the similarity between the feedback data and the data to be replaced and the product of the second coefficient; and replace the data to be replaced with the data corresponding to the first sorted data.
5. The value function continuous learning system based on online data as described in claim 4, characterized in that, The retraining module trains the value function using the following steps: Extract flow patterns from data samples; Extracting value scalars from the flow state; The pre-configured model is trained based on the value scalar.
6. The value function continuous learning system based on online data as described in claim 4, characterized in that, Analyze each data point in the data group to be replaced to determine the descriptive parameter set of the target object, the target task, and the corresponding task scenario, including: After feature extraction from the data, the input is fed into a pre-trained and converged neural network model to obtain the call factor; Based on the retrieval factors, the system retrieves the description parameter set of the target object, the target task, and the corresponding task scenario from the preset target object and task analysis library.