Integrated data management system and method based on reinforcement learning
By using a data integrated management system based on reinforcement learning and employing classification agents and reward mechanisms, the problem of insufficient dynamic adaptability in existing metadata management systems is solved. This enables adaptive hierarchical data management, improving the system's dynamic adaptability and management efficiency.
Patent Information
- Application Number
- CN202511035985.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-26
- Publication Date
- 2025-11-11
AI Technical Summary
Existing metadata management systems are insufficient in terms of dynamic adaptability and struggle to meet rapidly changing business needs.
A data management system based on reinforcement learning is adopted, which classifies K data sources through K classification agents. By combining soft classification strategies and reward mechanisms, an adaptive classification style is formed, thereby achieving hierarchical management of data.
This enables data management that adapts to dynamic changes under the guidance of a unified classification strategy, thereby improving the system's dynamic adaptability and management efficiency.
Smart Images

Figure CN120929982A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a data integration management system and method based on reinforcement learning, belonging to the field of artificial intelligence technology. Background Technology
[0002] Chinese invention patent application CN118643071A discloses a metadata management system and method integrating a large model. The system includes an information indexing module, a data processing unit, a model scheduling module, and a large model. The data processing unit receives industry source data transmitted from the large model, processes the data to generate a structured metadata dataset, and processes the metadata into tables that are easy to analyze. The model scheduling module receives questions, transforms the questions into a preset vector form, and executes a query mechanism on the information indexing module to extract key information. After obtaining the key information from the information indexing module, the key information is used to distribute the data in the knowledge base of the data warehouse using the model scheduling module. The system analyzes and generates prompt word templates, and schedules the API interface of the large model service to generate metadata information. The large model is used to generate metadata information after understanding the intent and parsing the content of the prompt word templates, and returns the generated metadata information to the data processing unit so that the data processing unit can execute multiple rounds of interactive responses. Among them, the large model sets the performance standard parameters of the initial large model according to the new business requirements; and fixes the weight of the performance standard parameters in the initial large model, obtains the knowledge and representation capabilities in the initial large model, fine-tunes the initial large model to integrate the knowledge of the preset domain, and trains it to generate the target large model of the preset domain after obeying human instructions, so as to provide accurate metadata for the data processing unit to execute interactive responses.
[0003] Although the technical solution disclosed in this invention patent application achieves a low-cost, high-quality, and high-efficiency real-time metadata management method, its dynamic adaptability is poor. Summary of the Invention
[0004] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide a data comprehensive management system and method based on reinforcement learning. This system uses K classification agents to classify K data sources, and manages the data in a hierarchical manner. It has both the guidance of a unified classification strategy and the ability to form its own classification style to adapt to dynamic changes.
[0005] To achieve the aforementioned objective, this invention provides a data integration management system based on reinforcement learning, comprising a centralized processor and K distributed classification agents. The centralized processor includes a rule distribution module and a data integration management module. The rule distribution module is configured to distribute soft classification strategies for each layer of hierarchical classification data to the K classification agents. The data integration management module is configured to classify the Kth classification agent according to the soft classification strategies. Classification results of the layers Evaluate and reward , ; L and K are both positive integers greater than or equal to 2; the k-th classifying agent is classified according to the reward. Obtain the par value portion of the return, obtain the total return based on the par value portion of the return, and obtain the target classification strategy based on the total return and the soft classification strategy.
[0006] To achieve the aforementioned objective, this invention also provides a data integration and management method based on reinforcement learning, which includes the following steps: Step 1: Distribute soft classification strategies for each layer and type of data in the hierarchical classification data to K classification agents through the rule distribution module of the centralized processor; Step 2: The k-th classification agent classifies and labels the data under its jurisdiction according to the soft classification strategy, and sends the classification results to the central processor; Step 3: Using the centralized processor data management module, classify the k-th classifying agent according to a soft classification strategy. Classification results of the layers Evaluate and reward , ; Both L and K are positive integers greater than or equal to 2. Step 4: The k-th classifying agent classifies according to the reward. Obtain the par value portion of the return, obtain the total return based on the par value portion of the return, and obtain the target classification strategy based on the total return and the soft classification strategy.
[0007] Compared with existing technologies, the data integrated management system and method based on reinforcement learning provided by this invention classifies K data sources by K classification agents, and performs hierarchical management of the data. It has the advantages of having a unified classification strategy and being able to form its own classification style, thus adapting to dynamic changes. Attached Figure Description
[0008] Figure 1 This is a block diagram of the data integration management system based on reinforcement learning provided in the first embodiment of the present invention. Detailed Implementation
[0009] To make the technical problems to be solved, the technical solutions, and the beneficial effects of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the present invention and are not intended to limit the present invention.
[0010] First Embodiment
[0011] Figure 1 This is a block diagram of the data integration management system based on reinforcement learning provided in the first embodiment of the present invention, as shown below. Figure 1 As shown, the first embodiment of the present invention provides a data integration management system based on reinforcement learning, which includes a centralized processor and K distributed classification agents. The centralized processor includes a rule distribution module and a data integration management module. The rule distribution module is configured to distribute soft classification strategies for each layer of hierarchical classification data to the K classification agents. The data integration management module is configured to classify the Kth classification agent according to the soft classification strategies. Classification results of the layers Evaluate and reward , ; L and K are both positive integers greater than or equal to 2; the k-th classifying agent is classified according to the reward. The process involves obtaining the par value portion of the return, then calculating the full return based on the par value portion, and finally determining the target classification strategy based on the full return and a soft classification strategy. In this invention, a positive reward is given if the classification result is both correct in category and labeling; otherwise, a negative reward is given.
[0012] In the first embodiment, the k-th classification agent performs the following process: S1-1: Based on the soft classification strategy Classify the k-th data source in the (l-1)-th layer of the data management system to obtain a column of data. In the formula, To label the classification result of the k-th data source in layer (l-1), in this invention, when l is 1, the k-th classification agent classifies the k-th data source, and its classification result is: .
[0013] S1-2: Obtain the field of view according to the following formula Par value partial return : In the formula, For the horizon, ; S1-3: Obtain the full return using the following formula : , In the formula, For the kth Discount factor for layered data sources; For the par value portion of the return at point T; S1-4: Update the following formula Layered returns The sum of corresponding weights: In the formula For weights; S1-5: Update the following formula The value function of the k-th state of the layer data source: ; S1-6: Generate the target classification strategy for the l-th layer according to the following formula: ; S1-7: Judgment: If Then output Otherwise, update the weights using the following formula: And return to step S1-2, For the first The kth importance level is sampled in the layer.
[0014] In the first embodiment, the k-th classification agent also performs the following process: S1-8: Classification Strategy Based on Target Classify the k-th data source in layer (l-1) to obtain class J data. J is a positive integer greater than or equal to 2, j = 1, ..., J. That is, the classification result... From J-class numbers to J-class data
[0015] In the first embodiment, the k-th classification agent also performs the following process: S1-9: Based on the soft classification strategy For the j-th type of data in layer l The data is categorized to obtain a column of data. In the formula, Label the classification result of the j-th data source in the l-th layer; S1-10: Obtained according to the following formula Par value partial return : In the formula, For the second horizon, ; S1-11: Obtain the full return using the following formula : , In the formula, For the first The j-th discount factor in the layer; To Affordable returns at the location; S1-12: Update the return using the following formula The sum of corresponding weights: In the formula For weights; S1-13: Update the value function of the j-th state in the (l+1)-th layer using the following formula: ; S1-14: Generate the target classification strategy for the (l+1)th layer according to the following formula: S1-15: Judgment: If Then output Otherwise, update the weights using the following formula: And return to steps S1-9, where, For the first The j-th importance of the layer is sampled.
[0016] In the first embodiment, the k-th classification agent also performs the following process: S1-16: Target Classification Strategy For the The j-th class of data in the layer is classified into P classes, where P is a positive integer greater than or equal to 2.
[0017] And so on, we can... The system categorizes data layer by layer, where r is a positive integer greater than or equal to 1. Each classification agent is guided by a unified classification strategy while also developing its own classification style. This allows for hierarchical management of data classification without the need for labeled data, and it adapts to dynamic changes.
[0018] In the first embodiment, the data integration management module implements the following process: S2-1: Based on the classification results of n similar data sources from K classification agents, form a data set. For each data set The data is fused to obtain the fusion result. , data fusion results The evaluation results were obtained. ; S2-2: Obtained from the following formula Evaluation value: , In the formula, These are the parameters of the evaluator D.
[0019] S2-3: Based on the evaluation value Generate data fusion strategy ,in, , In the formula, It's a hyperparameter; It is a data fusion strategy The optimization variables of the neural network These are the parameters after the data fusion strategy has been updated; These are the current parameters of the data fusion strategy; S2-4: Evaluate the results of data fusion based on the data fusion strategy to obtain sequence evaluation: ; For the m-th data group s m The fusion result obtained by performing data fusion; To analyze the fusion results The evaluation results; S2-5: Update the evaluator D according to the following formula. : , In the formula, The learning coefficient, To The gradient.
[0020] The first embodiment, through the above technical solution, can autonomously merge the classification results provided by n similar data sources, evaluate them through an intelligent evaluator, and adapt to the beneficial effects of dynamic changes.
[0021] The reinforcement learning-based data management system provided in the first embodiment also includes a data export module, which is configured to encrypt the classification result data group of the data source provided by the data management module, and then export it to local storage or cloud storage.
[0022] The data management system based on reinforcement learning provided in the first embodiment also includes a visualization query and analysis module, which is configured to query the required data group from local storage or cloud storage according to the user's instructions and display it on the display.
[0023] Second Embodiment
[0024] The second embodiment of the present invention only describes the contents that are different from those of the first embodiment; the contents that are the same will not be described again.
[0025] The data integration and management method based on reinforcement learning provided in the second embodiment of the present invention includes the following steps: Step 1: Distribute soft classification strategies for each layer and type of data in the hierarchical classification data to K classification agents through the rule distribution module of the centralized processor; Step 2: The k-th classification agent classifies and labels the data under its jurisdiction according to the soft classification strategy, and sends the classification results to the central processor; Step 3: Using the centralized processor data management module, classify the k-th classifying agent according to a soft classification strategy. Classification results of the layers Evaluate and reward , ; Both L and K are positive integers greater than or equal to 2. Step 4: The k-th classifying agent classifies according to the reward. Obtain the par value portion of the return, obtain the total return based on the par value portion of the return, and obtain the target classification strategy based on the total return and the soft classification strategy.
[0026] The beneficial effects of the second embodiment of the present invention are the same as those of the first embodiment, and will not be described again here.
[0027] Third Embodiment
[0028] To achieve the aforementioned objective, the present invention also provides a computer program product, characterized in that it includes computer program code, which is capable of being invoked by a processor to execute the method described in the second embodiment.
[0029] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified. "Several" means one or more, unless otherwise explicitly specified.
[0030] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. A data integration management system based on reinforcement learning, characterized in that, It includes a centralized processor and K distributed classification agents. The centralized processor includes a rule distribution module and a data integration and management module. The rule distribution module is configured to distribute soft classification strategies for each layer of hierarchical classification data to the K classification agents. The data integration and management module is configured to classify the Kth classification agent according to the soft classification strategies. Classification results of the layers Evaluate and reward , ; L and K are both positive integers greater than or equal to 2; the k-th classifying agent is classified according to the reward. Obtain the par value portion of the return, obtain the total return based on the par value portion of the return, and obtain the target classification strategy based on the total return and the soft classification strategy.
2. The data integration management system based on reinforcement learning according to claim 1, characterized in that, The k-th classifying agent performs the following process: S1-1: Based on the soft classification strategy Classify the k-th data source in the (l-1)-th layer of the data management system to obtain a column of data. In the formula, To classify and label the k-th data source in layer (l-1); S1-2: Obtain the field of view according to the following formula Par value partial return : In the formula, For the horizon, ; S1-3: Obtain the full return using the following formula : , In the formula, For the kth Discount factor for layered data sources; For the par value portion of the return at point T; S1-4: Update the following formula Layered returns The sum of corresponding weights: In the formula For weights; S1-5: Update the following formula The value function of the k-th state of the layer data source: ; S1-6: Generate the target classification strategy for the l-th layer according to the following formula: ; S1-7: Judgment: If Then output Otherwise, update the weights using the following formula: And return to step S1-2, For the first The kth importance level is sampled in the layer.
3. The data integration management system based on reinforcement learning according to claim 2, characterized in that, The k-th classifying agent also performs the following process: S1-8: Classification Strategy Based on Target Classify the k-th data source in layer (l-1) to obtain class J dataset. J is a positive integer greater than or equal to 2, where j = 1, ..., J.
4. The data integration management system based on reinforcement learning according to claim 3, characterized in that, The k-th classifying agent also performs the following process: S1-9: Based on the soft classification strategy For the j-th class dataset of layer l The data is categorized to obtain a column of data. In the formula, Label the classification results of the j-th data source in the l-th layer; S1-10: Obtained according to the following formula Par value partial return : In the formula, For the second horizon, ; S1-11: Obtain the full return using the following formula : , In the formula, For the first The j-th discount factor in the layer; To Affordable returns at the location; S1-12: Update the return using the following formula The sum of corresponding weights: In the formula For weights; S1-13: Update the value function of the j-th state in the (l+1)-th layer using the following formula: ; S1-14: Generate the target classification strategy for the (l+1)th layer according to the following formula: S1-15: Judgment: If Then output Otherwise, update the weights using the following formula: And return to steps S1-9, where, For the first The j-th importance of the layer is sampled.
5. The data integration management system based on reinforcement learning according to claim 4, characterized in that, The k-th classifying agent also performs the following process: S1-16: Target Classification Strategy For the first The j-th class of data in the layer is classified into P classes, where P is a positive integer greater than or equal to 2.
6. A data integration and management method based on reinforcement learning, characterized in that, Includes the following steps: Step 1: Distribute soft classification strategies for each layer and type of data in the hierarchical classification data to K classification agents through the rule distribution module of the centralized processor; Step 2: The k-th classification agent classifies and labels the data under its jurisdiction according to the soft classification strategy, and sends the classification results to the central processor; Step 3: Using the centralized processor data management module, classify the k-th classifying agent according to a soft classification strategy. Classification results of the layers Evaluate and reward , ; Both L and K are positive integers greater than or equal to 2. Step 4: The k-th classifying agent classifies according to the reward. Obtain the par value portion of the return, obtain the total return based on the par value portion of the return, and obtain the target classification strategy based on the total return and the soft classification strategy.
7. The data integration and management method based on reinforcement learning according to claim 6, characterized in that, The k-th classifying agent performs the following process: S1-1: Based on the soft classification strategy Classify the k-th data source in the (l-1)-th layer of the data management system to obtain a column of data. In the formula, To label the k-th data source in layer (l-1); S1-2: Obtain the field of view according to the following formula Par value partial return : In the formula, For the horizon, ; S1-3: Obtain the full return using the following formula : , In the formula, For the kth Discount factor for layered data sources; For the par value portion of the return at point T; S1-4: Update the following formula Layered returns The sum of corresponding weights: In the formula For weights; S1-5: Update the following formula The value function of the k-th state of the layer data source: ; S1-6: Generate the target classification strategy for the l-th layer according to the following formula: , S1-7: Judgment: If Then output Otherwise, update the weights using the following formula: And return to step S1-2, For the first The kth importance level is sampled in the layer.
8. The data integration and management method based on reinforcement learning according to claim 7, characterized in that, The k-th classifying agent also performs the following process: S1-8: Classification Strategy Based on Target Classify the k-th data source in layer (l-1) to obtain class J data. J is a positive integer greater than or equal to 2, j=1,...,J.
9. The data integration and management method based on reinforcement learning according to claim 3, characterized in that, The k-th classifying agent also performs the following process: S1-9: Based on the soft classification strategy For the j-th type of data in layer l The data is categorized to obtain a column of data. In the formula, For the j-th type of data, the first... Label the hierarchical classification results; S1-10: Obtained according to the following formula Par value partial return : In the formula, For the second horizon, ; S1-11: Obtain the full return using the following formula : , In the formula, For the first The j-th discount factor in the layer; To Affordable returns at the location; S1-12: Update the return using the following formula The sum of corresponding weights: In the formula For weights; S1-13: Update the value function of the j-th state in the (l+1)-th layer using the following formula: ; S1-14: Generate the target classification strategy for the (l+1)th layer according to the following formula: , S1-15: Judgment: If Then output Otherwise, update the weights using the following formula: And return to steps S1-9, where, For the first The j-th importance of the layer is sampled.
10. The data integration and management method based on reinforcement learning according to claim 9, characterized in that, The k-th classifying agent also performs the following process: S1-16: Target Classification Strategy Classify the j-th data source in the l-th layer to obtain P types of data, where P is a positive integer greater than or equal to 2.
Citation Information
Patent Citations
Metadata management system and method of integrated large model
CN118643071A