Data management method, system and equipment based on artificial intelligence, medium and product
Through the data governance method based on artificial intelligence, identifying data source categories, collecting and adjusting data, and outputting the optimal data service strategy, the problem of lack of flexibility and adaptability of data governance methods in the existing technology is solved, and efficient and intelligent data governance is achieved.
Patent Information
- Application Number
- CN202510050712.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-06-06
AI Technical Summary
Existing data governance methods lack flexibility and adaptability, cannot adapt to complex and changeable business needs, and can only manage data in a single field.
Using artificial intelligence-based data governance methods, we use the category of data sources to collect raw data by identifying the corresponding data collection strategies, and use large language models to detect data quality and generate standards, adjust and store data. At the same time, monitor the data usage status, output the optimal data service strategy, and perform data analysis and display.
It has achieved automatic adaptation to the data characteristics of different fields, improved the flexibility and adaptability of data governance methods, significantly improved the efficiency of data governance, and met the needs of efficient data governance in large-scale and complex data environments.
Smart Images

Figure CN120104599A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data governance technology, and in particular to an artificial intelligence-based data governance method, system, device, medium and product. Background Art
[0002] With the rapid development of information technology and the advent of the big data era, all walks of life have accumulated massive amounts of structured and unstructured data. How to efficiently use and manage these massive amounts of data has become an important challenge facing all organizations.
[0003] Existing data governance methods are generally governance construction strategies developed based on the data characteristics of a specific field. They can only be applied to a single field or a single task. Once the business field is changed, the data governance strategy needs to be re-formulated.
[0004] It can be seen that the existing data governance methods cannot adapt to complex and changing business needs and lack flexibility and adaptability. Summary of the invention
[0005] The present invention provides an artificial intelligence-based data governance method, system, device, medium and product to overcome the defect that the data governance method in the prior art can only perform data governance on data in a single field, realize automatic adaptation to the data characteristics of different fields, and improve the flexibility and adaptability of the data governance method.
[0006] The present invention provides a data governance method based on artificial intelligence, comprising the following steps: Identifying a data source; and collecting raw data from the data source using a data collection strategy corresponding to the category of the data source according to the identified category of the data source; Inputting the original data into a large language model to obtain a data quality detection result output by the large language model; the data quality detection result includes at least one of data integrity, data accuracy and data consistency; the large language model is also used to generate a data standard according to the data quality detection result; The original data is adjusted according to the data standard, and the adjusted original data is stored in a database.
[0007] According to an artificial intelligence-based data governance method provided by the present invention, after adjusting the original data according to the data standard and storing the adjusted original data in the database, the method further includes: Monitor the data usage status of the data user terminal; Based on the data usage status, the proximal policy optimization model is used to output the optimal data service policy to the data user, so that the data user obtains the adjusted original data according to the optimal data service policy.
[0008] According to an artificial intelligence-based data governance method provided by the present invention, after adjusting the original data according to the data standard and storing the adjusted original data in the database, the method further includes: Inputting the adjusted original data into the large language model to obtain a data analysis result output by the large language model; The data analysis result is input into a multimodal pre-training model to obtain a data display image for the data analysis result output by the multimodal pre-training model.
[0009] According to an artificial intelligence-based data governance method provided by the present invention, the data usage status includes the access requirements of the data user for target data; the optimal data service strategy includes a security strategy; based on the data usage status, the optimal data service strategy is output to the data user using a proximal strategy optimization model so that the data user obtains the adjusted original data according to the optimal data service strategy, including: Based on the access requirements for the target data, a proximal policy optimization model is used to output a security policy so that the data user obtains the target data according to the security policy.
[0010] According to an artificial intelligence-based data governance method provided by the present invention, the large language model is also used to output a data sensitivity analysis result for the original data; before adjusting the original data according to the data standard and storing the adjusted original data in the database, it also includes: Desensitizing or encrypting the original data according to the data sensitivity analysis result to obtain secure data; The adjusting the original data according to the data standard and storing the adjusted original data into the database comprises: The security data is adjusted according to the data standard, and the adjusted security data is stored in the database.
[0011] The present invention also provides a data governance system based on artificial intelligence, comprising the following modules: A data collection agent, used to identify a data source; and according to the identified category of the data source, adopt a data collection strategy corresponding to the category of the data source to collect raw data from the data source; A data quality agent, used to input the original data into a large language model to obtain a data quality detection result output by the large language model; the data quality detection result includes at least one of data integrity, data accuracy and data consistency; The data standard agent is used to input the data quality detection result into the large language model to obtain the data standard output by the large language model; adjust the original data according to the data standard, and store the adjusted original data into the database.
[0012] According to an artificial intelligence-based data governance system provided by the present invention, the system also includes an API service agent; The API service agent is used to monitor the data usage status of the data user end; The API service agent is also used to output the optimal data service strategy to the data user end based on the data usage status using the proximal strategy optimization model, so that the data user end obtains the adjusted original data according to the optimal data service strategy.
[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, an artificial intelligence-based data governance method as described above is implemented.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the artificial intelligence-based data governance methods described above.
[0015] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the artificial intelligence-based data governance methods described above.
[0016] The data governance method, system, device, medium and product based on artificial intelligence provided by the present invention, through different data sources; according to the category of the identified data source, adopt the data collection strategy corresponding to the category of the data source to collect raw data from the data source; input the raw data into the large language model, and obtain the data quality detection result output by the large language model; the data quality detection result includes at least one of data integrity, data accuracy and data consistency; the large language model is also used to generate data standards according to the data quality detection results; adjust the raw data according to the data standards, and store the adjusted raw data in the database. The method can select the most suitable data collection strategy to collect raw data according to the characteristics of different data sources, and generate the quality detection results and data standards of the raw data through the large language model, adjust and store the raw data according to the data standards. Compared with the traditional data governance strategy that can only be applied to a single data source, the method can automatically adapt to different data sources, realize the intelligence and automation of data management, significantly improve the efficiency of data governance, and meet the needs of efficient data governance in large-scale and complex data environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 It is a flowchart of the artificial intelligence-based data governance method provided by the present invention.
[0019] Figure 2 It is a schematic diagram of the software architecture of the artificial intelligence-based data governance system provided by the present invention.
[0020] Figure 3 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0022] Combine the following Figure 1-Figure 3 Specific embodiments of the present invention are described.
[0023] Figure 1 It is a flow chart of the data governance method based on artificial intelligence provided by the present invention, such as Figure 1 As shown, the method includes the following.
[0024] Step 101, identifying a data source; and collecting raw data from the data source using a data collection strategy corresponding to the category of the data source according to the category of the identified data source.
[0025] The data source refers to the port that outputs the original data, which can be real-time data or offline data imported by the user.
[0026] Specifically, this application implements the data governance method based on artificial intelligence by building a data governance platform. The platform can be implemented with a server or server cluster, or it can run in the cloud and have input and output ports. The structural diagram of the platform in the software architecture is as follows Figure 2 As shown, the platform structure includes the basic layer, core layer, collaboration layer and service layer.
[0027] The basic layer consists of data collection agents, which are responsible for real-time or batch collection and preprocessing of data from one or more data sources. The data sent by multiple data sources can be the same or different types of data. The data collection agent uses the PPO (Proximal Policy Optimization) algorithm to optimize the data collection strategy, automatically adapt to changes in data sources, and improve the speed and accuracy of data collection. The PPO algorithm is a reinforcement learning algorithm that is mainly based on improvements in the policy gradient direction and mainly includes the following key steps.
[0028] (1) Collecting data: By executing the current policy in the environment, such as the data collection policy, a set of interaction data is collected. This data includes state, action, reward, and possible next state.
[0029] (2) Computing advantage estimates: In order to evaluate how good an action is relative to the average, we need to compute an advantage function. This is usually done through some form of temporal difference estimation or generalized advantage estimation.
[0030] (3) Optimizing the objective function: The PPO algorithm uses a specially designed objective function that involves the probability ratio of the old strategy.
[0031] (4) Update strategy: Use the gradient ascent method to update the strategy parameters in the above objective function.
[0032] (5) Repeat the above steps using the updated policy parameters until certain stopping criteria are met, such as when the policy performance no longer improves or a certain number of iterations has been reached.
[0033] Through the above steps, the data collection strategy that best suits the characteristics of the current data source can be obtained, that is, the data collection strategy corresponding to the data category of the data source can be obtained.
[0034] Step 102, input the original data into the large language model to obtain the data quality detection result output by the large language model; the data quality detection result includes at least one of data integrity, data accuracy and data consistency; the large language model is also used to generate data standards based on the data quality detection result.
[0035] Specifically, after the base layer collects the original data, it transmits the original data to the core layer; Figure 2 As shown, the core layer includes metadata agent, data quality agent, data standard agent and data security agent.
[0036] The core layer is mainly responsible for inputting the above raw data into the trained large language model, such as the Qwen-7B (Tongyi Qianwen) model, to obtain the data quality test results output by the trained large language model, including at least one of data integrity, data accuracy and data consistency. Each agent in the core layer will also combine the PPO algorithm to optimize the data management strategy, including optimizing the data integrity management strategy, data accuracy management strategy and data consistency management strategy.
[0037] In addition, the large language model that has been trained will also regularly maintain data standards, for example, using the Qwen-7B (Tongyi Qianwen) model to generate data standards based on the above data quality test results to ensure data consistency and correctness, and dynamically adjust data standards in combination with the PPO algorithm. Among them, the Qwen-7B (Tongyi Qianwen) model is used to generate and interpret data standards, while the PPO algorithm is used to optimize the application and maintenance of data standards.
[0038] Step 103, adjusting the original data according to the above data standard, and storing the adjusted original data into the database.
[0039] Specifically, the data standard agent in the core layer also processes the above raw data according to the above data standard, and stores the processed raw data into the database.
[0040] The above embodiment, by identifying the data source; according to the category of the identified data source, adopting the data collection strategy corresponding to the category of the data source to collect the original data from the data source; inputting the original data into the large language model, obtaining the data quality detection result output by the large language model; the data quality detection result includes at least one of data integrity, data accuracy and data consistency; the large language model is also used to generate data standards according to the data quality detection results; adjusting the original data according to the data standards, and storing the adjusted original data in the database. This method can select the most suitable data collection strategy to collect the original data according to the characteristics of different data sources, and generate the quality detection results and data standards of the original data through the large language model, adjust and store the original data according to the data standards. Compared with the traditional data governance strategy that can only be applied to a single data source, this method can automatically adapt to different data sources, realize the intelligence and automation of data management, significantly improve the efficiency of data governance, and meet the needs of efficient data governance in large-scale and complex data environments.
[0041] In one embodiment, after the above step 103, it also includes: monitoring the data usage status of the data user end; based on the data usage status, using the proximal strategy optimization model to output the optimal data service strategy to the data user end, so that the data user end obtains the adjusted original data according to the optimal data service strategy.
[0042] Specifically, combined Figure 2 As shown, the software architecture provided by this application also includes a service layer, which includes an AI service agent, which is mainly used to monitor the data usage status of the user end (i.e., the data resource subscriber, or the data user end), including data usage requests, data usage formats, etc. Based on the above data usage status, the PPO optimization model is used to optimize service recommendations and subscription strategies to provide personalized data services. The agent will also use a large language model that has been trained (such as the above Qwen-7B model), where the large language model is used to generate API (Application Programming Interface) documents and service descriptions.
[0043] The above-mentioned embodiment generates API documents and services through a large language model, and optimizes the service recommendation strategy for users using the PPO optimization model, so as to provide personalized data service strategies and further improve the intelligent data governance strategy.
[0044] In one embodiment, after the above step 103, it also includes: inputting the adjusted original data into the large language model to obtain the data analysis results output by the large language model; inputting the data analysis results into the multimodal pre-training model to obtain a data display image for the data analysis results output by the multimodal pre-training model.
[0045] Specifically, Figure 2 As shown, the service layer of the software architecture also includes: data visualization agent. The data visualization agent can generate visualization reports and dashboards based on the large language model that has been trained (such as the Qwen-7B model) and the multimodal pre-training model that has been trained (such as CLIP (Contrastive Language–Image Pre-training, multimodal pre-training model)), and supports multi-dimensional data analysis and display. For example, after the adjusted original data is stored in the database, the above-mentioned original data is analyzed by the above-mentioned large language model and multimodal pre-training model to obtain the image-text analysis results, and a visualization report is established based on this. Among them, the Qwen-7B model is used to generate data descriptions and annotations, and the CLIP model is used to combine image and text data for comprehensive analysis and display.
[0046] In addition, the data visualization agent also optimizes the visualization configuration process through the PPO algorithm, allowing users to generate personalized visualization reports through simple operations.
[0047] The above embodiment, by combining a large language model to generate descriptions and annotations of the data, and using a multimodal pre-trained model to generate a visual report based on the above descriptions and annotations, is conducive to providing vivid and intuitive data management reports to the management end and the user end, thereby facilitating the user and the management end to optimize the data content and further improve the efficiency of data governance.
[0048] In one embodiment, the data usage status includes the data user's access requirements for the target data; the optimal data service policy includes a security policy; the above-mentioned outputting the optimal data service policy to the data user using a proximal policy optimization model based on the data usage status so that the data user obtains the adjusted original data according to the optimal data service policy includes: based on the access requirements for the target data, outputting the security policy using a proximal policy optimization model so that the data user obtains the target data according to the security policy.
[0049] Specifically, combined Figure 2 As shown in the software architecture diagram, the core layer of the software architecture also includes: data security agent. During the data export process, when the user end proposes access requirements for target data, the data security agent will analyze the current security policy of the target data according to the access requirements, for example, using the PPO algorithm, combined with the trial risk assessment results, dynamically adjust the security policy for the target data, and output the security policy so that the data user end can obtain the above target data according to the currently output security policy.
[0050] The above-mentioned embodiment improves the security of data use by providing a security policy for accessing target data according to the access requirements of users.
[0051] In one embodiment, the above-mentioned large language model is also used to output a data sensitivity analysis result for the original data. Before the above-mentioned step 103, it also includes: desensitizing or encrypting the original data according to the data sensitivity analysis result to obtain secure data; the above-mentioned step 103 includes: adjusting the secure data according to the data standard, and storing the adjusted secure data into the database.
[0052] Specifically, the above-mentioned data security intelligent agent uses a large language model (such as the Qwen-7B model) to analyze data sensitivity and access requirements, generate data desensitization and encryption strategies, desensitize or encrypt the original data to obtain secure data, adjust the secure data according to data standards, and store the adjusted secure data in the database.
[0053] The above-mentioned embodiment can further ensure the security of data by using a large language model to generate a desensitization and encryption strategy for the original data during the data import process, thereby automatically encrypting and desensitizing the original data.
[0054] The artificial intelligence-based data governance system provided by the present invention is described below. The artificial intelligence-based data governance system described below and the artificial intelligence-based data governance method described above can be referenced to each other.
[0055] Combination Figure 2 As shown in the software architecture diagram of the artificial intelligence-based data governance system shown, the artificial intelligence-based data governance system includes.
[0056] A data collection agent, used to identify a data source; and according to the identified category of the data source, adopt a data collection strategy corresponding to the category of the data source to collect raw data from the data source; A data quality agent, used to input the original data into a large language model to obtain a data quality detection result output by the large language model; the data quality detection result includes at least one of data integrity, data accuracy and data consistency; The data standard agent is used to input the data quality detection result into the large language model to obtain the data standard output by the large language model; adjust the original data according to the data standard, and store the adjusted original data into the database.
[0057] Furthermore, the above system also includes an API service agent; the API service agent is used to monitor the data usage status of the data user end; the API service agent is also used to output the optimal data service strategy to the data user end based on the data usage status using a proximal strategy optimization model, so that the data user end obtains the adjusted original data according to the optimal data service strategy.
[0058] Specifically, Figure 2 As shown, the data collection agent mainly realizes the following functions.
[0059] 1. Data source identification.
[0060] 1.1. Identify data sources: The data collection agent identifies and connects to different data sources, such as databases, file systems, API interfaces, etc. By configuring the connection parameters of the data source (such as URL, user name, password, etc.), the agent can automatically adapt to various data sources.
[0061] 1.2. Intelligent collection strategy: Using the PPO algorithm, the data collection agent dynamically adjusts the data collection strategy and optimizes the data collection frequency, data format and preprocessing method.
[0062] 2. Data preprocessing.
[0063] 2.1. Data cleaning: The data collection agent cleans the imported data, including removing duplicate data, processing missing values and outliers, etc.
[0064] 2.2. Data conversion: According to predefined rules, data format conversion (such as CSV to JSON, XML to JSON, etc.) is performed to ensure data consistency and availability.
[0065] 2.3. Data storage: The cleaned and converted data is stored in big data storage systems such as Hadoop, Hive, HBase, etc.
[0066] This AI-based data governance system automatically optimizes data collection strategies through intelligent collection agents and PPO algorithms. It can adapt to a variety of data sources and data formats, significantly improving the efficiency and accuracy of data collection.
[0067] Optionally, the above-mentioned artificial intelligence-based data governance system also includes: a metadata agent; the functions that can be realized by the metadata agent are as follows.
[0068] 3. Metadata management.
[0069] 3.1. Metadata generation: The metadata agent uses the Qwen-7B model to generate metadata descriptions, including field names, data types, data sources, etc.
[0070] 3.2. Metadata maintenance: Through the PPO algorithm, the metadata agent automatically maintains and updates metadata to ensure the consistency and integrity of metadata.
[0071] 3.3. Metadata recommendation: Based on data usage and user needs, intelligently recommend relevant metadata to facilitate users to understand and use data.
[0072] This AI-based data governance system uses the Qwen-7B model to generate and maintain metadata standards, optimizes the metadata management process through the PPO algorithm, provides intelligent metadata analysis and recommendations, and ensures the consistency and integrity of metadata.
[0073] Furthermore, the functions that can be realized by the data quality agent in the artificial intelligence-based data governance system include:
[0074] 4. Data quality management.
[0075] 4.1. Quality monitoring: The data quality agent monitors data quality in real time and uses the Qwen-7B model to identify data quality issues (such as data integrity, accuracy, consistency, etc.).
[0076] 4.2. Generation of quality audit rules: Generate and optimize data quality audit rules through the Qwen-7B model and PPO algorithm.
[0077] 4.3. Automatic repair: The data quality agent provides repair suggestions based on the audit results and automatically performs repair operations to ensure data quality.
[0078] In this AI-based data governance system, the data quality agent uses the Qwen-7B model to identify data quality issues, generate and optimize data quality audit rules, and provide automatic repair suggestions through the PPO algorithm to achieve efficient data quality monitoring and repair, and improve data accuracy and reliability.
[0079] The above-mentioned artificial intelligence-based data governance system also includes: a data standard agent, which can realize the following functions.
[0080] 5. Data standard management.
[0081] 5.1. Standard definition: The data standard agent uses the Qwen-7B model to generate and define data standards, such as naming conventions, data formats, storage conventions, etc.
[0082] 5.2. Standard maintenance: Through the PPO algorithm, the data standard agent dynamically adjusts the data standard to ensure that the data standard meets business requirements.
[0083] 5.3. Consistency check: Automatically check the consistency of data standards, identify and resolve standard conflicts, and ensure the consistency and correctness of data.
[0084] This AI-based data governance system uses the Qwen-7B model to define and interpret data standards, and uses the PPO algorithm to dynamically adjust data standards to ensure that data standards meet changing business needs and improve data consistency and standardization.
[0085] like Figure 2 As shown, the above-mentioned artificial intelligence-based data governance system also includes: a data security intelligent agent, whose achievable functions include:
[0086] 6. Data security management.
[0087] 6.1. Security policy generation: The data security agent uses the Qwen-7B model to analyze data sensitivity and access requirements and generate data desensitization and encryption strategies.
[0088] 6.2. Security policy optimization: Through the PPO algorithm, the security policy is dynamically adjusted to ensure high data security based on real-time risk assessment results.
[0089] 6.3. Access control and auditing: The intelligent agent monitors data access in real time, performs intelligent access control and auditing, and ensures data security.
[0090] It can be seen that the data security agent uses the Qwen-7B model to analyze data sensitivity and security requirements, generate and optimize data desensitization and encryption strategies, and dynamically adjusts security strategies through the PPO algorithm to ensure the security of data throughout its entire life cycle.
[0091] like Figure 2 As shown, the above-mentioned artificial intelligence-based data governance system also includes: API service agent, whose achievable functions include:
[0092] 7. Data service release.
[0093] 7.1. API service registration: The API service agent registers data resources as data assets and generates API interfaces for data users to subscribe to and call.
[0094] 7.2. Intelligent recommendation: Through the PPO algorithm, the most suitable data services and API interfaces are intelligently recommended based on user needs and data usage.
[0095] It can be seen that the API service agent uses the PPO algorithm to optimize service recommendations and subscription strategies, combines user needs and data usage, provides personalized data services and API interfaces, and improves user experience and data utilization.
[0096] like Figure 2 As shown, the above-mentioned artificial intelligence-based data governance system also includes: a data visualization agent, whose achievable functions include:
[0097] 8. Data visualization.
[0098] 8.1. Report generation: The data visualization agent uses Qwen-7B and CLIP (multimodal pre-trained model) models to generate visualization reports and dashboards, supporting multi-dimensional data analysis and display.
[0099] 8.2. Self-service configuration: Users can configure personalized visual reports through simple operations. The PPO algorithm optimizes the configuration process and improves the user experience.
[0100] 8.3. Real-time display: The visualization agent displays data analysis results in real time, supporting users to conduct in-depth data exploration and decision-making.
[0101] It can be seen that the data visualization agent combines the Qwen-7B and CLIP models to generate intelligent visualization reports and dashboards, supports multi-dimensional data analysis and display, and optimizes the configuration process through the PPO algorithm. Users can generate personalized visualization reports through simple operations to enhance data insights.
[0102] Alternatively, if Figure 2 As shown, the above-mentioned artificial intelligence-based data governance system also includes a collaboration layer, which includes collaborative agents and scheduling agents, and its achievable functions include:
[0103] 9. Collaborative agent: responsible for coordinating the work of the core layer agents to ensure information sharing and task collaboration. Use the PPO algorithm to achieve efficient task allocation and resource coordination. This agent understands the needs and status of each agent through the Qwen-7B model, and dynamically optimizes task scheduling and resource allocation through the PPO algorithm.
[0104] 10. Scheduling agent: responsible for the scheduling and resource allocation of system tasks, and performs intelligent scheduling and dynamic resource allocation based on the PPO algorithm to ensure efficient operation of the system. The Qwen-7B model is used to understand complex scheduling requirements, and the PPO algorithm is used to optimize the scheduling strategy and achieve optimal resource allocation.
[0105] It can be seen that the collaborative agents and scheduling agents achieve efficient task allocation and resource coordination through the PPO algorithm, ensuring the efficient operation and intelligent collaboration of the system and improving the overall performance and resource utilization of the data governance platform.
[0106] This application realizes the intelligence and automation of data governance by building an artificial intelligence-based data governance system, introducing advanced reinforcement learning algorithms and large-scale pre-training models.
[0107] Figure 3 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 3As shown, the electronic device may include: a processor 310, a communication interface 320, a memory 330 and a communication bus 340, wherein the processor 310, the communication interface 320 and the memory 330 communicate with each other through the communication bus 340. The processor 310 may call the logic instructions in the memory 330 to execute the data governance method based on artificial intelligence, which includes: identifying the data source; according to the category of the identified data source, using the data collection strategy corresponding to the category of the data source to collect raw data from the data source; inputting the raw data into the large language model to obtain the data quality detection result output by the large language model; the data quality detection result includes at least one of data integrity, data accuracy and data consistency; the large language model is also used to generate data standards according to the data quality detection results; adjusting the raw data according to the data standard, and storing the adjusted raw data in the database.
[0108] In addition, the logic instructions in the above-mentioned memory 330 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0109] On the other hand, the present invention also provides a computer program product, which includes a computer program, and the computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the artificial intelligence-based data governance method provided by the above-mentioned methods, and the method includes: identifying a data source; according to the category of the identified data source, using a data collection strategy corresponding to the category of the data source to collect original data from the data source; inputting the original data into a large language model to obtain a data quality detection result output by the large language model; the data quality detection result includes at least one of data integrity, data accuracy and data consistency; the large language model is also used to generate data standards based on the data quality detection results; adjusting the original data according to the data standard, and storing the adjusted original data in a database.
[0110] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the artificial intelligence-based data governance method provided by the above-mentioned methods, the method comprising: identifying a data source; according to the category of the identified data source, collecting original data from the data source using a data collection strategy corresponding to the category of the data source; inputting the original data into a large language model to obtain a data quality detection result output by the large language model; the data quality detection result includes at least one of data integrity, data accuracy and data consistency; the large language model is also used to generate data standards based on the data quality detection results; adjusting the original data according to the data standards, and storing the adjusted original data in a database.
[0111] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0112] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data governance method based on artificial intelligence, characterized in that: include: Identify data sources; According to the identified category of the data source, adopting a data collection strategy corresponding to the category of the data source to collect raw data from the data source; Inputting the original data into a large language model to obtain a data quality detection result output by the large language model; the data quality detection result includes at least one of data integrity, data accuracy and data consistency; the large language model is also used to generate a data standard according to the data quality detection result; The original data is adjusted according to the data standard, and the adjusted original data is stored in a database.
2. The artificial intelligence-based data governance method according to claim 1, characterized in that: After adjusting the original data according to the data standard and storing the adjusted original data in the database, the method further includes: Monitor the data usage status of the data user terminal; Based on the data usage status, the proximal policy optimization model is used to output the optimal data service policy to the data user, so that the data user obtains the adjusted original data according to the optimal data service policy.
3. The data governance method based on artificial intelligence according to claim 1 is characterized in that: After adjusting the original data according to the data standard and storing the adjusted original data in the database, the method further includes: Inputting the adjusted original data into the large language model to obtain a data analysis result output by the large language model; The data analysis result is input into a multimodal pre-training model to obtain a data display image for the data analysis result output by the multimodal pre-training model.
4. The data governance method based on artificial intelligence according to claim 2 is characterized in that: The data usage status includes the access requirements of the data user for the target data; the optimal data service strategy includes a security strategy; and based on the data usage status, outputting the optimal data service strategy to the data user using a proximal strategy optimization model so that the data user obtains the adjusted original data according to the optimal data service strategy includes: Based on the access requirements for the target data, a proximal policy optimization model is used to output a security policy so that the data user obtains the target data according to the security policy.
5. The data governance method based on artificial intelligence according to claim 1 is characterized in that: The large language model is also used to output a data sensitivity analysis result for the original data; before adjusting the original data according to the data standard and storing the adjusted original data in the database, the method further includes: Desensitizing or encrypting the original data according to the data sensitivity analysis result to obtain secure data; The adjusting the original data according to the data standard and storing the adjusted original data into the database comprises: The security data is adjusted according to the data standard, and the adjusted security data is stored in the database.
6. A data governance system based on artificial intelligence, characterized in that: include: Data collection agent, used to identify data sources; According to the identified category of the data source, adopting a data collection strategy corresponding to the category of the data source to collect raw data from the data source; A data quality agent, used to input the original data into a large language model to obtain a data quality detection result output by the large language model; the data quality detection result includes at least one of data integrity, data accuracy and data consistency; The data standard agent is used to input the data quality detection result into the large language model to obtain the data standard output by the large language model; adjust the original data according to the data standard, and store the adjusted original data into the database.
7. The artificial intelligence-based data governance system according to claim 6, characterized in that: The system also includes an API service agent; The API service agent is used to monitor the data usage status of the data user end; The API service agent is also used to output the optimal data service strategy to the data user end based on the data usage status using the proximal strategy optimization model, so that the data user end obtains the adjusted original data according to the optimal data service strategy.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, it implements the artificial intelligence-based data governance method as described in any one of claims 1 to 5.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the artificial intelligence-based data governance method as described in any one of claims 1 to 5.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the artificial intelligence-based data governance method as described in any one of claims 1 to 5.