Data management method based on large model and co-driving mode

Through the data governance methods of large models and co-pilot models, automated data collection, standardization and service provision have been solved, and the problem of low efficiency of existing data governance has been achieved, and efficient data management and quality evaluation have been achieved.

CN120296068APending Publication Date: 2025-07-11CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510140474.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing data governance methods mainly rely on manual operations, are low in efficiency, have a large workload for users, and the level of data governance needs to be improved.

Method used

Data governance methods based on large models and co-pilot models are adopted to collect, analyze, standardize and store data through the data sharing platform, and use the instruction input unit to automatically formulate data standards and generate data models, provide data services and quality evaluation information, and reduce user manual operations.

Benefits of technology

It improves data governance efficiency, reduces user workload, improves the automation level of data management, ensures the effectiveness and reliability of data assets, and quickly responds to upper-level application needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296068A_ABST
    Figure CN120296068A_ABST
Patent Text Reader

Abstract

The invention provides a data governance method based on a large model and a co-driving mode, which is applied to a data sharing platform, and in the method, the data sharing platform communicates with a data source through a data interface, data of the data source is collected, analyzed, standardized and forwarded into a warehouse for storage, and a user formulates a data standard through the data sharing platform. The data sharing platform performs data modeling according to a data standard, generates a data model, calls the data model by an ETL data development task, processes the stored data, registers the processed data as data assets, and develops the data assets into data services through an internal API interface of the data sharing platform; the data sharing platform provides the data assets and the data services for the subscribers to check and use through the data opening portal, meanwhile, the quality evaluation information of the data assets is provided for the subscribers, and the method can guarantee the validity and reliability of the data assets and is beneficial to improving the enterprise data governance level and efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data governance, and particularly to a data governance method based on a large model and a co-pilot mode. Background Art

[0002] With the development of enterprise business, higher requirements are put forward for digital transformation. At present, the data governance levels of many enterprises need to be improved. After a large amount of data is put into the lake, there is a lack of control and application, resulting in insufficient external empowerment of the data center and the inability to provide convenient data opening and processing capabilities. Data governance refers to the management process of data within an organization, which includes formulating policies, standards, and procedures to ensure data quality, security, and compliance. The main goal of data governance is to improve the value of data while reducing data-related risks.

[0003] Artificial intelligence large models refer to artificial intelligence models with extremely large-scale parameters and data, which can process complex natural language, images, audio, and other multi-modal information to achieve intelligent generation, understanding, reasoning, and other capabilities. The co-pilot mode refers to an application mode that uses artificial intelligence large models as auxiliary tools for humans to provide services such as information, advice, feedback, and guidance to help humans complete certain tasks or improve certain capabilities.

[0004] Data governance usually involves multiple aspects such as data standard management, metadata management, data asset management, and indicator management. However, currently, these managements often require manual operations by users, with a large workload and room for further improvement in efficiency. The digital governance level still needs to be further improved. Therefore, it is necessary to study a data governance method that combines artificial intelligence large models and the co-pilot mode to make full use of the powerful data processing capabilities of artificial intelligence large models and improve the efficiency and effect of data governance. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a data governance method based on a large model and a co-pilot mode to solve the problems that current data governance mainly relies on manual management, with low efficiency, large user workload, and room for further improvement in the data governance level.

[0006] To achieve the above-mentioned invention purpose, the present invention provides a data governance method based on a large model and a co-pilot mode, and the method includes:

[0007] S101. The data sharing platform communicates with the data source through a data interface, collects, parses, standardizes, and forwards the data from the data source for storage in the database;

[0008] S102. The user formulates data standards through the data sharing platform;

[0009] S103. The data sharing platform performs data modeling according to the data standard to generate a data model;

[0010] S104. The ETL data development task calls the data model to process the data stored in the warehouse;

[0011] S105. Register the processed data as a data asset, and develop the data asset into a data service through the internal API interface of the data sharing platform;

[0012] S106. The data sharing platform provides data assets and data services through the data open portal for subscribers to view and use, and at the same time provides quality evaluation information of the data assets for subscribers.

[0013] Further, the data sharing platform includes an instruction input unit. In step S102, the user formulates a data standard through the data sharing platform, which specifically includes the following operations:

[0014] S201. The user inputs a dialogue instruction regarding formulating a data standard through the instruction input unit;

[0015] S202. The instruction input unit understands the dialogue instruction input by the user through a large language model and identifies the user's intention;

[0016] S203. The instruction input unit generates a data standard management instruction according to the user intention recognition result;

[0017] S204. The data sharing platform creates a data standard according to the data standard management instruction.

[0018] Further, the data sharing platform includes an instruction input unit. Subscribers obtain data assets or data services through the data open portal, which specifically includes the following operations:

[0019] S301. The user inputs a dialogue instruction regarding obtaining a data asset or data service through the instruction input unit;

[0020] S302. The instruction input unit understands the dialogue instruction input by the user through a large language model and identifies the user's intention;

[0021] S303. The instruction input unit generates a subscription instruction for a data asset or a data service according to the user intention recognition result;

[0022] S304. The data sharing platform obtains the corresponding data asset or data service according to the subscription instruction for the data asset or the data service, and provides the obtained data asset or data service to the subscriber through the data open portal.

[0023] Further, the instruction input unit uses a large language model to understand the conversation instructions input by the user and identify the user's intention, which specifically includes the following operations:

[0024] S401. Construct a concept context graph based on the conversation instructions and a preset knowledge base;

[0025] S402. Convert each node in the concept context graph into a vector to obtain the embedding representation of the vector of each node;

[0026] S403. Based on the embedding representation of the node vectors, use a graph attention network to calculate the attention features of the nodes to obtain the node feature representation;

[0027] S404. Inject global information into all nodes of the concept context graph through a BERT encoder, and output the global representation of the concept context graph and the sequence output result of the hidden layer;

[0028] S405. Input the global representation of the concept context graph into an intention classifier to obtain the probabilities of different intentions, and input the sequence output result of the hidden layer into a slot classifier to obtain the probability distribution of the slots;

[0029] S406. Perform vector connection on the global representation of the concept context graph and the representation of the slots to obtain the slot-based intention recognition result, and calculate based on the output of the intention classifier and the slot-based intention recognition result to obtain the final user intention recognition result.

[0030] Further, before step S106, there is also a step: evaluating the quality of the registered data assets.

[0031] Further, evaluating the quality of the registered data assets specifically includes the following operations:

[0032] S501. Construct a quality evaluation index system for data assets, where the quality evaluation index system includes several primary indicators, and each primary indicator includes several secondary indicators;

[0033] S502. Calculate the weights of each level of indicators;

[0034] S503. Based on the weights of each level of indicators, determine the quality of the target data assets by constructing a cloud model, and output the quality evaluation result of the target data assets.

[0035] Further, calculating the weights of each level of indicators specifically includes the following steps:

[0036] S601. Calculate the first-class weights of the primary indicators and secondary indicators through the entropy method;

[0037] S602. Calculate the second-class weights of the primary indicators and secondary indicators through the order relation analysis method;

[0038] S603. Calculate their respective final weights based on the first - class weights and the second - class weights of the primary indicators and the secondary indicators respectively.

[0039] Further, the step S503 specifically includes the following operations:

[0040] S701. Generate standard cloud models of different evaluation levels according to the evaluation set and value range of the data asset quality evaluation. The standard cloud models of the evaluation levels are used to represent the value ranges, expected values, entropy values, and hyper - entropy values corresponding to different evaluations;

[0041] S702. Calculate the cloud model characteristic values of the primary indicators and the secondary indicators respectively;

[0042] S703. Based on the cloud model characteristic values and weights of the secondary indicators, perform information aggregation to calculate the evaluation cloud matrix of the primary indicators;

[0043] S704. Based on the cloud model characteristic values and weights of the primary indicators, perform information aggregation to calculate the comprehensive quality evaluation cloud model of the target data asset;

[0044] S705. Compare the comprehensive quality evaluation cloud model of the target data asset with the standard cloud models of different evaluation levels, determine the standard cloud model that best matches it, and use the evaluation level corresponding to the best - matching standard cloud model as the final quality evaluation level of the target data asset.

[0045] Compared with the prior art, the beneficial effects of the present invention are:

[0046] 1) After the method provided by the present invention collects, parses, standardizes, and forwards the data of the data source for storage in the database, it performs data modeling according to the data standards formulated by the user, calls the data model through the ETL data development task, processes the data stored in the database, registers the processed data as data assets, develops the data assets into data services through the API interface, and provides data assets and data services through the data development portal for subscribers to view and use, so as to quickly respond to the requirements of upper - layer applications. At the same time, it can also provide the quality evaluation information of the data assets for subscribers as a reference, so as to facilitate the control of the quality of data assets and ensure the effectiveness and reliability of data assets;

[0047] 2) In the method provided by the present invention, the user can input a conversation instruction through the instruction input unit. The instruction input unit can understand the user's conversation instruction through the large language model, identify the user's intention, and thus generate a series of instructions such as data standard management, data asset or data service acquisition based on the user's intention, helping the user quickly realize data standard management, or the acquisition and application of data assets and data services. While effectively reducing the user's workload and improving management efficiency, it can retain the user's initiative and creativity, which is helpful for further improving the enterprise data governance level. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only the preferred embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0049] Figure 1 It is a schematic diagram of the overall process of a data governance method based on a large model and a co-pilot mode provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] The following describes the principles and features of the present invention with reference to the drawings. The listed embodiments are only used to explain the present invention and are not used to limit the scope of the present invention.

[0051] Refer to Figure 1 , this embodiment provides a data governance method based on a large model and a co-pilot mode. The method is applied to a data sharing platform, and the data sharing platform is used to realize unified data collection, unified storage, unified processing, unified sharing, and unified governance. Internally, it serves as a data processing factory, and externally provides a unified data consumption entrance. The method includes the following steps:

[0052] S101. The data sharing platform communicates with the data source through the data interface, collects, parses, standardizes, and forwards the data from the data source for storage in the database.

[0053] Exemplarily, the data interface can include various types such as FTP / SFTP, SDTP, JDBC, Kafka / MQ, SNMP, Socket, OGG, Corba, MML, API, etc. The collected data from the data source finally enters the distributed file system for storage.

[0054] S102. The user formulates data standards through the data sharing platform.

[0055] A data standard refers to a normative constraint that ensures the consistency and accuracy of internal and external use and exchange of data. Exemplarily, data standards can include multiple aspects such as data layering, data partitioning, field libraries, word segmentation libraries, time dimensions, space dimensions, other dimensions, data models, asset themes, professional categories, data types, etc. Users can perform a series of management operations on data standards through a data sharing platform, including but not limited to adding, modifying, deleting, adding versions, deleting versions, copying versions, importing and exporting data standards, enabling and disabling versions, etc.

[0056] S103. The data sharing platform performs data modeling according to the data standard to generate a data model.

[0057] A data model is an abstraction of data characteristics, which describes the static characteristics, dynamic behaviors, and constraint conditions of a system at an abstract level, providing an abstract framework for information representation and operations of a database system.

[0058] S104. The ETL data development task invokes the data model to process the data stored in the repository.

[0059] S105. Register the processed data as a data asset, and develop the data asset into a data service through the internal API interface of the data sharing platform.

[0060] S106. The data sharing platform provides data assets and data services through a data open portal for subscribers to view and use, and at the same time provides quality evaluation information of the data assets for subscribers.

[0061] The data open portal is a data service publishing platform based on data assets, which develops data services through a front-end page, provides various data sharing methods such as real-time data sharing, data service (API) sharing, and batch file sharing for data users to subscribe to data, and quickly responds to the requirements of upper-layer applications.

[0062] As a preferred implementation, the data sharing platform includes an instruction input unit. In step S102, the user formulates a data standard through the data sharing platform, which specifically includes the following operations:

[0063] S201. The user inputs a dialogue instruction regarding formulating a data standard through the instruction input unit.

[0064] S202. The instruction input unit understands the dialogue instruction input by the user through a large language model and identifies the user's intention.

[0065] S203. The instruction input unit generates a data standard management instruction according to the user intention recognition result.

[0066] S204. The data sharing platform creates a data standard according to the data standard management instruction.

[0067] In this embodiment, the user can input management instructions for data standards in a conversational manner through the instruction input unit. The instruction input unit understands and recognizes the actual intention of the user through a large language model, and automatically generates data standard management instructions according to the recognition result of the user's intention, enabling the data sharing platform to create data standards according to the data standard management instructions, or perform other management operations on the data standards, thus eliminating the need for the user to manually configure and manage the data standards and reducing the user's workload.

[0068] Correspondingly, subscribers can obtain data assets or data services through the data opening portal in a similar manner, which specifically includes the following operations:

[0069] S301. The user inputs a conversational instruction for obtaining data assets or data services through the instruction input unit.

[0070] S302. The instruction input unit understands the conversational instruction input by the user through a large language model and recognizes the user's intention.

[0071] S303. The instruction input unit generates a subscription instruction for data assets or data services according to the user intention recognition result.

[0072] S304. The data sharing platform obtains the corresponding data assets or data services according to the subscription instruction for data assets or data services, and provides the obtained data assets or data services to the subscriber through the data opening portal.

[0073] In the above steps, the user as a subscriber can input a subscription instruction for data assets or data services in a conversational manner through the instruction input unit. The instruction input unit understands and recognizes the actual intention of the user through a large language model, and automatically generates a subscription instruction for data assets or data services according to the recognition result of the user's intention, enabling the data sharing platform to further provide the corresponding data assets or data services to the subscriber according to the subscription instruction. Subscribers can quickly achieve batch subscriptions of various data assets and data services of different types and contents in this way, thereby improving the efficiency of upper-layer data analysis applications.

[0074] On the basis of the foregoing embodiment, as a further possible embodiment, the instruction input unit understands the conversational instruction input by the user through a large language model and recognizes the user's intention, which specifically includes the following operations:

[0075] S401. Construct a concept context graph based on the conversational instruction and a preset knowledge base.

[0076] In this embodiment, the preset knowledge base is a knowledge graph that describes data governance-related knowledge in natural language, including multiple tuples, and each tuple is composed of a head concept, a relationship, a tail concept, and a confidence level. The concept context graph is a data structure used to represent and process natural language text concepts and their context relationships, capturing concepts, entities, and their relationships in the text in the form of a graph, so as to provide richer context information for the dialogue instruction processing task.

[0077] When constructing the concept context graph, first filter out the tuples in the preset knowledge base that are not relevant to the dialogue instruction according to the confidence level and the relationship, then select multiple tuples that are most relevant to the word elements of the dialogue instruction text to form a concept subgraph, and construct the concept context graph based on the concept subgraph. The nodes of the concept context graph are composed of the word elements of the dialogue instruction text and their corresponding concept tuples, and the edges of the concept context graph include the edges between two adjacent word elements, the edges between the word elements and their corresponding concept tuples, and the edges between the status markers and other nodes.

[0078] S402. Convert each node in the concept context graph into a vector to obtain the embedding representation of the vector of each node.

[0079] In this step, the embedding representation of the vector of the node can be expressed as:

[0080] n i =V w (n i )+V p (n i )+V t (n i )

[0081] Among them, n i represents the i-th node in the concept context graph, V w (n i ) represents the vector representation of the word element of the dialogue text of the i-th node, V p (n i ) represents the vector representation of the position of the i-th node, and V t (n i ) represents the vector representation of the type of the i-th node.

[0082] S403. Based on the embedding representation of the vector of the node, use the graph attention network to calculate the attention feature of the node to obtain the node feature representation.

[0083] Exemplarily, the attention feature of the node, that is, the attention score between the node and its adjacent nodes, is calculated as follows:

[0084] af ij =sa(n i ,nj )

[0085] where af ij represents the attention score between node i and node j, and n j represents the j-th node in the concept context graph, and sa() represents a self-attention based neural network. The final node feature representation is obtained by concatenating the attention features of the node and its own features, and the expression is as follows:

[0086]

[0087] where is the final node feature representation, and ad i is the set of all adjacent nodes of node i, is the self-feature of node i.

[0088] S404. Inject global information into all nodes of the concept context graph through a BERT encoder, and output the global representation of the concept context graph and the sequence output result of the hidden layer.

[0089] Among them, the global representation of the concept context graph is a summary of the input semantic content, which can be used for the intent recognition task in the future; the sequence output result of the hidden layer is used for the slot filling task in the future.

[0090] S405. Input the global representation of the concept context graph into an intent classifier to obtain the probabilities of different intents, and input the sequence output result of the hidden layer into a slot classifier to obtain the probability distribution of slots.

[0091] In this implementation, the recognition of the user's intent is completed by the intent classifier. After the intent classifier receives the global representation of the concept context graph, it uses the sigmoid activation function to classify the intent and obtain the probabilities of multiple intent labels. At the same time, in the slot filling task, after the sequence output result of the hidden layer is input into the slot classifier, a normalization operation is performed to obtain the probability distribution of slots.

[0092] S406. Vectorially connect the global representation of the concept context graph and the representation of the slot to obtain the slot-based intent recognition result, and calculate based on the output of the intent classifier and the slot-based intent recognition result to obtain the final user intent recognition result.

[0093] In this step, calculating based on the output of the intent classifier and the slot-based intent recognition result specifically means multiplying the output result of the intent classifier and the slot-based intent recognition result element by element, and using the calculation result as the final user intent recognition result.

[0094] This embodiment introduces a preset knowledge base based on the dialogue instructions input by the user, constructs a concept context graph according to the dialogue instructions and the knowledge content of the preset knowledge base, then processes the concept context graph based on the graph attention network and the BERT encoder, and inputs the output results of the BERT encoder into the intent classifier and the slot classifier respectively. The final user intent recognition result is obtained by combining the classification results of the intent classifier and the classification results of the slots, enabling the instruction input unit to understand the complex concepts and intents in the user's dialogue instructions, enabling the data sharing platform to accurately execute the user's instructions, and improving the data governance efficiency.

[0095] To ensure the reliability and effectiveness of data assets, in this embodiment, the data sharing platform provides quality evaluation information of the data assets to the subscribers while providing the data assets. High-quality data is the basis for enterprise decision-making. Based on the analysis of the accuracy, integrity, consistency, and timeliness of the data, enterprises can optimize business processes and improve work efficiency. Before providing the evaluation information of the data assets to the subscribers, it is necessary to evaluate the quality of the registered data assets.

[0096] As a preferred embodiment, evaluating the quality of the registered data assets specifically includes the following operations:

[0097] S501. Construct a quality evaluation index system for the data assets. The quality evaluation index system includes several primary indicators, and each primary indicator includes several secondary indicators.

[0098] For example, the primary indicators may include integrity, consistency, interpretability, security, etc., and the secondary indicators under the consistency primary indicator include data source consistency, time consistency, format consistency, etc. The establishment of the quality evaluation index system can be formulated by combining various methods such as literature data analysis and expert consultation. The specific content of the primary and secondary indicators can be formulated according to actual needs, and this embodiment does not make specific limitations on this.

[0099] S502. Calculate the weights of each level of indicators.

[0100] As a further possible embodiment, calculating the weights of each level of indicators specifically includes the following operations:

[0101] S601. Calculate the first-class weights of the primary indicators and secondary indicators by the entropy method.

[0102] The entropy value method is a multi-index decision-making weight assignment method that objectively determines weights by measuring the dispersion of indicators. When the numerical values of indicators vary greatly, the entropy value is small, indicating that the amount of information is large, and a larger weight should be assigned; conversely, a smaller weight should be assigned. First, it is necessary to standardize the evaluation values of the evaluation object for primary indicators or secondary indicators to eliminate the influence of dimensions. Then, calculate the information entropy according to the probability distribution of each indicator, and calculate the weight of each indicator based on the information entropy.

[0103] S602. Calculate the two types of weights of primary indicators and secondary indicators through the order relation analysis method.

[0104] The order relation analysis method is a decision analysis method that improves the analytic hierarchy process. By combining quantitative and qualitative factors, the indicators to be decided are classified and graded. Finally, considering the influence of each indicator comprehensively, the weights of each indicator are adjusted according to the situation.

[0105] In this step, first, it is necessary to determine the order relation among various indicators. The order relation can be determined by experts according to the importance of each indicator. The higher the importance of the indicator, the more forward it is in the order relation. Then, determine the ratio of the importance between two adjacent indicators in the order relation, and determine the weights of each indicator based on this ratio.

[0106] S603. Calculate their corresponding final weights based on the first type of weights and the second type of weights of primary indicators and secondary indicators respectively.

[0107] In this implementation manner, the final weights are calculated by combining the first type of weights and the second type of weights of primary indicators and secondary indicators through the combined weight assignment method. The combined weight assignment method comprehensively reflects the subjective weight assignment that reflects the experience and preference of the decision maker and the objective weight assignment based on the statistical characteristics and internal correlations of indicators, and can improve the scientificity of weight decision-making. The calculation formula for the final weight is:

[0108]

[0109] In the above formula, W represents the final weight, w 1j represents the first type of weight of the jth indicator, and w 2j represents the second type of weight of the jth indicator.

[0110] S503. Based on the weights of secondary indicators, determine the quality of the target data asset by constructing a cloud model, and output the quality evaluation result of the target data asset.

[0111] The cloud model is a model that realizes the bidirectional conversion between qualitative concepts and quantitative data, and is composed of the expected value Ex, entropy value En, and hyper-entropy value He. In this embodiment, the expected value represents the overall evaluation level of the quality evaluation index, and the higher the expected value, the higher the importance of the index. The entropy value reflects the degree of dispersion of the evaluation of the quality evaluation index, and the greater the uncertainty of the evaluation result, the larger its value. The hyper-entropy value is a further manifestation of the uncertainty, reflecting the deeper uncertainty in the evaluation process. The larger its value, the more discrete and uncertain the evaluation result of the corresponding index.

[0112] As a further possible implementation, step S503 specifically includes the following operations:

[0113] S701. Generate standard cloud models for different evaluation levels according to the evaluation set and value range of the data asset quality evaluation. The evaluation level standard cloud model is used to represent the value range, expected value, entropy value, and hyper-entropy value corresponding to different evaluations.

[0114] S702. Calculate the cloud model characteristic values of the primary indicators and secondary indicators respectively.

[0115] In this implementation, the data sharing platform scores n indicators respectively to obtain the decision evaluation matrix X = (x i ), (i = 1, 2,..., n), and x n represents the evaluation result of the platform for the i-th indicator. Then the characteristic value of the cloud model of the i-th indicator can be calculated by the following formula: i In the above formula,

[0116]

[0117]

[0118] where represents the mean value. The calculated cloud model characteristic value can be tested for its reliability through the consistency test index, and the expression is:

[0119]

[0120] If Ce is less than 1, it means that the result is reliable.

[0121] S703. Based on the cloud model characteristic values and weights of the secondary indicators, perform information aggregation to calculate the evaluation cloud matrix of the primary indicators.

[0122] In this step, based on the cloud model characteristic values of the secondary indicators, perform a first information aggregation on the cloud of the secondary indicators according to the weight values of the secondary indicators to calculate the evaluation cloud matrix of the primary indicators. The calculation formula is:

[0123]

[0124] In the above formula, w ij represents the weight of the j-th secondary index under the i-th primary index. Correspondingly, Ex ij , En ij and He ij are the cloud model characteristic values of the j-th secondary index under the i-th primary index. Ex i , En i and He i are the cloud model characteristic values of the i-th primary index.

[0125] S704. Information aggregation is performed based on the cloud model characteristic values and weights of the primary indicators, and the comprehensive quality evaluation cloud model of the target data asset is calculated.

[0126] In this step, based on the same principle as S703, information aggregation operations are performed again on the cloud model characteristic values and weights of the primary indicators, and the comprehensive quality evaluation cloud model of the target data asset can be calculated.

[0127] S705. The comprehensive quality evaluation cloud model of the target data asset is compared with the standard cloud models of different evaluation levels, the most matching standard cloud model is determined, and the evaluation level corresponding to the most matching standard cloud model is used as the final quality evaluation level of the target data asset.

[0128] Exemplarily, let the comprehensive quality evaluation cloud model be C1(Ex1, En1, He1) and the standard cloud model be C2(Ex2, En2, He2), then the matching degree MD(C1, C2) is calculated by the following formula:

[0129]

[0130] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A data governance method based on a large model and a co-pilot mode, characterized in that, The method is applied to a data sharing platform, and the method includes the steps: S101. The data sharing platform communicates with the data source through a data interface, collects, parses, standardizes, and forwards the data from the data source for storage in the database; S102. The user formulates data standards through the data sharing platform; S103. The data sharing platform performs data modeling according to the data standards to generate a data model; S104. The ETL data development task calls the data model to process the data stored in the database; S105. Register the processed data as a data asset, and develop the data asset into a data service through the internal API interface of the data sharing platform; S106. The data sharing platform provides data assets and data services through a data opening portal for subscribers to view and use, and at the same time provides quality evaluation information of the data assets for subscribers.

2. The data governance method based on the large model and the co-pilot mode according to claim 1, wherein The data sharing platform includes an instruction input unit. In step S102, when the user formulates data standards through the data sharing platform, it specifically includes the following operations: S201. The user inputs a dialogue instruction for formulating data standards through the instruction input unit; S202. The instruction input unit understands the dialogue instruction input by the user through a large language model and identifies the user's intention; S203. The instruction input unit generates a data standard management instruction according to the user intention recognition result; S204. The data sharing platform creates a data standard according to the data standard management instruction.

3. The data governance method based on a large model and a co-pilot mode according to claim 1, characterized in that, The data sharing platform includes an instruction input unit. When a subscriber obtains a data asset or a data service through the data opening portal, it specifically includes the following operations: S301. The user inputs a dialogue instruction for obtaining a data asset or a data service through the instruction input unit; S302. The instruction input unit understands the dialogue instruction input by the user through a large language model and identifies the user's intention; S303. The instruction input unit generates a subscription instruction for the data asset or the data service according to the user intention recognition result; S304. The data sharing platform obtains the corresponding data asset or data service according to the subscription instruction for the data asset or the data service, and provides the obtained data asset or data service to the subscriber through the data opening portal.

4. A data governance method based on a large model and a co-pilot mode according to any one of claims 2 or 3, characterized in that, The instruction input unit understands the dialogue instruction input by the user through a large language model and identifies the user's intention, which specifically includes the following operations: S401. Construct a concept context graph based on the dialogue instruction and a preset knowledge base; S402. Convert each node in the concept context graph into a vector to obtain the embedding representation of the vector of each node; S403. Based on the embedding representation of the vector of the node, use a graph attention network to calculate the attention feature of the node to obtain the node feature representation; S404. Inject global information into all nodes of the concept context graph through a BERT encoder, and output the global representation of the concept context graph and the sequence output result of the hidden layer; S405. Input the global representation of the concept context graph into an intention classifier to obtain the probabilities of different intentions, and input the sequence output result of the hidden layer into a slot classifier to obtain the probability distribution of the slots; S406. Vector-connect the global representation of the concept context graph with the representation of the slot to obtain the slot-based intention recognition result. Calculate based on the output of the intention classifier and the slot-based intention recognition result to obtain the final user intention recognition result.

5. A data governance method based on a large model and a co-pilot mode according to claim 1, characterized in that Before step S106, there is also a step of evaluating the quality of the registered data assets.

6. The data governance method based on the large model and the co-pilot mode according to claim 5, wherein, Evaluate the quality of the registered data assets, specifically including the following operations: S501. Construct a quality evaluation index system for data assets. The quality evaluation index system includes several primary indicators, and each primary indicator includes several secondary indicators. S502. Calculate the weights of each level of indicators. S503. Based on the weights of each level of indicators, determine the quality of the target data assets by constructing a cloud model, and output the quality evaluation result of the target data assets.

7. The data governance method based on the large model and the co-pilot mode according to claim 6, wherein, Calculating the weights of each level of indicators specifically includes the following steps: S601. Calculate the first-class weights of the primary indicators and secondary indicators by the entropy value method. S602. Calculate the second-class weights of the primary indicators and secondary indicators by the order relation analysis method. S603. Calculate their corresponding final weights based on the first-class weights and second-class weights of the primary indicators and secondary indicators respectively.

8. The data governance method based on large models and co-pilot mode according to claim 6, wherein, The specific operations of step S503 include the following: S701. Generate standard cloud models for different evaluation levels according to the evaluation set and value range of data asset quality evaluation. The evaluation level standard cloud model is used to represent the value range, expected value, entropy value, and hyper-entropy value corresponding to different evaluations. S702. Calculate the cloud model characteristic values of the primary indicators and secondary indicators respectively. S703. Conduct information aggregation based on the cloud model characteristic values and weights of the secondary indicators, and calculate the evaluation cloud matrix of the primary indicators. S704. Conduct information aggregation based on the cloud model characteristic values and weights of the primary indicators, and calculate the quality comprehensive evaluation cloud model of the target data assets. S705. Compare the quality comprehensive evaluation cloud model of the target data assets with the standard cloud models of different evaluation levels, determine the standard cloud model that best matches it, and use the evaluation level corresponding to the best-matching standard cloud model as the final quality evaluation level of the target data assets.