Data management method and system
By applying the large language model (LLM) technology in data governance, semantic retrieval and vector database optimization, building a domain knowledge base, and generating SQL code to process data, the problem of existing data governance methods is solved, and the efficiency and effectiveness of data governance is improved.
Patent Information
- Application Number
- CN202411993677.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
The existing data governance methods require manual development of data standards and data cleaning, which is time-consuming and labor-intensive, and it is difficult to cope with rapidly changing business needs. The lack of objective and unified data quality evaluation standards and it is difficult to accurately identify the correlation between data, resulting in unsatisfactory data integration results. Traditional anomaly detection methods are difficult to adapt to complex and changeable data scenarios, affecting the reliability of data governance.
The large language model (LLM) technology is used to optimize data query through semantic search and vector database, build a domain knowledge base for search, generate SQL code to process data using LLM, and perform exploratory data analysis of data sets.
It improves the efficiency and effectiveness of data governance, reduces manual intervention, improves the objectivity and consistency of data quality evaluation, enhances the ability to identify data association relationships, and improves the quality and reliability of data integration.
Smart Images

Figure CN119938879A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data governance, and in particular to a data governance method and system. Background Art
[0002] With the advent of the big data era, enterprises have accumulated massive amounts of data assets. Data governance, as an important means to ensure data quality, consistency and availability, has become a key link in the digital transformation of enterprises. However, existing data governance methods have the following technical problems:
[0003] Traditional data governance requires manual formulation of data standards and data cleaning, which is a time-consuming and labor-intensive process. Especially when dealing with unstructured data, the accuracy of manual standardization is low and it is difficult to cope with rapidly changing business needs. Existing data quality assessment methods often rely on manual experience judgment and lack objective and unified assessment standards. Different assessors may come to different quality judgment results for the same data, affecting the consistency of data governance. When dealing with cross-departmental and cross-domain data, due to the lack of in-depth understanding of business semantics, it is difficult to accurately identify the relationship between data, resulting in unsatisfactory data integration results. Traditional rule-based anomaly detection methods are difficult to adapt to complex and changing data scenarios, and often have omissions and false alarms, affecting the reliability of data governance. Existing metadata management requires a lot of manual maintenance and updates, and it is difficult to accurately describe the complex relationship between data, which restricts the effective management and utilization of data assets.
[0004] The above technical problems have seriously affected the efficiency and effectiveness of enterprise data governance. With the development of large language model (LLM) technology, its powerful natural language understanding and knowledge reasoning capabilities provide a new technical approach to solve these problems. Summary of the invention
[0005] To achieve the above objectives and other related objectives, the present invention discloses a data governance method, including: Receive user search information, perform semantic retrieval, and use Embedding to optimize semantic retrieval; Build a domain knowledge base and search within the domain knowledge base; Use the LLM large model to generate SQL code and process the data; Exploratory data analysis of the dataset was performed using a large model.
[0006] Furthermore, the method of optimizing semantic retrieval by using Embedding includes: Generate semantic vectors for pre-stored indicator information and store them in the vector database as a benchmark; After vectorizing the user search index information, search the vector database; Calculate the vector distance between the two and find the vector whose vector distance with the user's search term is less than a preset threshold.
[0007] Furthermore, the building of the domain knowledge base includes: Extract the text content from the original document, cut it into pieces according to semantics, and generate multiple text blocks; The embedding model processes the text block and generates a semantic vector for the text block; Store the semantic vectors and text blocks into the vector database.
[0008] Furthermore, in the process of generating text blocks, metadata extraction and sensitive information detection are performed on the text content.
[0009] Furthermore, the vector distance is cosine similarity.
[0010] In another aspect, the present invention provides a data governance system, comprising: The retrieval module is used to receive user search information, perform semantic retrieval, and optimize semantic retrieval using Embedding; The knowledge base building module is used to build a domain knowledge base and search within the domain knowledge base; The data processing module is used to generate SQL code using the LLM large model to process the data; Analysis module for exploratory data analysis of datasets using large models.
[0011] By adopting the above technical solutions, the large model LLM is applied to the field of data governance, which improves the efficiency of data governance. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, among which:
[0013] Figure 1 Flowchart of this application. DETAILED DESCRIPTION
[0014] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0015] Reference Figure 1, an embodiment of the present invention provides a data governance method, including: S1: Receive user search information, perform semantic retrieval, and use Embedding to optimize semantic retrieval.
[0016] Specifically include: Generate semantic vectors for pre-stored indicator information and store them in the vector database as a benchmark; After vectorizing the user search index information, search the vector database; Calculate the vector distance between the two and find the vector whose vector distance with the user's search term is less than the preset threshold. The vector distance is the cosine similarity.
[0017] In the above, the threshold is set as a dynamic threshold, including: T = T0 × (1 + α × log (n) + β × R); Where: T0 is the initial threshold; n is the number of historical queries; R is the correlation feedback coefficient; α and β are adjustment parameters.
[0018] As mentioned above, by adjusting the search threshold in real time, when there are too many results, the threshold is raised to ensure high-quality matching, which reduces the interference of irrelevant results and greatly improves the relevance of the query. It can automatically adjust the optimal threshold for different business fields, maintain a high threshold for professional terminology-intensive queries to ensure accurate matching, and appropriately relax the threshold for general descriptive queries to improve the recall rate, thus realizing intelligent adaptation of the scene. In terms of load, the threshold is automatically raised during peak hours to reduce computing resource consumption, and the threshold is appropriately lowered during off-peak hours to provide more comprehensive search results, achieving the optimal configuration of resource utilization. By continuously learning user interaction records, continuously optimizing threshold parameters, building a query-click behavior model, and extracting user preferences, the fit of the search is further improved.
[0019] S2: Build a domain knowledge base and search within the domain knowledge base.
[0020] As mentioned above, when industry knowledge is relatively professional, large models cannot ensure accurate and efficient provision. In addition, in the process of using large model capabilities, our internal data and environment cannot be exposed to the outside world and must be fully controllable to avoid any data privacy leakage and security risks. Therefore, we use the method of building a domain knowledge base for retrieval, usually using the method of Embedding + vector search engine + LLM. The processing process is as follows: Extract the text content from the original document, cut it into pieces according to semantics, and generate multiple text blocks; The embedding model processes the text block and generates a semantic vector for the text block; Store the semantic vectors and text blocks into the vector database.
[0021] In the process of generating text blocks, metadata extraction and sensitive information detection are performed on the text content.
[0022] Specifically, when detecting sensitive information, it includes: S=α∑(Wi×Di)+β∑(Cj×Tj)+γ∑(Pk×Ek)×(1+λR); in: S is the final sensitivity score; Wi is the weight of the i-th data dimension; Di is the basic sensitivity score of the i-th data dimension; Cj is the weight of the j-th context factor; Tj is the value of the j-th context influence factor; Pk is the weight of the k-th transmission risk factor; Ek is the assessment value of the k-th transmission risk; R is the time attenuation function; α, β, γ are dimension adjustment coefficients; λ is the time influence factor.
[0023] As mentioned above, the sensitivity detection comprehensively considers the three dimensions of data itself (D), context environment (T) and propagation risk (E). The sensitivity is dynamically adjusted through the time decay function R to reflect the timeliness of information. The context influence factor T is introduced to consider the impact of the data environment on sensitivity. The possibility of information diffusion is evaluated through the propagation risk factor E. The dynamic weight adjustment of different dimensions is realized through α, β, and γ, which can evaluate the sensitivity of information more comprehensively and accurately.
[0024] S3: Use the LLM large model to generate SQL code and process the data; As mentioned above, the big model can quickly generate SQL code snippets based on natural language input and display the results in a visual way, thereby assisting data personnel in their daily work. This reduces the time spent on writing complex queries, so more time can be invested in understanding the business and analyzing query results, thereby obtaining decision support from data results.
[0025] For example: You can create a SQL query from the big model to get a specific set of data, for example: "Show the average income per month in 2023."
[0026] Large models can convert this into SQL queries.
[0027] S4: Exploratory data analysis of the dataset using a large model.
[0028] Data analysts often need to spend a lot of time preparing and cleaning data before analysis. Using big models can provide data preprocessing techniques, such as handling missing values, handling outliers, variable correlation analysis, and suggestions for solving user data quality issues. Data preprocessing suggestions can help simplify the data preparation process and ensure analysis quality.
[0029] An embodiment of the present invention further provides a system, comprising: The retrieval module is used to receive user search information, perform semantic retrieval, and optimize semantic retrieval using Embedding; The knowledge base building module is used to build a domain knowledge base and search within the domain knowledge base; The data processing module is used to generate SQL code using the LLM large model to process the data; Analysis module for exploratory data analysis of datasets using large models.
[0030] Those skilled in the art will appreciate that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art in the art to which the present invention belongs. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with the meanings in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless specifically defined.
[0031] For the method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0032] It can be known from the description of the above implementation modes that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be essentially or partly contributed to the prior art in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute the methods described in the various implementation modes of the present application or certain parts of the implementation modes.
[0033] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data governance method, characterized in that: include: Receive user search information, perform semantic retrieval, and use Embedding to optimize semantic retrieval; Build a domain knowledge base and search within the domain knowledge base; Use the LLM large model to generate SQL code and process the data; Exploratory data analysis of the dataset was performed using a large model.
2. The method according to claim 1, characterized in that The method of optimizing semantic retrieval by using Embedding includes: Generate semantic vectors for pre-stored indicator information and store them in the vector database as a benchmark; After vectorizing the user search index information, search the vector database; Calculate the vector distance between the two and find the vector whose vector distance with the user's search term is less than a preset threshold.
3. The method according to claim 1, characterized in that The construction of the domain knowledge base includes: Extract the text content from the original document, cut it into pieces according to semantics, and generate multiple text blocks; The embedding model processes the text block and generates a semantic vector for the text block; Store the semantic vectors and text blocks into the vector database.
4. The method according to claim 3, characterized in that In the process of generating text blocks, metadata extraction and sensitive information detection are performed on the text content.
5. The method according to claim 2, characterized in that: The vector distance is cosine similarity.
6. A data governance system, characterized in that: include: The retrieval module is used to receive user search information, perform semantic retrieval, and optimize semantic retrieval using Embedding; The knowledge base building module is used to build a domain knowledge base and search within the domain knowledge base; The data processing module is used to generate SQL code using the LLM large model to process the data; Analysis module for exploratory data analysis of datasets using large models.