Construction method, data processing method and device, electronic equipment and storage medium
By building a semantic knowledge base and generating a knowledge graph, the problem of inefficiency of traditional data management methods when facing a large number of heterogeneous data is solved, and data quality is improved and resource optimization is optimized.
Patent Information
- Application Number
- CN202510591712.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional data management methods are inefficient when facing large and heterogeneous data, and are difficult to ensure data quality. They require complex data development and integration processes when integrating and aggregating data, resulting in waste of resources and inconsistency in data.
By obtaining the data set of the business, using the pre-trained data processing model to build a semantic knowledge base, determine the mapping relationship between the metadata content and business terms, and generate a knowledge graph to align business terms and metadata content.
It realizes the conversion of technical metadata content into a completely business-oriented knowledge graph, eliminates the problem of inconsistency in business terms in different departments, reduces the cost of computing power hardware and time, and improves data processing efficiency.
Smart Images

Figure CN120216705A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to a construction method and apparatus, a data processing method and apparatus, an electronic device, and a storage medium. Background Art
[0002] With the popularization of big data technology, the amount of enterprise data has grown exponentially. These data are scattered in different databases, data warehouses, or even distributed file systems. When facing such a large amount of heterogeneous data, traditional data management methods are inefficient, difficult to ensure data quality, and it is difficult to align business calibers when facing multiple data source systems. Complex data development and integration processes are required for data integration and aggregation. Manual data governance is carried out through building data warehouses, constructing multi-level data systems, and adding a series of platforms such as data standards, data quality, and data security. Complex data development and governance consume a large number of data developers and time cycles. In the data application stage, professional data analysts are still required to sort out business data, define metrics, and perform visual analysis. Summary of the Invention
[0003] At least one embodiment of the present disclosure provides a method for constructing a knowledge graph, including: obtaining a data set of a service, where the data set includes at least one type of data source and a plurality of metadata contents corresponding to the at least one type of data source; obtaining a semantic knowledge base of the service based on the data set and a pre-trained data processing model, where the semantic knowledge base includes a plurality of service terms; analyzing the data set to obtain an analysis result; determining a mapping relationship between the plurality of metadata contents and the plurality of service terms according to the analysis result; and generating a knowledge graph of the plurality of service terms based on the semantic knowledge base and the mapping relationship.
[0004] For example, in the construction method provided by an embodiment of the present disclosure, it further includes: monitoring the data set, and in response to obtaining that a plurality of metadata contents of the data source change, retraining the data processing model by using a training engine and the changed plurality of metadata contents, and updating the semantic knowledge base.
[0005] At least one embodiment of the present disclosure provides a data processing method, including: obtaining a data processing request; querying a knowledge graph based on the data processing request to obtain query information; and obtaining response data for responding to the data processing request based on the query information, where the knowledge graph is constructed according to the following operations: obtaining a data set of a service, the data set including at least one type of data source and a plurality of metadata contents corresponding to the at least one type of data source; obtaining a semantic knowledge base of the service based on the data set and a pre-trained data processing model, the semantic knowledge base including a plurality of service terms; analyzing the data set to obtain an analysis result; determining a mapping relationship between the plurality of metadata contents and the plurality of service terms according to the analysis result; and generating a knowledge graph of the plurality of service terms based on the semantic knowledge base and the mapping relationship.
[0006] For example, in the data processing method provided by an embodiment of the present disclosure, it further includes: obtaining a vector library based on the data set and the data processing model, the vector library including vectors of the plurality of metadata contents respectively, and querying the knowledge graph based on the data processing request to obtain query information, including: querying the knowledge graph and the vector library based on the data processing request to obtain the query information.
[0007] For example, in the data processing method provided by an embodiment of the present disclosure, querying the knowledge graph and the vector library based on the data processing request to obtain the query information includes: querying the knowledge graph based on the data processing request to obtain the target service term of the response data; determining the target metadata content corresponding to the target service term based on the mapping relationship; and determining the query information based on the target metadata content and the vector library.
[0008] For example, in the data processing method provided by an embodiment of the present disclosure, the query information includes: the database, data table, and field where the response data is located.
[0009] For example, in the data processing method provided by an embodiment of the present disclosure, obtaining response data for responding to the data processing request based on the query information includes: constructing a database query statement of the query information using a language model based on the query information; and querying the at least one type of data source using the database query statement to obtain the response data.
[0010] For example, in the data processing method provided by an embodiment of the present disclosure, it further includes: analyzing the response data using a language model to generate a response report.
[0011] For example, in the data processing method provided in an embodiment of the present disclosure, it further includes: monitoring the data set, and in response to obtaining that multiple metadata contents of the data source have changed, retraining the data processing model by using a training engine and the changed multiple metadata contents, and updating the semantic knowledge base.
[0012] For example, in the data processing method provided in an embodiment of the present disclosure, it further includes: generating a business glossary based on the semantic knowledge base.
[0013] At least one embodiment of the present disclosure provides a knowledge graph construction device, including: a data source acquisition unit configured to acquire a data set of a service, the data set including at least one type of data source and multiple metadata contents of the at least one type of data source; a semantic knowledge base generation unit configured to obtain a semantic knowledge base of the service based on the data set and a data processing model pre-trained, wherein the semantic knowledge base includes multiple business terms; a metadata content analysis unit configured to analyze the data set to obtain an analysis result; a determination unit configured to determine a mapping relationship between the multiple metadata contents and the multiple business terms according to the analysis result; and a knowledge graph generation unit configured to generate a knowledge graph of the multiple business terms based on the semantic knowledge base and the mapping relationship.
[0014] At least one embodiment of the present disclosure provides a data processing device, including: a request acquisition unit configured to acquire a data processing request; a query unit configured to query a knowledge graph based on the data processing request to obtain query information; and a response data acquisition unit configured to acquire response data for responding to the data processing request based on the query information, where the knowledge graph is constructed according to the method provided in any embodiment of the present disclosure.
[0015] At least one embodiment of the present disclosure provides an electronic device, including: a processor; and a memory including one or more computer program instructions; wherein, when the one or more computer program instructions are run by the processor, they execute the method provided in any embodiment of the present disclosure.
[0016] At least one embodiment of the present disclosure provides a computer-readable storage medium that non-temporarily stores computer-readable instructions, wherein when the computer-readable instructions are executed by a processor, the method provided in any embodiment of the present disclosure is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings described below only relate to some embodiments of the present disclosure and do not limit the present disclosure.
[0018] Figure 1 Shows a flowchart of a method for constructing a knowledge graph provided by at least one embodiment of the present disclosure; Figure 2 Shows a flowchart of a data processing method provided by at least some embodiments of the present disclosure; Figure 3 Shows an architecture diagram of a generative AI for data processing provided by at least one embodiment of the present disclosure; Figure 4 Shows a flowchart of a data processing method provided by at least one embodiment of the present disclosure; Figure 5 Shows an architecture diagram of a data processing model provided by at least one embodiment of the present disclosure; Figure 6 Shows an architecture diagram of an enhanced artificial intelligence governance unit provided by at least one embodiment of the present disclosure; Figure 7 Shows a block diagram of a device for constructing a knowledge graph provided by at least one embodiment of the present disclosure; Figure 8 Shows a block diagram of a data processing device provided by at least one embodiment of the present disclosure; Figure 9 Shows a block diagram of an electronic device provided by at least one embodiment of the present disclosure; Figure 10 Shows a block diagram of another electronic device provided by at least one embodiment of the present disclosure; and Figure 11 Schematically shows a schematic diagram of a computer-readable storage medium according to an embodiment of the present disclosure. Detailed implementation manners
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0020] Unless otherwise defined, technical terms or scientific terms used in this disclosure shall have the ordinary meanings as understood by those of ordinary skill in the field to which this disclosure pertains. The terms "first", "second" and similar terms used in this disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Similarly, terms such as "a", "an" or "the" do not denote a quantity limitation, but mean that there is at least one. Terms such as "comprising" or "including" mean that the elements or items appearing before this term cover the elements or items listed after this term and their equivalents, without excluding other elements or items. Terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. Terms such as "upper", "lower", "left" and "right" are only used to indicate relative position relationships. When the absolute position of the object being described changes, the relative position relationship may also change accordingly.
[0021] The current mainstream data development and governance solution is to use the Data Middle Platform architecture, including data integration, data development, and an operation and maintenance center for data flow and processing. It mainly uses computing engines such as Flink and Spark; Hadoop Distributed File System (HDFS), Simple Storage Service (S3), etc. as storage for massive data; and data standards, data quality, data security, metadata content management, etc. to build the work steps for offline data governance tasks. The processing link is relatively long, consuming a large amount of computing and storage resources. Moreover, it is difficult to avoid problems such as sensitive data leakage, data association loss, and standard alignment errors relying on manual governance, reducing data accuracy and credibility.
[0022] Data analysis services mainly rely on a separate Customer Data Platform or a self-built data analysis system (such as data metrics, data tags, visualization reports, and BI, etc.). Professional business analysts are still needed to identify and sort out the data processed by the data middle platform, perform business meaning conversion and data analysis application construction again, and finally provide a data decision support system for front-line users. When facing different business metrics of different departments, it is necessary to customize metrics and create an analysis system according to the corresponding business scenarios.
[0023] At least one embodiment of the present disclosure provides a method for constructing a knowledge graph, including: obtaining a data set of a business, where the data set includes at least one type of data source and multiple metadata contents corresponding to at least one type of data source; obtaining a semantic knowledge base of the business based on the data set and a pre-trained data processing model, where the semantic knowledge base includes multiple business terms; analyzing the data set to obtain an analysis result; determining a mapping relationship between the multiple metadata contents and the multiple business terms according to the analysis result; and generating a knowledge graph of the multiple business terms based on the semantic knowledge base and the mapping relationship. The knowledge graph constructed by this method aligns business terms with the metadata contents of the database, converts the metadata contents into a knowledge graph that is completely business-oriented, eliminates the inconsistency of business terms among different departments, reduces the differences before business understanding and metadata content definition, and can reduce the computing power hardware cost and time cost.
[0024] Figure 1 The flowchart of a method for constructing a knowledge graph provided by at least one embodiment of the present disclosure is shown.
[0025] As Figure 1 shown, the construction method includes steps S101 to S105.
[0026] Step S101: Obtain a data set of a business, where the data set includes at least one type of data source and multiple metadata contents corresponding to at least one type of data source.
[0027] Step S102: Obtain a semantic knowledge base of the business based on the data set and a pre-trained data processing model, where the semantic knowledge base includes multiple business terms.
[0028] Step S103: Analyze the data set to obtain an analysis result.
[0029] Step S104: Determine a mapping relationship between the multiple metadata contents and the multiple business terms according to the analysis result.
[0030] Step S105: Generate a knowledge graph of the multiple business terms based on the semantic knowledge base and the mapping relationship.
[0031] For step S101, the business is, for example, the operation activities, tasks, etc. of an enterprise, organization or industry. During the daily operation or task execution of an enterprise, organization or industry, business data (such as tables, conversation records, etc.) will be generated, and the enterprise or organization stores these business data in at least one type of data source such as a database, data lake, etc.
[0032] In some embodiments of the present disclosure, each data source corresponds to metadata content to describe the data in the data source. The metadata content includes, for example, data that describes the data in the data source, which provides descriptions and definitions of the database structure, content, and operations to better understand and manage the data source. The metadata content includes, for example, table structures, database object relationships, data item definitions, index information, and permissions. For example, table structures include table names, column names, data types, primary keys, foreign keys, constraints, etc. Taking a "student" table as an example, the metadata content will record that the table name is "student", and it contains column information such as "student ID" (data type is integer, which is the primary key), "name" (data type is string), "age" (data type is integer), etc. Database object relationships describe the relationships between tables, such as association relationships (implemented through foreign keys), inheritance relationships, etc. For example, the "student" table and the "course" table establish a many-to-many association relationship through the "course selection" table, and the metadata content will record this relationship information. Data dictionary data item definitions are used to elaborate on the meaning, source, value range, etc. of each data item in the database. For example, the data item definition of the "gender" column in the "student" table may be that the value is "male" or "female", representing the gender information of the student. Data source descriptions are used to record the generation methods, collection channels, etc. of the data. For example, the "grade" data in the "student" table may be entered through an exam system, and the metadata content will record its source as the exam system. Index information records, for example, the indexes created in the database, including index names, index columns, index types, etc. Indexes can improve data query efficiency, and the index information in the metadata content helps the database management system optimize the query execution plan. Permissions are used to define the access permissions of different users or user roles to database objects, such as query, insert, update, delete, etc. permissions. For example, the "student" role may only have the permission to query the "student" table, while the "administrator" role has full control permissions over all tables.
[0033] For example, a big data platform is used to obtain data sources. A big data platform refers to a system or platform that can store, process, and analyze large amounts of data. These platforms usually contain multiple data sources, such as relational databases, non-relational databases, log files, social media data, etc., and provide rich data processing and analysis tools. The data structures of multiple data sources vary greatly, the quality is uneven, they are isolated from each other, the data is scattered, and it is very difficult to establish global associations and deeply mine the data value contained therein.
[0034] For example, configure the data source information that needs to be connected on the big data platform, including local databases, data warehouses, data lakes, or data application APIs. When configuring, it is necessary to ensure correct data access permissions; the domain or industry type can also be specified during configuration to formulate a business ontology.
[0035] For step S102, in some embodiments of the present disclosure, the data processing model is, for example, obtained by training multiple metadata contents using a pre-training engine.
[0036] In some embodiments of the present disclosure, in the process of training the data processing model using the pre-training engine, it may include not only multiple metadata contents in the database, but also private domain data in enterprises, organizations, or industries to construct a private domain data processing model.
[0037] In some embodiments of the present disclosure, the active metadata content technology can be adopted to transform from a passive data governance process to automatic metadata content collection and semantic vectorization analysis of metadata content, so as to generate context association relationships between different fields and data sources, unify the consistency between multi-source heterogeneous data, and through the automatically generated data processing model, unify the differences between metadata content and semantics, match it with the user's intention, and adapt to the different expressions of different business departments and personnel and provide accurate answers.
[0038] For example, automatically collect metadata content, align data fields, tables, and business logics from different sources through the data processing model, and integrate them into a unified semantic knowledge base, that is, the unique business ontology of the enterprise. The semantic knowledge base includes, for example, multiple business terms and the relationships between multiple business terms, etc. In some embodiments of the present disclosure, the semantic knowledge base can be represented in the form of a knowledge graph.
[0039] The data processing model is, for example, a model with powerful computing capabilities and high intelligence constructed using advanced artificial intelligence technologies such as deep learning. The data processing model usually contains a large number of parameters and complex neural network structures, can process and analyze a large amount of data, and has excellent capabilities in natural language processing, image recognition, speech recognition, etc.
[0040] For step S103, for example, in a big data platform, data usually contains a large amount of information, but not all information is useful. To obtain metadata content, we need to identify key information from this data. This usually involves technologies such as natural language processing and information extraction. Large model technologies can utilize these technologies to automatically extract useful information from data, such as entities, attributes, relationships, etc. Metadata content is information that describes aspects such as the structure, meaning, source, and quality of data. Step S103, for example, includes data preprocessing, semantic analysis, key information extraction, etc. The main purpose of data preprocessing is to clean data, unify formats, remove redundancy and outliers, and perform possible data conversions to ensure the accuracy and efficiency of subsequent processing steps. Semantic analysis can, for example, be to use natural language processing algorithms or large models to understand and analyze the metadata content in the database. Key information extraction can, for example, adopt natural language processing algorithms and models, such as deep learning models, neural network models, etc. Using natural language processing algorithms and models, entity information is identified and extracted from the preprocessed data, the relationship information between entities is identified and extracted, and by understanding the semantic structure in the preprocessed data, event information in the preprocessed data is identified and extracted. The entity information, blood relationship, association relationship, and event information constitute key information.
[0041] For step S104, for example, by analyzing the semantics of the metadata content and the meaning of business terms, it is judged whether the two are semantically equivalent or related, thereby suggesting a mapping. Another example is that when the name or definition of the metadata content is exactly the same as the business term, a mapping relationship can be directly established. Another example is that if there is a synonymous expression of the metadata content with the business term, a mapping relationship is established. For example, the "ID number" field in the metadata content can be mapped to the "citizen identification number" in the business term table because they represent the same concept.
[0042] For example, automatically mark and align tables and column names, and associate them with semanticized business terms. Mark duplicates, errors, and sensitive information to ensure the consistency of business terms. Provide clear and standardized business language to simplify the data access operation curve and facilitate use by non-technical personnel.
[0043] For step S105, after suggesting the mapping relationship between business terms and metadata content, and determining the blood relationship, association relationship of the metadata content, and the association relationship between business terms based on the metadata content, a knowledge graph is constructed. The knowledge graph includes an entity set and an association set used to represent the relationships between each entity in the entity set. The elements in the entity set include multiple business terms, business rules, etc.
[0044] This method converts technical metadata content into a fully business-oriented knowledge graph, eliminates the inconsistent understanding of terms within an enterprise or organization, and enhances readability.
[0045] As Figure 1 shown, in addition to steps S101 to S105, this construction method further includes step S106.
[0046] Step S106: Monitor the data set. In response to obtaining that multiple metadata contents have changed, retrain the data processing model using the training engine and the changed multiple metadata contents, and update the semantic knowledge base.
[0047] In some embodiments of the present disclosure, the training engine may include, for example, deep learning frameworks such as TensorFlow, PyTorch, Keras, etc. or machine learning libraries such as Weka. In some embodiments of the present disclosure, a suitable data processing model (also referred to as a "training algorithm") can be selected according to the type of metadata content and the goal of the semantic knowledge base, such as including word vector models, deep learning models (such as convolutional neural networks, recurrent neural networks, Transformers), clustering algorithms, classification algorithms, etc.
[0048] After obtaining that the metadata content of the data source has changed, retrain the data processing model using the updated metadata content to obtain an updated semantic knowledge base.
[0049] On the premise of the rapid development of artificial intelligence, enterprises are no longer satisfied with the long data preparation process and traditional descriptive statistical analysis, and are eager to obtain more forward-looking and creative insights from data. In the field of product R & D, simply relying on data warehouses and traditional data analysis systems can no longer meet the requirements. Some embodiments of the present disclosure provide a data processing method, which can simplify the data usage process through generative AI, correctly and intelligently identify the user's query intention, and improve the efficiency of data processing by building data trust and business semantic understanding.
[0050] This data processing method includes: obtaining a data processing request; querying the knowledge graph based on the data processing request to obtain query information; and obtaining response data for responding to the data processing request based on the query information. The knowledge graph is based on Figure 1Constructed by the described method. That is, constructed according to the following operations: obtaining a data set of the service, the data set including at least one type of data source and multiple metadata contents corresponding to at least one type of data source; based on the data set and a pre-trained data processing model, obtaining a semantic knowledge base of the service, the semantic knowledge base including multiple service terms; analyzing the data set to obtain an analysis result; according to the analysis result, determining a mapping relationship between the multiple metadata contents and the multiple service terms; and based on the semantic knowledge base and the mapping relationship, generating a knowledge graph of the multiple service terms. The following combines Figure 2 to illustrate the data processing method provided by some embodiments of the present disclosure.
[0051] Figure 2 FIG. shows a flowchart of a data processing method provided by at least some embodiments of the present disclosure.
[0052] As Figure 2 shown, the data processing method includes steps S201 to S203.
[0053] Step S201: Obtain a data processing request.
[0054] Step S202: Based on the data processing request, query the knowledge graph to obtain query information.
[0055] Step S203: Based on the query information, obtain response data for responding to the data processing request.
[0056] In step S202, the knowledge graph is constructed according to Figure 1 the described method.
[0057] This data processing method constructs an automated training-generated semantic knowledge graph based on metadata content, which can avoid the problem of large language model hallucinations. The questions, context, and intentions of user queries will all be accurately interpreted and executed, providing users with accurate answers and a display of the solution process, at least partially alleviating the hallucination problem of current generative AI.
[0058] For step S201, the data processing request is, for example, a query request put forward by the user in natural language form. For example, the data processing request is: How is Zhang San's comprehensive performance since enrollment and whether there is a bias in subjects.
[0059] For step S202, for example, after obtaining a query request put forward in natural language form, perform semantic analysis on the query request to parse out the user's query intention. For example, use a large language model (LLM) to understand and analyze the query request to obtain a semantic analysis result, and then query the knowledge graph according to the semantic analysis result to obtain query information.
[0060] In some embodiments of the present disclosure, in addition to querying the knowledge graph to obtain query information, the knowledge graph and the vector library can also be combined to obtain query information. The query information includes: the database, data table, and fields where the response data is located. The response data refers to the data required to process the response data request. For example, the response data includes Zhang San's scores in various subjects, award records, etc., then the query information may include the data tables where the scores in various subjects are located, the score table for the first year of high school, the score table for the second year of high school, and the score table for the third year of high school, etc.
[0061] For example, the data processing method further includes obtaining a vector library based on a data set and a data processing model, and the vector library includes vectors of respective multiple metadata contents. For example, a semantic vector library is obtained through pre-training based on multiple metadata contents and processing by a data processing model, and the semantic vector library includes vectors of respective multiple metadata contents. Step S202 includes: based on the data processing request, querying the knowledge graph and the vector library to obtain query information. For example, a data processing model is used to preprocess, extract features, and generate semantic vectors for multiple metadata contents in the data source to obtain a semantic vector library. Combining the semantic vector library can supplement the knowledge graph. For example, in the case where a certain business term in the semantic analysis result is lacking in the knowledge graph, the vector library can be relied on to identify it through context semantic matching. For example, for low-frequency fields (such as "CLV"), the knowledge graph may lack synonyms. Relying on the vector library to determine the high similarity between the description vectors of "customer value" and "CLV" through context semantic matching, then CLV is identified as customer value.
[0062] In some embodiments of the present disclosure, step S202 includes: based on the data processing request, querying the knowledge graph to obtain the target business term of the response data; determining the target metadata content corresponding to the target business term based on the mapping relationship; and determining the query information based on the target metadata content and the vector library.
[0063] In some embodiments of the present disclosure, the vector library technology and the graphic library technology are combined. The database schema (Schema), library, table, and field contents in the metadata content are stored in the vector library, and the terms, attributes, metrics, analyses, labels, etc. in the business term library are stored in the knowledge graphic library. A separate node is added to the knowledge graphic library to store the mapping relationship between the business term library and the metadata content, and the knowledge graph is stored in the knowledge graphic library.
[0064] For example, after obtaining the semantic analysis result of the data processing request, query the knowledge graph to obtain relevant target business terms, then obtain the target metadata content corresponding to the target business terms according to the mapping relationship, and then determine the table name, field name, field type, and the relationship between tables where the target metadata content is located in the vector library. By querying the metadata content knowledge graph, the system can obtain table and field information related to the query request, including the relationship between tables (such as foreign key association), the data type of fields, etc. The query information includes: the database, data table, and fields where the response data is located.
[0065] In some embodiments of the present disclosure, step S203 includes: based on the query information, using a language model to construct a database query statement for the query information; and using the database query statement to query the data source to obtain response data.
[0066] For example, input the obtained query information into a language model (LLM) for analysis and processing. The language model is a trained deep learning model that can understand complex query requests and construct corresponding database query statements (Structured Query Language, SQL) according to the metadata content (such as table name, field name, data type, etc.). The SQL query statement may include, for example, a Select clause, a From clause, a Where clause, etc. At the same time, the large model will also consider the optimization of the query, such as selecting appropriate indexes and avoiding unnecessary full table scans. The analysis and processing process of the language model may include understanding the context of the query, determining the logical structure of the query, and selecting appropriate SQL functions and operators. The SQL query statement can accurately reflect the user's query request and can be executed in the database to obtain the required data. In this embodiment, combined with the powerful generation ability of the language model, it can handle complex query requests, including multi-table queries, nested queries, etc. The language model can generate a SQL query statement that conforms to the syntax specification and is logically correct based on the query intent, information in the knowledge graph, and vector library.
[0067] As Figure 2 shown, the data processing method may further include: using a language model to analyze the response data to generate a response report.
[0068] For example, the language model analyzes the response data and integrates the response data to obtain a response report. For example, for the above example, the data processing request is: How is Zhang San's comprehensive performance since enrollment and whether he is partial to certain subjects. Then, integrate and analyze the response data (Zhang San's scores in each subject, award records, etc.) to finally obtain a response report. The response report may include, for example, a summary description of the comprehensive performance and a conclusion on whether there is a partiality. The response may also include a summary and details of the scores for each year of the three years of high school.
[0069] In some embodiments of the present disclosure, the processing method further includes generating a business glossary based on a semantic knowledge base. The business glossary is, for example, displayed in a user interface for the user to view and use. For example, the business glossary can be displayed in response to an operation by the user in the user interface.
[0070] In some embodiments of the present disclosure, the data processing method further includes: displaying a knowledge graph in a user interface. Displaying the knowledge graph on the user interface facilitates data exploration by the user.
[0071] Figure 3 Shows an architecture diagram of a generative AI for data processing provided by at least one embodiment of the present disclosure.
[0072] As Figure 3 shown, the architecture diagram of the generative AI includes various types of data sources 301, a data processing model 302, a metadata activation unit 303, a semantic embedding and knowledge graph unit 304, and a generative artificial intelligence unit 305.
[0073] As Figure 3 shown, various types of data sources 301 may include various databases connected within the domain, company, or organization. For example, Mysql, Postgresql, the distributed data warehouse tool Hive, Oracle database, Starrocks database, and the data lake Iceberg.
[0074] The metadata activation unit 303 collects and analyzes the metadata content in the data source 301, and performs data lineage and correlation analysis on the metadata content.
[0075] The data processing model 302 analyzes and processes the metadata content of various types of data sources 301 to obtain a semantic knowledge base, and generates a knowledge graph in combination with the data lineage and correlation analysis of the metadata content obtained by the metadata activation unit 303.
[0076] The semantic embedding and knowledge graph unit 304 creates a knowledge graph library and a vector library by combining business terms and metadata content, and provides functions such as term library management, metric and analysis library management, data lineage tracing, data analysis, and insights.
[0077] The generative artificial intelligence unit 305 uses a language model to empower analysis, decision-making, and recommendations by integrating semanticized data.
[0078] Figure 4 Shows a method flow diagram for data processing provided by at least one embodiment of the present disclosure.
[0079] As Figure 4 shown, the method includes a preparation stage 401, an automated processing stage 402, and an application stage 403.
[0080] Preparation stage 401, for example, obtaining data sources at the user layer. The data sources include, for example, not only various databases but also business records, conversations, etc. At the service layer, the data sources at the user layer are automatically collected, and the collected data sources are identified and analyzed using a data processing model to obtain a semantic knowledge base.
[0081] Automated processing stage 402, after the data processing model processes the semantic knowledge base, the enhanced artificial intelligence governance unit performs business logic alignment, generates a unified thesaurus, generates lineage and knowledge graphs, performs metrics, analysis and extraction, maps the metadata content to the semantic knowledge base spectrum, performs similarity detection and merging, deduplication, performs metrics, analysis and extraction, extracts lineage relationships, extracts association relationships, and finally converts the technical metadata content into a completely business-oriented knowledge base, eliminating the problem of inconsistent understanding of terms among different departments within the enterprise.
[0082] The enhanced artificial intelligence governance unit converts the metadata content and the semantic knowledge base into a semantic vector library and a knowledge graph library. For the enhanced artificial intelligence governance unit, please refer to Figure 6 the description of
[0083] At the user layer, a business glossary can be constructed based on the knowledge base, a data exploration function can be provided, and metrics and analysis libraries can be constructed.
[0084] Application stage 403, at the user layer, an input box for the user to enter questions can be provided. In response to the user entering a question in the input box, the LLM performs semantic understanding and context analysis on the question. After obtaining the semantic analysis result, the LLM queries the semantic vector library and the knowledge graph library to determine which data table and which field the response data for answering the question is stored in. After determining the location of the response data, the LLM uses the natural language to SQL conversion technology (Natural Language to SQL, NL2SQL) to generate a database query statement.
[0085] After querying resources such as databases and data lakes using the database query statement to obtain the response data, the LLM generates an answer (i.e., a response report) based on the response data.
[0086] Figure 5 FIG. shows an architecture diagram of a data processing model provided by at least one embodiment of the present disclosure.
[0087] As Figure 5 shown, the data processing model 500 automatically discovers metadata content from data sources and uses a proprietary pre-training engine 501 to train the data processing model 504 to obtain a semantic knowledge base 503.
[0088] In this architecture, for example, a metadata monitor 502 may also be included, configured to monitor in real time whether the metadata content changes. If the metadata content changes, the metadata content is retrained to obtain an updated data processing model 504.
[0089] Figure 6 The figure shows an architecture diagram of an enhanced artificial intelligence governance unit provided by at least one embodiment of the present disclosure.
[0090] As Figure 6 shown, the enhanced artificial intelligence governance unit 600 maps the business terms in the semantic knowledge base 503 to the metadata content. The mapping relationship between the metadata content and the business terms is obtained by analyzing the metadata content directory and the lineage of the metadata content. After performing the semantic mapping, operations such as deduplication, semantic merging, and metric extraction are performed, and finally a semantic knowledge graph 601 is obtained.
[0091] Figure 7 The figure shows a schematic block diagram of a knowledge graph construction device 700 provided by at least one embodiment of the present disclosure.
[0092] For example, as Figure 7 shown, the knowledge graph construction device 700 includes a data source acquisition unit 710, a semantic knowledge base generation unit 720, a metadata content analysis unit 730, a determination unit 740, and a knowledge graph generation unit 750.
[0093] The data source acquisition unit 710 is configured to acquire a business data set, and the data set includes at least one type of data source and the metadata content corresponding to at least one type of data source. The data source acquisition unit 710, for example, executes Figure 1 step S101 in
[0094] The semantic knowledge base generation unit 720 is configured to obtain a business semantic knowledge base based on the data set and a pre-trained data processing model. The semantic knowledge base includes multiple business terms. The semantic knowledge base generation unit 720, for example, executes Figure 1 step S102 in
[0095] The metadata content analysis unit 730 is configured to analyze the data set to obtain an analysis result. The metadata content analysis unit 730, for example, executes Figure 1 step S103 in
[0096] The determination unit 740 is configured to determine the mapping relationship between the metadata content and multiple business terms according to the analysis result. The determination unit 740, for example, executes Figure 1 step S104 in
[0097] The knowledge graph generation unit 750 is configured to generate a knowledge graph of multiple business terms based on a semantic knowledge base and mapping relationships. The knowledge graph generation unit 750, for example, executes Figure 1 step S105 in
[0098] Figure 8 FIG. 6 shows a schematic block diagram of a data processing device 800 provided by at least one embodiment of the present disclosure.
[0099] For example, as Figure 8 shown, the data processing device 800 includes a request acquisition unit 810, a query unit 820, and a response data acquisition unit 830.
[0100] The request acquisition unit 810 is configured to acquire a data processing request. The request acquisition unit 810, for example, executes Figure 2 step S201 described in
[0101] The query unit 820 is configured to query the knowledge graph based on the data processing request to obtain query information. The knowledge graph is constructed according to the construction method provided by any embodiment of the present disclosure. The query unit 820, for example, executes Figure 2 step S202 described in
[0102] The response data acquisition unit 830 is configured to acquire response data for responding to the data processing request based on the query information. The response data acquisition unit 830, for example, executes Figure 2 step S203 described in
[0103] For example, the data source acquisition unit 710, the semantic knowledge base generation unit 720, the metadata content analysis unit 730, the determination unit 740, and the knowledge graph generation unit 750, the request acquisition unit 810, the query unit 820, and the response data acquisition unit 830 can be hardware, software, firmware, and any feasible combination thereof. For example, the data source acquisition unit 710, the semantic knowledge base generation unit 720, the metadata content analysis unit 730, the determination unit 740, and the knowledge graph generation unit 750, the request acquisition unit 810, the query unit 820, and the response data acquisition unit 830 can be dedicated or general-purpose circuits, chips, or devices, etc., and can also be a combination of a processor and a memory. Regarding the specific implementation forms of the above-mentioned various units, the embodiments of the present disclosure do not limit this.
[0104] In some embodiments of the present disclosure, generative AI is used to simplify the data usage process, correctly and intelligently identify the user's query intention, and improve work efficiency by building data trust and business semantic understanding. For example, when answering questions, generative AI provides users with high-quality analysis and decision-making answers after precise summarization and refinement based on real and trustworthy data contexts and industry knowledge bases.
[0105] In some embodiments of the present disclosure, while automatically discovering and annotating metadata content, a semantic knowledge graph is automatically constructed based on the metadata content, aligning the metadata content structure with the business semantic background. This avoids the centralized storage and calculation of data, reducing the consumption of hardware resources and the time cost of developers. For example, aligning the data logic of an enterprise with unified business terms, such as associating the aggregation segments in SQL with common aggregation metrics, simplifies the identification and processing of metrics during data analysis.
[0106] In some embodiments of the present disclosure, a data processing model dynamic update mechanism is adopted to adjust model parameters in real time according to business requirements without manual intervention, reducing the usage cost, supporting more flexible expansion of business requirements, and synchronously updating the dynamically changing metadata content and the semantic knowledge graph without spending a large amount of time and resources on manual configuration or adjustment, greatly shortening the window period from data change to analysis and decision-making. For example, automatically discovering and mapping all structured data sources from local and data warehouses without migrating the data from the original sources, and automatically identifying information such as collection modes, tables, fields, descriptions, and association relationships.
[0107] The business ontology knowledge base created through semantic embedding technology is like a translator that matches the user's questions with definite and governed answers. It can understand and parse natural business language and user intentions, and align them with specific organized logic and business background, at least partially avoiding the generation of hallucinations and increasing credibility.
[0108] It should be noted that in the embodiments of the present disclosure, each unit of the knowledge graph construction device 700 corresponds to each step of the foregoing construction method. For the specific functions of the knowledge graph construction device 700, reference can be made to the relevant descriptions of the construction method, which will not be elaborated here. Figure 7 The components and structures of the shown knowledge graph construction device 700 are exemplary rather than restrictive. According to needs, the knowledge graph construction device 700 may further include other components and structures. In the embodiments of the present disclosure, each unit of the data processing device 800 corresponds to each step of the foregoing construction method. For the specific functions of the data processing device 800, reference can be made to the relevant descriptions of the construction method, which will not be elaborated here. Figure 8 The components and structures of the shown data processing device 800 are exemplary rather than restrictive. According to needs, the data processing device 800 may further include other components and structures.
[0109] At least one embodiment of the present disclosure also provides an electronic device, which includes a processor and a memory. The memory includes one or more computer program instructions. When the one or more computer program instructions are run by the processor, they execute the method provided in any embodiment of the present disclosure. This electronic device can automatically construct a knowledge graph based on metadata content, unify the consistency between multi-source heterogeneous data, unify the metadata content structure and business terms, and reduce the consumption of hardware resources and time costs. By using generative AI to simplify the data usage process, correctly and intelligently identify the user's query intention, and provide work efficiency by building data trust and business semantic understanding.
[0110] Figure 9 It is a schematic block diagram of an electronic device provided in some embodiments of the present disclosure. As Figure 9 shown, the electronic device 900 includes a processor 910 and a memory 920. The memory 920 is used to store non-transitory computer-readable instructions (such as one or more computer program modules). The processor 910 is used to run the non-transitory computer-readable instructions, and when the non-transitory computer-readable instructions are run by the processor 910, they can execute one or more steps in the above-mentioned method. The memory 920 and the processor 910 can be interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0111] For example, the processor 910 can be a central processing unit (CPU), a graphics processing unit (GPU), or other forms of processing units with data processing capabilities and / or program execution capabilities. For example, the central processing unit (CPU) can be of the X86 or ARM architecture, etc. The processor 910 can be a general-purpose processor or a dedicated processor, and can control other components in the electronic device 900 to perform desired functions.
[0112] For example, the memory 920 can include any combination of one or more computer program products. The computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory can include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer program modules can be stored on the computer-readable storage media, and the processor 910 can run one or more computer program modules to implement various functions of the electronic device 900. Various application programs and various data, as well as various data used and / or generated by the application programs, can also be stored in the computer-readable storage media.
[0113] It should be noted that in the embodiments of the present disclosure, for the specific functions and technical effects of the electronic device 900, reference may be made to the description of the above method in the foregoing text, which will not be elaborated herein.
[0114] Figure 10 FIG. is a schematic block diagram of another electronic device provided by some embodiments of the present disclosure. The electronic device 1000 is, for example, suitable for implementing the above method provided by the embodiments of the present disclosure. The electronic device 1000 may be a terminal device or the like. It should be noted that Figure 10 The illustrated electronic device 1000 is merely an example and will not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0115] As Figure 10 shown, the electronic device 1000 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 1010, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1020 or a program loaded from a storage device 1080 into a random access memory (RAM) 1030. In the RAM 1030, various programs and data required for the operation of the electronic device 1000 are also stored. The processing device 1010, the ROM 1020, and the RAM 1030 are connected to each other via a bus 1040. An input / output (I / O) interface 1050 is also connected to the bus 1040.
[0116] Generally, the following devices may be connected to the I / O interface 1050: an input device 1060 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1070 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1080 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1090. The communication device 1090 may allow the electronic device 1000 to communicate with other electronic devices wirelessly or wirelesly to exchange data. Although Figure 10 the illustrated electronic device 1000 has various devices, it should be understood that it is not required to implement or include all the illustrated devices, and the electronic device 1000 may alternatively implement or include more or fewer devices.
[0117] For example, according to an embodiment of the present disclosure, the above method can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for executing the above method. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 1090, or installed from a storage device 1080, or installed from a ROM 1020. When the computer program is executed by a processing device 1010, the functions defined in the method provided by the embodiment of the present disclosure can be implemented.
[0118] At least one embodiment of the present disclosure further provides a computer-readable storage medium, which is used to store non-temporary computer-readable instructions, and when the non-temporary computer-readable instructions are executed by a computer, the above method can be implemented. Using this computer-readable storage medium, the medium can automatically construct a knowledge graph based on metadata content, unify the consistency between multi-source heterogeneous data, unify the metadata content structure and business terms, reduce hardware resource consumption and time costs. By generative AI to simplify the data usage process, correctly and intelligently identify the user's query intention, and through building data trust and business semantic understanding, work efficiency can be provided.
[0119] At least some embodiments of the present disclosure further provide a non-transitory storage medium. Figure 11 Schematically shows a schematic diagram of a computer-readable storage medium provided by an embodiment of the present disclosure. For example, as Figure 11 shown, the storage medium 1100 stores non-temporary computer-readable instructions 1101, and when the non-temporary computer-readable instructions 1101 are executed by a computer (including a processor), the method provided by any embodiment of the present disclosure can be executed. This method can automatically construct a knowledge graph based on metadata content, unify the consistency between multi-source heterogeneous data, unify the metadata content structure and business terms, reduce hardware resource consumption and time costs. By generative AI to simplify the data usage process, correctly and intelligently identify the user's query intention, and through building data trust and business semantic understanding, work efficiency can be provided.
[0120] For example, one or more computer instructions can be stored on the storage medium 1100. Some of the computer instructions stored on the storage medium 1100 can be, for example, instructions for implementing one or more steps in the above method.
[0121] For example, the storage medium may include the storage component of a tablet computer, the hard disk of a personal computer, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a compact disc read-only memory (CD-ROM), a flash memory, or any combination of the above storage media, and may also be other applicable storage media.
[0122] The technical effects of the storage medium provided by the embodiments of the present disclosure can refer to the corresponding descriptions of the processing methods in the above embodiments, and will not be elaborated here.
[0123] Regarding the present disclosure, the following points need to be noted: (1) In the accompanying drawings of the embodiments of the present disclosure, only the structures related to the embodiments of the present disclosure are involved, and other structures can refer to the general design.
[0124] (2) Without conflict, the features in the same embodiment and different embodiments of the present disclosure can be combined with each other.
[0125] The above are only the specific implementation manners of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A method for constructing a knowledge graph, comprising: Acquire a data set of a business, wherein the data set includes at least one type of data source and a plurality of metadata contents corresponding to the at least one type of data source; Based on the data set and the pre-trained data processing model, a semantic knowledge base of the business is obtained, wherein the semantic knowledge base includes a plurality of business terms; Analyzing the data set to obtain analysis results; Determining mapping relationships between the plurality of metadata contents and the plurality of business terms according to the analysis result; and Based on the semantic knowledge base and the mapping relationship, a knowledge graph of the multiple business terms is generated.
2. The construction method according to claim 1, further comprising: The data set is monitored, and in response to obtaining changes in the multiple metadata contents, the data processing model is retrained using a training engine and the changed multiple metadata contents, and the semantic knowledge base is updated.
3. A data processing method, comprising: Get data processing request; Based on the data processing request, query the knowledge graph to obtain query information; as well as Based on the query information, obtaining response data for responding to the data processing request, The knowledge graph is constructed according to the following operations: Acquire a data set of a business, wherein the data set includes at least one type of data source and a plurality of metadata contents corresponding to the at least one type of data source; Based on the data set and the pre-trained data processing model, a semantic knowledge base of the business is obtained, wherein the semantic knowledge base includes a plurality of business terms; Analyzing the data set to obtain analysis results; and Determining mapping relationships between the plurality of metadata contents and the plurality of business terms according to the analysis result; and Based on the semantic knowledge base and the mapping relationship, a knowledge graph of the multiple business terms is generated.
4. The method according to claim 3, further comprising: A vector library is obtained based on the data set and the data processing model, wherein the vector library includes vectors of the plurality of metadata contents respectively. Wherein, based on the data processing request, querying the knowledge graph to obtain query information includes: Based on the data processing request, the knowledge graph and the vector library are queried to obtain the query information.
5. The method according to claim 4, wherein: Based on the data processing request, query the knowledge graph and the vector library to obtain the query information, including: Based on the data processing request, query the knowledge graph to obtain the target business term of the response data; Based on the mapping relationship, determining target metadata content corresponding to the target business term; and The query information is determined based on the target metadata content and the vector library.
6. The method according to claim 5, wherein: The query information includes: the database, data table and field where the response data is located.
7. The method according to claim 4, wherein: Acquiring the response data for responding to the data processing request based on the query information includes: Based on the query information, construct a database query statement of the query information using a language model; and The at least one type of data source is queried using the database query statement to obtain the response data.
8. The method according to claim 3, further comprising: The response data is analyzed using a language model to generate a response report.
9. The method according to claim 3, further comprising: The data set is monitored, and in response to obtaining changes in the multiple metadata contents, the data processing model is retrained using a training engine and the changed multiple metadata contents, and the semantic knowledge base is updated.
10. The method according to claim 3, further comprising: Based on the semantic knowledge base, a business term list is generated.
11. The method according to claim 3, further comprising: Display the knowledge graph on a user interface.
12. A knowledge graph construction device, comprising: A data source acquisition unit, configured to acquire a data set of a business, the data set including at least one type of data source and a plurality of metadata contents of the at least one type of data source; A semantic knowledge base generating unit, configured to obtain a semantic knowledge base of the business based on the data set and a pre-trained data processing model, wherein the semantic knowledge base includes a plurality of business terms; A metadata content analysis unit, configured to analyze the data set to obtain an analysis result; a determining unit configured to determine a mapping relationship between the plurality of metadata contents and the plurality of business terms according to the analysis result; and The knowledge graph generating unit is configured to generate a knowledge graph of the plurality of business terms based on the semantic knowledge base and the mapping relationship.
13. A data processing device, comprising: a request acquisition unit configured to acquire a data processing request; A query unit, configured to query the knowledge graph to obtain query information based on the data processing request; as well as a response data acquisition unit, configured to acquire response data for responding to the data processing request based on the query information, Wherein, the knowledge graph is constructed according to the construction method according to any one of claims 1-2.
14. An electronic device comprising: processor; as well as a memory including one or more computer program instructions; The one or more computer program instructions are executed by the processor to perform the method according to any one of claims 1 to 11.
15. A computer-readable storage medium non-transitorily storing computer-readable instructions, wherein: When the computer-readable instructions are executed by a processor, the method according to any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
Data analysis method and system and related equipment
CN117668242A
Knowledge graph construction method and knowledge graph application method
CN118069856A
Generative AI large language model-based knowledge base construction method, system and equipment
CN118885465A
Multi-table SQL (Structured Query Language) generation method and device based on large model and metadata knowledge graph
CN119719145A
Energy information query method, system and equipment based on large language model and medium
CN119719274A
Cited By
Data consanguinity analysis method of e-commerce data warehouse, product, equipment and medium
CN120725717A
Data lineage analysis method, product, device and medium of e-commerce data warehouse
CN120725717B