Data asset management method and system and related equipment

By establishing a knowledge graph, the problem of enterprise users needing to access multiple systems in a chain is solved, thus improving data reading efficiency and user experience.

CN120950528APending Publication Date: 2025-11-14HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410601037.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-14
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Enterprise users need to access multiple business systems in a chain to obtain data, resulting in low data reading efficiency.

Method used

By establishing a knowledge graph-based data asset management system, the association and querying of data from multiple business systems can be realized, and users can get the answer simply by entering a question.

Benefits of technology

It improves data reading efficiency, reduces user operations and time costs, and enhances user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950528A_ABST
    Figure CN120950528A_ABST
Patent Text Reader

Abstract

The invention provides a data asset management method and system and related equipment, and the method comprises the following steps: a data asset management system receives a question sentence which is input by a user and is expressed in a natural language, determines a question sentence entity contained in the question sentence, carries out the retrieval in a knowledge graph based on the question sentence entity, obtains a query result, and stores the query result in a database; the knowledge graph in the data asset management system is established based on the business data in the plurality of business systems, and the business data comprises structured data, semi-structured data and unstructured data, so that the knowledge graph is established based on the business data in the plurality of business systems; according to the method, information of different departments and different data sources can be correlated and inquired, a user does not need to have chained access to multiple service systems to obtain answers, the answers can be obtained only by inputting simple questions, and the use experience of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of big data, and in particular to a data asset management method, system and related equipment. Background Technology

[0002] In the era of big data, data has become one of the most important assets for enterprises. Companies store and manage data from various sources to support their analysis and decision-making. These data assets from different sources are often stored in different systems; for example, the supply chain department uses a supply chain system, the finance department uses a finance system, and the marketing department uses a marketing system.

[0003] However, different departments using their own systems to manage specific business functions can lead to users needing to chain through multiple different business systems to obtain the data they need. For example, if a sales manager wants to understand the sales and inventory of a product, they need to first access the supply chain system to view the inventory report and understand the product's inventory level, and then access the financial system to view the sales records and understand the product's sales performance. If the sales manager needs more data, they may need to chain through even more business systems, increasing the user's workload and time costs, resulting in low data retrieval efficiency. Summary of the Invention

[0004] This application provides a data asset management method, system, and related equipment to solve the problem of low data reading efficiency caused by business personnel needing to access multiple business systems in a chain when reading data.

[0005] Firstly, a data asset management method is provided, which is applied to a data asset management system that establishes communication connections with multiple business systems. The method includes the following steps: receiving a query request input by a user, the query request including a question expressed in natural language; determining the question entities contained in the question and the question category of the question, the question category including data reading type and information query type; data reading type questions are used to read business data from multiple business systems, and information query type questions are used to obtain information query results; retrieving query results in a knowledge graph based on the question category and question entities; wherein the knowledge graph in the data asset management system is built based on business data, including structured data, semi-structured data, and unstructured data; and determining the answer corresponding to the question based on the query results.

[0006] Implementing the method described in the first aspect, the data asset management system establishes a knowledge graph based on business data from multiple business systems, enabling data from different business systems to be interconnected and queried. Thus, when a user sends a question to the data asset management system through a client, the system can search the knowledge graph based on the question entity contained in the question to obtain query results related to that question entity, and then generate the answer to the question based on the query results. This eliminates the need for users to access multiple business systems in a chain to obtain the answer; they only need to enter a simple question to get the answer, improving the user experience.

[0007] In one possible implementation, when retrieving query results in the knowledge graph based on question category and question entity, multiple question templates corresponding to the question category can be obtained first. The question is then matched with multiple question templates to determine the question template corresponding to the question. Based on the question template corresponding to the question, the query statement template corresponding to the question is obtained. Each question template corresponds to one query statement template. Then, the query statement template is filled based on the question entity to obtain the graph query statement corresponding to the question.

[0008] Furthermore, when the question type is a query type, a graph query statement is used to search for the business data corresponding to the question entity in the knowledge graph to obtain the query results; when the question type is an information query type, a graph query statement is used to search for the related entities, relationships, and attributes of the question entity in the knowledge graph to obtain the query results.

[0009] The above-described implementation method of this application can fulfill two different needs: data reading and information querying. For the field of enterprise data asset management, data reading and information querying are the primary needs of relevant business personnel. Data reading refers to retrieving data that meets certain conditions from raw business data based on user needs, such as extracting order records within a specific time period or extracting specific types of time records from log files. Information querying refers to analyzing and inferring information based on raw business data to obtain the information needed by the user. This information can be an analysis result or an inference result to meet the needs of daily business operations and decision-making, such as asset owner queries, organizational structure queries, personnel information queries, and equity structure queries. This type of information query does not require complete raw data; the system only needs to answer the question. Based on knowledge graphs, entities or relationships can be quickly located, supporting daily operational decisions. The data asset management method provided in this application can meet the above-mentioned data reading and information querying needs, and the implementation process is simple and fast. Users only need to input a simple question, and the retrieval efficiency based on knowledge graphs is also very high, improving the user experience.

[0010] In one possible implementation, structured data includes table data stored in a fixed format, semi-structured data includes non-table data stored in a fixed format, and unstructured data includes text data with an irregular format. Before receiving a query request from a user, the method further includes the following steps: obtaining business data from multiple business systems; obtaining first entity information and first relationship information based on the fixed format features of structured and semi-structured data; performing named entity recognition on entities in unstructured data based on semantic features of unstructured data to obtain second entity information and second relationship information; establishing nodes in a knowledge graph based on the first and second entity information; establishing edges between nodes in the knowledge graph based on the first and second relationship information; and establishing a knowledge graph.

[0011] The above implementation allows the data asset management system to extract entities from structured and semi-structured data in business systems based on their fixed format characteristics, thus building a knowledge graph. Similarly, unstructured data can be used to identify entities based on its semantic features, also building a knowledge graph. This results in a knowledge graph encompassing the knowledge features of structured, semi-structured, and unstructured data from multiple business systems. This provides users with more comprehensive answers, enabling not only the reading of business data from multiple systems but also complex information queries, thereby improving the work efficiency of business personnel.

[0012] In one possible implementation, when obtaining the first entity information and the first relationship information based on the fixed format characteristics of structured data, multiple entities of the structured data are determined based on the fields and field values ​​in the structured data, and the first entity information of the structured data is obtained. The multiple entities of the structured data include fields and field values ​​with keywords, fields and field values ​​with attribute types of preset types, and fields and field values ​​contained in the entity template. Based on the relationship between the fields and field values, and the relationship between the fields, the first relationship information between the multiple entities of the structured data is determined.

[0013] In practical implementation, fields or field values ​​that begin or end with specific keywords may represent specific types of entities. For example, a field starting with "product_" may represent a product entity. Alternatively, the entity represented by a field can be inferred from its data type. For instance, a string type may represent a name or description, a numeric type may represent quantity or value, and a date type may represent time or date. Alternatively, the entity represented by a field or field value can be determined based on expert knowledge. Field templates that could potentially serve as entities can be configured, and the fields or field values ​​in the table data that can function as entities can be determined based on these templates, thereby obtaining entity information. The above examples are for illustrative purposes only and are not intended to limit the scope of this application.

[0014] In practical implementation, correlation analysis can also be performed on fields that are not recorded as having a relationship, or between field values, to determine the relationships between entities. For example, analyzing the frequency of two fields appearing simultaneously in the data can help determine their relationship. It should be understood that if a product appears frequently in an order, it can be assumed that the product was likely purchased in that order. Alternatively, expert knowledge can be used to determine the relationships between entities. Association templates can be configured to identify potential relationships between entities, and the associated fields in the table data can be determined based on these templates to obtain the association information. The above examples are for illustrative purposes only and are not intended to limit the scope of this application.

[0015] The above implementation method obtains the first entity information and the first relationship information based on the fixed format characteristics of structured data. This method is simple, fast and effective, realizing the automated and systematic extraction of entity information and relationship information from structured data, reducing the complexity and difficulty of data processing, and improving the efficiency of knowledge graph construction.

[0016] In one possible implementation, when obtaining the first entity information and the first relationship information based on the fixed format characteristics of semi-structured data, for an XML file, firstly, based on the tags and attributes in the semi-structured data, multiple entities of the semi-structured data are determined to obtain the first entity information of the semi-structured data. Based on the relationship between tags and attributes and the nesting relationship between tags, the first relationship information between the multiple entities of the semi-structured data is determined. Here, the tags in the semi-structured data are used to separate the hierarchical structure and relationships of the data.

[0017] For example, XML uses angle brackets (e.g. <person>Elements are identified by angle brackets () and the names within these brackets are tags. XML tags can identify entities and attributes in data, as well as the relationships between them. For an XML file, you can search for the content between the start and end tags. Typically, the start tag indicates the beginning of an entity, and the end tag indicates the end. Therefore, by using the start and end tags, you can determine the entity corresponding to each tag and then identify the nesting relationships between tags. Outer tags represent more general, higher-level entities, while inner tags represent more specific, more detailed entities. Based on nesting relationships, you can generate association information between entities. Furthermore, you can also determine entities based on tag attributes; for example, you can determine the association information between entity 1 corresponding to a tag and entity 2 corresponding to an attribute. This method allows you to obtain entity relationship data corresponding to an XML file.

[0018] For JSON files, based on the key-value pairs in the semi-structured data, multiple entities within the semi-structured data are identified, obtaining the first entity information. Based on the correspondence between keys and values, and the nesting relationship between the first and second keys, the second relationship information between the multiple entities in the semi-structured data is determined. Here, the second key is a subkey of the first key, and the second key is an attribute or member of the object described by the first key. For example, in JSON, key-value pairs are used to represent data attributes and values; the key identifies the attribute of the entity, and the value represents the specific value.

[0019] The above implementation method is based on the fixed format characteristics of semi-structured data. For example, the fixed format characteristics of XML are that there are many tags, and the fixed format characteristics of JSON files are that there are many key-value pairs. By using these fixed formats, entity information and relationship information can be effectively extracted, reducing the complexity and difficulty of data processing, realizing the automated and systematic extraction of entity information and relationship information in semi-structured data, and improving the efficiency of knowledge graph construction.

[0020] In one possible implementation, when performing named entity recognition on entities in unstructured data based on the semantic features of unstructured data to obtain second entity information and second relation information, the entity recognition naming model is first used to perform named entity recognition on entities in unstructured data to obtain second entity information and second relation information. The entity recognition naming model includes a word embedding layer, a feature extraction layer, and a feature classification layer. The word embedding layer is used to convert unstructured data into a word vector matrix with word-level features. The feature extraction layer is used to extract semantic features from the word vector matrix to obtain a sentence vector matrix with sentence-level features. The feature classification layer is used to classify the word segments in unstructured data according to the sentence vector matrix to obtain second entity information and second relation information.

[0021] The above implementation captures the linguistic similarity and correlation between words through the word embedding layer, and extracts sentence-level semantic features through feature extraction to better represent the semantic and contextual information of the entire sentence. In this way, the word embedding layer and the feature extraction layer deeply mine the features in the text data, and the feature classification layer completes entity recognition and naming based on the deeply mined features, making the entity recognition and naming effect more accurate.

[0022] In one possible implementation, the word embedding layer includes a text input layer and a vector transformation layer. The text input layer is used to perform word segmentation on unstructured data to obtain multiple words. The vector transformation layer inputs the multiple embedding vectors of each word into the transformer structure to extract the word-level features of each word and obtain a word vector matrix. The multiple embedding vectors include word embedding vectors, sentence embedding vectors, and position embedding vectors.

[0023] Furthermore, the transformer structure comprises multiple transformer blocks. The output of each transformer block can be input into the next transformer block, allowing the model to process, learn, and represent the input data layer by layer, thereby better capturing the semantics and structure of the input data. Each transformer block includes a multi-head attention network and a feedforward neural network. The output of the multi-head attention network is input into the feedforward neural network. The multi-head attention network consists of multiple attention heads, each generating an attention weight matrix. These attention weights represent the importance of a position, guiding the subsequent feedforward neural network to prioritize certain positions when processing the data. Moreover, each attention head determines its attention weight matrix by calculating the similarity between that position and other positions. Therefore, each attention weight matrix also includes the relationship between the word segment and its adjacent segments, enabling the extraction of deep bidirectional semantic features of the word segment. Feedforward neural networks are used to perform non-linear transformations and feature extraction on word segmentation at each position, helping the model learn richer and more abstract feature representations, thereby better extracting the deep semantic features of text data. Combined with the attention weights output by the attention head, semantic features can be extracted with emphasis, so that the final output word vector matrix is ​​a word vector matrix containing deep features.

[0024] In the above implementation, the word vector matrix generated by the embedding layer is a vector representation generated after deep semantic feature extraction of unstructured data. This word vector matrix is ​​input into the subsequent feature extraction layer, which can extract richer and more accurate semantic information. Based on this, entity recognition and naming can be performed, making the final entity information and association information more accurate, and the knowledge graph established is also more accurate.

[0025] In one possible implementation, the feature extraction layer includes a context feature extraction network and a multi-head attention network. The context feature extraction network is used to extract context features from unstructured data based on the word vector matrix, generating a multi-dimensional feature sequence containing context features. The multi-head attention network is used to determine the attention score of each element in the multi-dimensional feature sequence to obtain a sentence vector matrix, where the attention score is used to indicate the degree of influence of each element on entity recognition and naming.

[0026] The multi-head attention network can include multiple attention heads, each of which is used to process a portion of the multidimensional feature sequence. In a specific implementation, the multidimensional feature sequence can be divided equally according to the number of attention heads to obtain multiple segmented data, and each segmented data is assigned to an attention head for processing.

[0027] Furthermore, each attention head can first pass through a linear transformation layer to obtain a linear transformation sequence corresponding to the segmented data, which is used for subsequent attention calculations. This linear transformation is typically implemented by multiplying a weight matrix. Based on the linear transformation sequence, the attention score corresponding to the segmented data is calculated, usually through a single scaling dot product calculation. This attention score includes the score for each position in the segmented data, used to measure the importance of each element at each position. Then, a scaling operation is performed on the attention score to stabilize the subsequent softmax operation. Scaling ensures numerical stability in the subsequent softmax calculation, preventing gradient explosion or vanishing problems. Next, a softmax operation is performed based on the scaled attention score, mapping the attention score to the range of 0-1 to obtain the attention weights corresponding to the segmented data. This is done so that the model can focus on different positions in the input data in a probability distribution. Finally, the attention weights are multiplied by the segmented data to obtain the attention sequence of the segmented data. The purpose of this step is to apply the attention weights to the multidimensional feature sequence generated by the upper and lower feature extraction networks, emphasizing the important parts in the multidimensional feature sequence. Such attention training participates in subsequent entity recognition and naming, which enables the model to complete the task more effectively and improves the performance of entity recognition and naming.

[0028] Furthermore, each attention head can obtain the attention sequence corresponding to the segmented data according to the above process. The attention sequences of multiple attention heads can be input into the fully connected layer to obtain sentence-level feature vectors. In the fully connected layer, the attention sequences obtained by multiple attention heads can be concatenated or weighted to obtain richer sentence representations and obtain the sentence vector matrix corresponding to the word vector matrix.

[0029] In the above implementation, the context feature extraction network in the feature extraction layer processes text data in both forward and backward directions, which can effectively capture the positional information of word segments in the sentence and the relationship between word segments. The multi-head attention network can identify the dependency relationship between different positions in the sentence and assign different attention weights to word segments at different positions in the sentence, which can help the model better understand the semantic relationship between entities in the sentence. Such sentence vector features not only include the semantic information of the sentence, but also the relationship features between word segments, so that the subsequent feature classification layer can well identify entities in the sentence and obtain entity information and relationship information.

[0030] In one possible implementation, the feature classification layer includes an annotation network and an optimization network. The annotation network is used to generate a label prediction result for each word based on the sentence vector matrix. The optimization network is used to adjust the label of each word in the unstructured data based on the label prediction result for each word, combined with sentence semantics and contextual information, to obtain second entity information and second relation information in the unstructured data. The label is used to indicate whether the word is an entity and the category to which the entity belongs.

[0031] Optionally, the annotation network can predict labels for each segment based on a preset labeling method (e.g., BIO (begin, inside, outside) labeling method). Labels can include B, I, and O tags. A segment labeled with a B tag indicates it is the beginning of an entity; a segment labeled with an I tag indicates it is a non-beginning part of an entity (the remaining part after removing the beginning); and a segment labeled with an O tag indicates it is not an entity. Furthermore, the labels also include the entity's category. Based on the category in the label, the entity can be named. For example, a B-per label indicates the beginning of an entity, with the type "person"; an I-per label indicates the non-beginning part of an entity, with the type "person"; a B-loc label indicates the beginning of an entity, with the type "location"; and an I-loc label indicates the non-beginning part of an entity, with the type "location". The above examples are for illustration only and are not intended to be specific. By expanding the labels, the entity parts and the entity type can be clearly identified.

[0032] The above implementation uses an annotation network to label each word to determine whether it is an entity. Then, an optimization network adjusts the label of each word based on the semantics of the entire sentence. Since the label prediction result of each word obtained by the annotation network is based on local features, a word may represent different entities in different contexts. For example, "ink" can represent writing ink or knowledge. If the sentence is "I go to buy ink", "ink" represents an object. If the sentence is "I have ink in my belly", "ink" represents knowledge. Therefore, label prediction based solely on the annotation network may result in errors. By using an optimization network to analyze the label prediction result of each word in the entire sentence, considering the contextual information of the entire sentence, and understanding the meaning and structure of the sentence, the entity recognition result obtained is more accurate.

[0033] Secondly, a data asset management system is provided. This system establishes communication connections with multiple business systems. The data asset management system includes: an acquisition unit for receiving query requests input by users, including questions expressed by users in natural language; a retrieval unit for determining the question entities contained in the question and the question category of the question, including data reading type and information query type. Data reading type questions are used to read business data from multiple business systems, while information query type questions are used to obtain information query results; a retrieval unit for performing retrieval in a knowledge graph based on the question category and question entities to obtain query results. The knowledge graph in the data asset management system is built based on business data, including structured data, semi-structured data, and unstructured data; and a generation unit for determining the answer corresponding to the question based on the query results.

[0034] Implementing the system described in the second aspect, the data asset management system establishes a knowledge graph based on business data from multiple business systems, enabling data from different business systems to be interconnected and queried. Thus, when a user sends a question to the data asset management system through a client, the system can search the knowledge graph based on the question entity contained in the question to obtain query results related to that question entity, and then generate the answer to the question based on the query results. This eliminates the need for users to access multiple business systems in a chain to obtain the answer; they only need to enter a simple question to get the answer, improving the user experience.

[0035] In one possible implementation, a retrieval unit is used to obtain multiple question templates corresponding to a question category, match the question with the multiple question templates, and determine the question template corresponding to the question. The retrieval unit is also used to obtain a query statement template corresponding to the question based on the question template, wherein one question template corresponds to one query statement template. Furthermore, the retrieval unit is used to populate the query statement template based on the question entity to obtain the graph query statement corresponding to the question. When the question category is a query type question, the retrieval unit uses the graph query statement to search for the business data corresponding to the question entity in the knowledge graph to obtain query results. Finally, when the question category is an information query type question, the retrieval unit uses the graph query statement to search for the related entities, relationships, and attributes of the question entity in the knowledge graph to obtain query results.

[0036] In one possible implementation, structured data includes table data stored in a fixed format, semi-structured data includes non-table data stored in a fixed format, and unstructured data includes text data with an irregular format. The system also includes a graph building unit, which is used to obtain business data from multiple business systems, obtain first entity information and first relationship information based on the fixed format features of structured and semi-structured data, perform named entity recognition on entities in unstructured data based on semantic features of unstructured data to obtain second entity information and second relationship information, and build nodes in the knowledge graph based on the first and second entity information, and build edges between nodes in the knowledge graph based on the first and second relationship information to build the knowledge graph.

[0037] In one possible implementation, the graph building unit is used to determine multiple entities in the structured data based on the fields and field values ​​in the structured data, and to obtain the first entity information of the structured data. The multiple entities in the structured data include fields and field values ​​with keywords, fields and field values ​​with attribute types of preset types, and fields and field values ​​contained in the entity template. The graph building unit is used to determine the first relationship information between the multiple entities in the structured data based on the relationship between the fields and field values, and the relationship between the fields.

[0038] In one possible implementation, the graph building unit is used to determine multiple entities in the semi-structured data based on the tags and attributes of the tags, obtain first entity information of the semi-structured data, and determine first relationship information between the multiple entities in the semi-structured data based on the relationship between tags and attributes and the nesting relationship between tags, wherein the tags in the semi-structured data are used to separate the hierarchical structure and relationships of the data; or, the graph building unit is used to determine multiple entities in the semi-structured data based on key-value pairs in the semi-structured data, obtain first entity information of the semi-structured data, and determine second relationship information between the multiple entities in the semi-structured data based on the correspondence between keys and values ​​and the nesting relationship between internal keys and external keys.

[0039] In one possible implementation, the graph building unit is used to perform named entity recognition on entities in unstructured data based on an entity recognition naming model to obtain second entity information and second relation information. The entity recognition naming model includes a word embedding layer, a feature extraction layer, and a feature classification layer. The word embedding layer is used to convert unstructured data into a word vector matrix with word-level features. The feature extraction layer is used to extract semantic features from the word vector matrix to obtain a sentence vector matrix with sentence-level features. The feature classification layer is used to classify the word segments in the unstructured data according to the sentence vector matrix to obtain second entity information and second relation information.

[0040] In one possible implementation, the word embedding layer includes a text input layer and a vector transformation layer. The text input layer is used to perform word segmentation on unstructured data to obtain multiple words. The vector transformation layer inputs the multiple embedding vectors of each word into the transformer structure to extract the word-level features of each word and obtain a word vector matrix. The multiple embedding vectors include word embedding vectors, sentence embedding vectors, and position embedding vectors.

[0041] In one possible implementation, the feature extraction layer includes a context feature extraction network and a multi-head attention network. The context feature extraction network is used to extract context features from unstructured data based on the word vector matrix, generating a multi-dimensional feature sequence containing context features. The multi-head attention network is used to determine the attention score of each element in the multi-dimensional feature sequence to obtain a sentence vector matrix, where the attention score is used to indicate the degree of influence of each element on entity recognition and naming.

[0042] In one possible implementation, the feature classification layer includes an annotation network and an optimization network. The annotation network is used to generate a label prediction result for each word based on the sentence vector matrix. The optimization network is used to adjust the label of each word in the unstructured data based on the label prediction result for each word, combined with sentence semantics and contextual information, to obtain second entity information and second relation information in the unstructured data. The label is used to indicate whether the word is an entity and the category to which the entity belongs.

[0043] Thirdly, a computing device is provided, the computing device including a processor and a memory, the memory for storing instructions and the processor for executing the instructions, such that the computing device implements the method described in the first aspect.

[0044] Fourthly, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium, and the instructions are executed by a computing device or a cluster of computing devices to implement the method described in the first aspect.

[0045] Fifthly, a computing device cluster is provided, the computing device cluster including at least one computing device, each computing device including a processor and a memory, the processor of the at least one computing device being configured to execute instructions stored in the memory of the at least one computing device, such that the computing device cluster implements the method described in the first aspect.

[0046] In a sixth aspect, a computer program product comprising instructions is provided, the computer program product including instructions capable of running on a computing device or stored in any available medium, and when the computer program product is run on a computing device or a cluster of computing devices, causing the computing device or cluster of computing devices to perform the method described in the first aspect.

[0047] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description

[0048] Figure 1 This is an architecture diagram of a data asset management system provided in this application;

[0049] Figure 2 This is an example diagram of a data asset management system deployed in a cloud environment, as provided in this application;

[0050] Figure 3 This is a flowchart illustrating the steps of a data asset management method provided in this application during the configuration phase.

[0051] Figure 4 This is a schematic diagram of the word embedding layer of an entity recognition naming model provided in this application;

[0052] Figure 5 This is a schematic diagram of the feature extraction layer of an entity recognition naming model provided in this application;

[0053] Figure 6 This is a schematic diagram of the feature classification layer of an entity recognition naming model provided in this application;

[0054] Figure 7 This is a flowchart illustrating the steps of a data asset management method provided in this application during the application phase.

[0055] Figure 8 This is a sample diagram of a client interface provided in this application;

[0056] Figure 9 This is a schematic diagram of the structure of a data asset management system provided in this application;

[0057] Figure 10 This is a schematic diagram of the structure of a computing device provided in this application;

[0058] Figure 11 This is an example diagram of a computing device cluster provided in this application;

[0059] Figure 12 This is another example diagram of a computing device cluster provided in this application. Detailed Implementation

[0060] A company's data assets (or data warehouse assets) include data related to its operations, such as customer data, sales data, financial data, production data, and employee data. These data assets are typically stored in the company's business systems. These data assets are crucial to the company. By analyzing data assets, companies can understand market trends, customer needs, product performance, financial status, etc., thereby making more informed decisions and optimizing business processes. Therefore, companies need to effectively manage, protect, and utilize data assets to achieve continuous business growth and competitive advantage.

[0061] Typically, a company's data assets originate from different business systems. For example, the supply chain department uses a supply chain system, the finance department uses a finance system, and the marketing department uses a marketing system. Different departments using their own systems to manage specific business functions leads to users needing to chain through multiple systems when retrieving data. For instance, if a sales manager wants to understand the sales and inventory of a product, they need to first access the supply chain system to view inventory reports and understand the product's inventory level, and then access the finance system to view sales records and understand the product's sales performance. If the sales manager needs more data, they may need to chain through even more systems, increasing the workload and time costs for operators and resulting in low data retrieval efficiency.

[0062] To address the problem that different departments using their own systems to manage specific business functions leads to users needing to chain through multiple systems when retrieving data, increasing workload and time costs, and resulting in low data retrieval efficiency, this application provides a knowledge graph-based data asset management method. In this method, the data asset management system establishes communication connections with multiple business systems and builds a knowledge graph based on business data from these systems. This allows information from different departments and data sources to be interconnected and queried. When a user sends a question to the data asset management system through a client, the system can search the knowledge graph based on the question entity contained in the question to obtain query results related to that entity. Then, based on the query results, it generates the answer to the question. This eliminates the need for users to chain through multiple business systems to obtain answers; they only need to input a simple question to get the answer, improving the user experience.

[0063] Figure 1 This is an architecture diagram of a data asset management system provided in this application, such as... Figure 1 As shown, the architecture includes a client 100, a data asset management system 200, and a business system 300. The client 100, the data asset management system 200, and the business system 300 establish a communication connection via a network. This communication connection can be wired or wireless. The network can be the public internet, an internal local area network (LAN), a virtual private network (VPN), a dedicated line such as fiber optic lines, copper wires, or satellite connections, or a wireless network such as wireless fidelity (Wi-Fi) or a cellular network. This application does not impose specific limitations. The number of clients 100 establishing a communication connection with the data asset management system 200 can be one or more, and the number of business systems 300 establishing a communication connection with the data asset management system 200 can be one or more. Figure 1 (This will be illustrated using supply chain systems, financial systems, and sales systems as examples; no specific limitations are set forth in this application.)

[0064] Client 100 is deployed on terminal devices or computing devices to enable human-computer interaction. Terminal devices include personal computers, smartphones, wearable devices, handheld processors, tablets, mobile laptops, augmented reality (AR) devices, virtual reality (VR) devices, smart conferencing devices, etc., without specific limitations here. The description of computing devices can be found above and will not be repeated here. Computing devices can be bare metal servers (BMS), virtual machines, containers, or storage devices. BMS refers to a general-purpose physical server, such as an ARM server or an x86 server; a virtual machine refers to a complete computer system simulated by software, possessing full hardware system functionality and running in a completely isolated environment. Any task that can be performed on a physical computer can also be performed in a virtual machine. When creating a virtual machine on a computing device, a portion of the physical machine's hard drive and memory capacity needs to be used as the virtual machine's hard drive and memory capacity. Each virtual machine has its own independent basic input / output system (BIOS), hard disk, and operating system, and can be operated like a physical machine. A container is a portable software unit that can combine an application and all its dependencies into a single software package. This package is not limited by the underlying host operating system, thus eliminating the need to build complex environments and simplifying the application development and deployment process.

[0065] The data asset management system 200 can be deployed on computing devices or on a cluster of computing devices. The description of the computing devices is as described above and will not be repeated here. The computing device cluster may include multiple computing devices, or at least one computing device and at least one storage device. The storage device may include a single hardware storage device, such as a hard disk drive (HDD), a solid-state drive (SSD), a mechanical hard disk (HDD), a universal serial bus (USB), flash memory, an SD card, a memory stick, etc., which are not specifically limited in this application. It may also include a storage array composed of multiple hardware storage devices. The storage array may be a redundant array of independent disks (RAID), network attached storage (NAS), a storage area network (SAN), etc., which are not specifically limited in this application. The storage device may also include virtual storage devices, such as cloud storage services provided by a cloud data center, which are not specifically limited in this application.

[0066] Business system 300 includes systems with various specific business functions, such as Figure 1 As shown, the business system 300 includes a supply chain system, a financial system, a sales system, etc., and may also include more business systems, such as a human resources system, a customer relationship management system, etc., which are not specifically limited in this application. Each business system 300 can be deployed on a computing device or a cluster of computing devices. The description of computing devices and computing device clusters can be found in the above description, and will not be repeated here.

[0067] Optionally, the data asset management system 200 can be as follows: Figure 1 As shown, the data asset management system 200 can establish communication connections with multiple business systems 300, or it can establish communication connections with a storage system. The storage system establishes communication connections with multiple business systems 300. The storage system is used to store data from multiple business systems 300, such as supply chain data from the supply chain system, financial data from the financial system, and sales data from the sales system. All of these are stored in the storage system. Then, the storage system establishes communication connections with the data asset management system 200. This application does not make any specific limitations.

[0068] Optionally, such as Figure 1 As shown, the data asset management system 200 and multiple business systems 300 are deployed on different computing devices or different computing device clusters. Alternatively, the data asset management system 200 and some of the business systems 300 may be deployed on the same computing device or the same computing device cluster; this application does not impose specific limitations.

[0069] Optionally, the client 100 may be software or an application running on a terminal device or computing device controlled by the user, such as a personal computer (PC) client, a World Wide Web (web) client accessed through a browser, an application (APP) client running on a mobile terminal, or a console of a cloud platform. This application does not make any specific limitations.

[0070] It should be noted that users holding Client 100 can be business personnel within an enterprise, such as data analysts, sales personnel, product managers, and financial personnel. It should be understood that many business personnel need to retrieve data from multiple business systems 300 to ensure smooth business operations. For example, financial personnel may need not only data from the financial system but also inventory data from the supply chain system and sales data from the sales system for financial statements, budgets, and other financial analyses. The above examples are for illustrative purposes only and are not intended to impose specific limitations.

[0071] Optionally, client 100 can be a client specifically designed for data asset management. Here, assets refer to data from multiple business systems 300, such as supply chain data from the supply chain system and financial data from the finance system. This type of client can be used by users to perform operations such as querying and analyzing data assets. This module can read data from multiple business systems and return the answer based on the user's input question. The question is usually a natural language question used to obtain data from multiple business systems, such as "I want this year's sales report, including information such as inventory, cost, and sales." The returned answer can be the data table corresponding to the question, so that business personnel do not need to continuously jump from multiple business systems to read the desired data when processing related business. They can obtain the data through a simple natural language question, thus improving the user experience.

[0072] Optionally, client 100 can also be a comprehensive client that includes a data asset management function module. For example, this comprehensive client can be a financial management software client, which includes financial management-related modules such as accounting, financial reporting, and tax management modules, as well as a data asset management module. This module can read data from multiple business systems and return it based on the user's input, so that financial personnel do not need to continuously switch between multiple business systems to read the desired data when processing related business.

[0073] Optionally, client 100 can also be a client of a cloud platform, used for users to purchase and rent various cloud services. The data asset management solution provided in this application can be one of the cloud services, and users can purchase the cloud service separately for data asset management; or, the cloud platform provides users with a comprehensive service, and the data asset management solution provided in this application can be a sub-service of the comprehensive cloud service. For example, the data asset management solution provided in this application can be a sub-service of the financial management cloud service. This application does not make any specific limitations.

[0074] The preceding text has described in detail the possible deployment methods for the data asset management system 200, client 100, and business system 300. In actual deployment, flexible deployment can be carried out based on specific application scenarios and business needs. The following section provides examples of actual deployment methods for the data asset management system 200, client 100, and business system 300 using specific application scenarios.

[0075] For example, in one application scenario, the data asset management system 200, client 100, and business system 300 can be deployed on the company's internal office equipment. For instance, the data asset management system 200 and multiple business systems 300 can be deployed on the services or server clusters purchased by the company, while the client 100 can be deployed on the company's office computers. Company employees can use their office computers to run the client 100 and use the functions of the data asset management system 200 through the client 100 to achieve data asset management. Specifically, company employees can input natural language through the client 100, and the data asset management system 200 can read the corresponding data from multiple business systems based on the natural language, so that company employees do not need to access multiple business systems 300 in a chain to read the data they want.

[0076] In another application scenario, the data asset management system 200 can be deployed in a cloud environment. For example... Figure 2 This is an example diagram of a data asset management system deployed in a cloud environment, as provided in this application. Figure 2 As shown, users can initiate a purchase request for data asset management cloud services through client 100. After client 100 sends the purchase request to the cloud platform, the cloud platform can provide client 100 with the cloud service access rights of data asset management system 200, enabling users to use data asset management system 200 to complete enterprise data asset management through client 100.

[0077] The cloud platform also maintains various basic resources, including computing resources, storage resources, network resources, and security resources, to meet the computing needs of the data asset management system 200 under different scales and loads. Furthermore, these computing resources can be dynamically scaled according to the usage needs of the data asset management system 200 to ensure the stable operation of the data asset management system 200 and provide users with reliable data asset management services.

[0078] Alternatively, business system 300 can also be a cloud service. The cloud service corresponding to business system 300 and the data asset management cloud service can be provided by the same cloud platform. Figure 2 The cloud platform also includes a business system 300. Similarly, users can initiate a purchase request for cloud services corresponding to the business system 300 through the client 100. After the client 100 sends the purchase request to the cloud platform, the cloud platform can grant the client 100 access to the cloud services of the business system 300, allowing users to use the business system 300 to complete related business through the client 100. Furthermore, the business system 300 can establish a communication connection with the data asset management system 200, so that when users input natural language through the client 100, the data asset management system 200 can read the corresponding data from the business system 300 based on the natural language, eliminating the need for enterprise employees to access multiple business systems 300 in a chain to read the desired data.

[0079] Optionally, the cloud services corresponding to the business system 300 and the data asset management cloud services can also be services provided by different cloud platforms. In this case, a hybrid cloud architecture can be used to realize data communication between the data asset management system 200 and the business system 300. That is, the data asset management system 200 is deployed in data center A, and the business system 300 is deployed in data center B. Data center A and data center B realize data communication between the data asset management system 200 and the business system 300 through their respective cloud platforms. When the user inputs natural language through the client 100, the data asset management system 200 can read the corresponding data from the business system 300 according to the natural language, so that enterprise employees do not need to access multiple business systems 300 in a chain to read the data they want.

[0080] It should be understood that the above application scenarios are for illustrative purposes only. The data asset management system 200, business system 300, and client 100 can be flexibly deployed according to actual business needs. These will not be illustrated one by one here.

[0081] Furthermore, such as Figure 1 As shown, the data asset management system 200 can be divided into a question-answering device 220 and a knowledge graph device 210 according to its functions. It should be understood that this application has divided the data asset management system 200 for the convenience of describing the scheme. Figure 1 The division shown is one possible implementation. In the specific implementation process, the data asset management system 200 may include more or fewer devices, units or modules, depending on the actual business scenario. This application does not make any specific limitations.

[0082] The question-answering device 220 and the knowledge graph device 210 can be implemented in software or hardware. For example, the implementation of the question-answering device 220 will be described below. Similarly, the implementation of the knowledge graph device 210 can refer to the implementation of the question-answering device 220.

[0083] As an example of a software functional unit, the question-answering device 220 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, the question-answering device 220 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.

[0084] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0085] As an example of a hardware functional unit, the question-answering device 220 may include at least one computing device, such as a server. Alternatively, the question-answering device 220 may be implemented using a central processing unit (CPU), an application-specific integrated circuit (ASIC), or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), a data processing unit (DPU), a neural network processing unit (NPU), a system-on-a-chip (SoC), an offload card, an accelerator card, or any combination thereof.

[0086] The question-answering device 220 includes multiple computing devices that can be distributed in the same region or in different regions. Similarly, the multiple computing devices in the question-answering device 220 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices in the question-answering device 220 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, GALs, DPUs, NPUs, SoCs, offloading cards, and accelerator cards.

[0087] It should be noted that, in other embodiments, the question-answering device 220 can be used to execute any step in the data asset management method provided in this application, and the knowledge graph device 210 can be used to execute any step in the data asset management method provided in this application. The steps implemented by the question-answering device 220 and the knowledge graph device 210 can be specified as needed. By implementing different steps in the data asset management method provided in this application through the question-answering device 220 and the knowledge graph device 210 respectively, all functions of the data asset management system 200 can be realized.

[0088] The functions of the question-answering device 220 and the knowledge graph device 210 will be explained below.

[0089] The knowledge graph device 210 is used to build a knowledge graph based on the business data provided by the business system 300, and search the knowledge graph to obtain query results based on the question entities provided by the question answering device 220, and then feed the query results back to the question answering device 220.

[0090] In practical implementation, the business data provided by multiple business systems 300 may include structured data, semi-structured data, and unstructured data. Structured data includes table data with a strictly defined organizational structure (strictly defined by rows and columns). Semi-structured data includes data stored in a fixed format but without a strictly defined organizational structure, such as XML files and JSON files. Unstructured data refers to data that does not conform to a fixed format, such as text data like emails, reports, news articles, and social media comments; it may also include images and videos, but this application does not specifically limit it.

[0091] In this embodiment, the knowledge graph device 210 can obtain entity information and association information from structured and semi-structured data according to their fixed formats, extract entity-relationship triples (entity 1, relation, entity 2) from the structured and semi-structured data, then identify semantic features of the unstructured data, extract entities from the unstructured data and name them, then determine the association relationships between entities, thereby generating entity-relationship triples from the unstructured data, and finally establish a knowledge graph based on the entity-relationship triples. The knowledge graph may include multiple nodes, each node corresponding to an entity, and the edges between nodes are used to represent the association relationships between entities.

[0092] It should be noted that an entity refers to a single, referable event, object, or concept. This can be a concrete event in the real world, such as a person, place, organization, or product, or an abstract concept, such as an event, date, or number. In a knowledge graph, entities are nodes, and the relationships between entities are edges between nodes. For example, in a social network knowledge graph, entities could be users, posts, comments, etc., and relationships could be between users, between users and posts, between users and comments, etc. These examples are for illustrative purposes only and are not intended to impose specific limitations.

[0093] It should be understood that by integrating data from multiple business systems into a knowledge graph through a knowledge graph device, information from different departments and data sources can be interconnected and queried. Enterprises can achieve more intelligent data asset search and recommendation functions. When users query data assets, they do not need to access multiple business systems in a chain. They can directly complete the data reading by searching the knowledge graph, which is more efficient and improves the user experience.

[0094] The question-answering device 220 is used to receive the question input by the user sent by the client 100, obtain the question entity that exists in the question, and then send the question entity to the knowledge graph device 210 for searching, obtain the query results, and generate the answer corresponding to the question based on the query results.

[0095] In its specific implementation, the question-answering device 220 can obtain the question entity in the question through entity recognition, and then use the question entity as a keyword to search for information related to the question entity in the knowledge graph, such as querying the entity's attributes, related entities, etc., to obtain the query results. Then, based on the query results, it generates the answer corresponding to the question. For example, if the question asks to view a certain chart, the question-answering device 220 can generate the corresponding line chart, bar chart, or table based on the query results to meet the user's needs.

[0096] It should be understood that in the application scenario of enterprise data asset management, there are characteristics of large data volume and wide range of data sources. Without using a knowledge graph to build a question-answering system, relying solely on semantic recognition of user questions and generating SQL to search data across multiple business systems 300 can only handle simple questions. This is because user input is diverse, and semantic recognition may be misunderstood or incorrect, leading to inaccurate query results. Furthermore, if data needs to be searched from multiple tables across multiple business systems 300, relational search operations are required, resulting in slow query speeds. This application builds a knowledge graph based on data from business systems 300. The knowledge graph can easily model complex relationships between data assets. During queries, the knowledge graph can quickly search for data assets related to the question entity, improving query speed. Simultaneously, because the entities in the knowledge graph are named, user input does not require semantic recognition; only entity recognition is needed. Entity recognition does not involve complex semantic understanding, resulting in higher efficiency and accuracy, thus improving query accuracy.

[0097] Furthermore, in the application scenario of data asset management, when users use the data asset management system, the questions they input may include queries for table data, such as "I want to see the 2013 financial annual report," as well as more complex information queries, such as "Who is the person in charge of asset X?" Without using a knowledge graph to build a question-answering system, such information queries through semantic recognition would require connecting and filtering across multiple tables, making the SQL statement generation process very complex. However, a question-answering system based on a knowledge graph can handle these types of questions very well. The relationships between assets can be expressed and stored in the form of a knowledge graph, allowing users to quickly obtain answers to these complex information queries.

[0098] In summary, the data asset management system provided in this application establishes communication connections with multiple business systems and builds a knowledge graph based on business data from these systems. This allows information from different departments and data sources to be interconnected and queried. When a user sends a question to the data asset management system via a client, the system can search the knowledge graph based on the question entity contained in the question to obtain query results related to that entity. Then, it generates the answer to the question based on the query results. This eliminates the need for users to chain through multiple business systems to obtain answers; they only need to input a simple question to get the answer, improving the user experience. Furthermore, structured and semi-structured data in the business data can have entities extracted based on their fixed format characteristics to build a knowledge graph, while unstructured data can have entities identified based on its semantic features to build a knowledge graph. The resulting knowledge graph encompasses the knowledge features contained in structured, semi-structured, and unstructured data from multiple business systems, providing users with more comprehensive answers. This not only enables the reading of business data from multiple systems but also allows for complex information queries, improving the work efficiency of business personnel.

[0099] The data asset management system provided in this application has been described in detail above. The following section will combine... Figures 3-8 This application provides an explanation of the data asset management method provided, which can be applied to, for example... Figure 1 and Figure 2 The data asset management system shown. Among them, Figures 3-6 The explanation focuses on the map generation stage. Figures 7-8 Explanation and instructions are provided for the map lookup stage.

[0100] Figure 3 This is a flowchart illustrating the steps of a data asset management method provided in this application during the configuration phase, as shown below. Figure 3 As shown, the method may include the following steps:

[0101] S310: Client 100 sends a map creation request to data asset management system 200.

[0102] This step can be performed by Figure 1 The knowledge graph device 210 in the embodiment is implemented.

[0103] In a specific implementation, the knowledge graph creation request may include data source information, which refers to business system 300. It should be understood that the data asset management system 200 has established communication connections with multiple business systems 300. When generating the knowledge graph, the system can select some or all of the business data from the multiple business systems 300 according to the user's needs to complete the knowledge graph creation. Specifically, the data source information may include the name and address of the business system 300, as well as information such as the scope, type, and format of the available business data in the business system 300, and may also include other data source information; this application does not specifically limit this.

[0104] Optionally, the graph creation request may also include some custom information required for knowledge graph creation. For example, the graph creation request may also include graph scale information, such as the number of nodes and relationships; it may also include structural information, such as the graph's hierarchical structure and visualization requirements, to ensure that the knowledge graph generated by the system meets the user's expectations; and it may also include security information, such as access permissions and encrypted transmission methods, to ensure the security and confidentiality of the knowledge graph. The graph creation request may also include more content, which can be determined according to the actual business scenario, and this application does not impose specific limitations.

[0105] S320: Data asset management system 200 obtains business data from business system 300.

[0106] This step can be performed by Figure 1 The knowledge graph device 210 in the embodiment is implemented.

[0107] In practice, the data asset management system 200 can obtain business data from the corresponding business system 300 based on the data source information in the knowledge graph creation request. The data asset management system 200 can obtain data source information in batches and build and update the knowledge graph based on the data source information obtained in each batch.

[0108] In specific implementations, data source information includes structured data, semi-structured data, and unstructured data. Structured data refers to table data organized and stored according to a fixed format. Semi-structured data refers to data stored in a fixed format but not conforming to the characteristics of table data (or relational databases), such as XML files and JSON files. Unstructured data refers to text data that does not conform to a fixed format, such as emails, reports, news articles, and social media comments; it may also include images and videos, though this application does not specifically limit it.

[0109] For example, in enterprise data asset management scenarios, structured data can be various forms stored in business systems. For instance, customer information forms can be stored in the database of a customer relationship management system, transaction record forms can be stored in an enterprise resource planning (ERP) system, and profit and loss statements, balance sheet statements, etc., can be stored in the financial system. Semi-structured data can be web page data (XML files), user profiles (JSON files), etc. Unstructured data can be emails, documents, reports, social media content, product images, etc. The above examples are for illustrative purposes only and are not intended to be specific.

[0110] S330: The data asset management system 200 generates the first entity relationship data corresponding to the structured data and semi-structured data.

[0111] This step can be performed by Figure 1 The knowledge graph device 210 in the embodiment is implemented.

[0112] In its implementation, the first entity relation data includes entity-relation triples corresponding to structured and semi-structured data. An entity-relation triple is a data structure used to describe and store knowledge, and is employed to build knowledge graphs. A triple contains two entities and the relationship between them, i.e., (entity 1, relation, entity 2). This structure clearly represents the semantic associations between entities. Entity 1 and entity 2 are specific objects or concepts, such as people, places, organizations, and events, while the relation is the specific connection between these two entities, such as belonging to or being located at. For a description of entities and relations, please refer to [reference needed]. Figure 1 The relevant descriptions of the embodiments will not be repeated here. For example, if the employee department relationship table shows that Zhang San is an employee of the Human Resources Department, then Entity 1 can be Zhang San, Entity 2 can be the Human Resources Department, and the relationship is "belongs to", that is, (Zhang San, belongs to, Human Resources Department). The above examples are for illustration only and are not intended to limit the scope of this application.

[0113] In its implementation, the data asset management system 200 first determines the entity information in structured and semi-structured data, then determines the relationship information between entities in the structured data, and finally determines the first entity relationship data based on the entity and relationship information. Since structured data is organized and stored according to a fixed format, when determining entity and relationship information, the system can extract entity and relationship information from the structured and semi-structured data based on their fixed format characteristics, thereby obtaining the first entity relationship data. As mentioned above, structured data includes table data, and semi-structured data includes XML, JSON files, etc. These data have certain differences in their fixed format characteristics. The following sections explain how entity and relationship information are determined for structured and semi-structured data.

[0114] In this embodiment, when determining entity information in structured data, such as table data, some fields and field values ​​in the table data can be identified as entities. Specifically, fields or field values ​​that begin or end with specific keywords may represent specific types of entities; for example, a field starting with "product_" may represent a product entity. Alternatively, the entity represented by a field can be inferred from its data type; for example, a string type may represent a name or description, a numeric type may represent quantity or value, and a date type may represent time or date. Alternatively, the entity represented by a field or field value can be determined based on expert knowledge. Field templates that may serve as entities can be configured, and the fields or field values ​​in the table data that can serve as entities can be determined based on these templates, thereby obtaining entity information. The above examples are for illustrative purposes only and are not intended to limit the scope of this application.

[0115] Optionally, when determining the association information in table data, the association between entities can be determined based on the relationships between records of fields and field values ​​in the table data. The relationships between records in table data are diverse. For example, there are relationships between fields and multiple field values ​​under a field. For instance, in the employee department table, if the field value under the "Department A" field includes "Employee X", then the entity "Employee X" and the entity "Department A" will also have a relationship, and the corresponding entity relationship triple could be (Employee X, Belongs to, Department A). For example, there are relationships between primary keys and foreign keys. A primary key is a field or combination of fields that uniquely identifies an entity; its value is unique within the table. For example, in the order table, the order number is the primary key, and each order has a unique order number. A foreign key is a field in one table that points to the primary key of another table, used to establish relationships between tables. For example, in the order table, "Product ID" is a foreign key that points to "Product ID" in the product table, indicating a relationship between the order number and the product ID. Therefore, relationships between some entities can be determined based on primary keys and foreign keys. Of course, the table data may contain more relationships, which will not be listed here, and this application does not limit this.

[0116] Furthermore, correlation analysis can be performed on fields that are not recorded as having relationships, or between field values, to determine the relationships between entities. For example, analyzing the frequency of two fields appearing simultaneously in the data can help determine their relationship. It should be understood that if a product appears frequently in an order, it can be assumed that the product was likely purchased in that order. Alternatively, expert knowledge can be used to determine the relationships between entities. Association templates can be configured to identify potential relationships between entities, and the associated fields in the table data can be determined based on these templates to obtain the association information. The above examples are for illustrative purposes only and are not intended to limit the scope of this application.

[0117] The above provides an example of how to determine entity information and related information in structured data. In the actual implementation of the solution, other methods can also be used to determine entity information and related information, and this application does not impose specific limitations.

[0118] In this embodiment, entity information and related information in semi-structured data can be determined based on markers or tags. Specifically, some entities can be determined based on markers and their attributes, and related information can be determined based on the relationship between markers and attributes, and the nesting structure between markers. Markers can be XML tags, JSON keys, etc. Semi-structured data typically contains tags describing the data structure, which helps in understanding the data's structure and content. For example, XML uses angle brackets (e.g., ...). <person>Elements are identified by angle brackets (); the names within these angle brackets are tags. XML tags can identify entities and attributes in data, as well as the relationships between them. For example, in JSON, key-value pairs are used to represent data attributes and values; the key identifies the attribute of an entity, and the value represents the specific value.

[0119] Specifically, for XML files, the content between the start and end tags can be searched. Typically, the start tag indicates the beginning of an entity, and the end tag indicates its end. Therefore, by using the start and end tags, the entity corresponding to each tag can be identified, and the nesting relationship between tags can be identified. Outer tags represent more general, higher-level entities, while inner tags represent more specific, more detailed entities. Based on the nesting relationship, the association information between entities can be generated. Furthermore, entities can also be identified based on tag attributes; the association information between entity 1 (corresponding to a tag) and entity 2 (corresponding to an attribute) can also be determined. This method allows the acquisition of entity relationship data corresponding to the XML file.

[0120] Specifically, for JSON files, key-value pairs can be found within the file. Typically, one key corresponds to one entity. The nesting relationships between keys can then be identified. Here, the second key is a subkey of the first key, and the second key is an attribute or member of the object described by the first key. The first key represents a higher-level, more general entity, while the second key represents a more specific entity. Based on these nesting relationships, association information between entities can be generated. Furthermore, entities can be determined based on the values ​​corresponding to keys; a relationship can also be established between entity 1 (key) and entity 2 (value). This method allows the acquisition of entity relationship data within a JSON file.

[0121] It should be noted that some tags in semi-structured data may not be able to be considered entities. By pre-setting tag templates, some tags or attribute values ​​that may be entities can be configured. Entities in semi-structured data can be determined based on the tag templates, and then the association information can be determined based on the nesting relationship and attributes of the tags. This will not be elaborated on here.

[0122] For example, some tags in an XML file <person> 、 <order> 、 <product>"" can be used as user entities, order entities, and product entities. Each tag can include multiple attributes, such as<person id="123"> The ID attribute in XML can represent a person's unique identity. Therefore, entity 1 is "person", entity 2 is "123", and the relationship is "id". XML tags also have nested structures, for example, a... <order>May contain multiple <product>In this case, entity 1 is order and entity 2 is product, and the relationship is containment. The above example is for illustration only and is not intended to be specific.

[0123] The above provides an example of how to determine entity information and related information in semi-structured data. In the actual implementation of the solution, other methods can also be used to determine entity information and related information, and this application does not impose specific limitations.

[0124] Furthermore, after determining the entity and relation information of structured and semi-structured data, the first entity relation data, namely the entity relation triple (entity 1, relation, entity 2), can be generated according to the format required by the knowledge graph.

[0125] It should be understood that S330 generates the first entity relationship data for structured and semi-structured data, and S340 to S360 generate the second entity relationship data for unstructured data. Therefore, S330 and S340 to S360 can be executed in parallel. Whether they are executed in parallel depends on the processing capacity of the data asset management system 200. This application does not make specific limitations.

[0126] S340: The Data Asset Management System 200 performs entity recognition and naming on unstructured data to obtain entity information of the unstructured data.

[0127] This step can be performed by Figure 1 The knowledge graph device 210 in the embodiment is implemented.

[0128] In its implementation, the data asset management system 200 uses an entity recognition model to obtain entity information from unstructured data. Specifically, the entity recognition model includes a word embedding layer, a feature extraction layer, and a feature classification layer. The word embedding layer converts the unstructured data (i.e., text data) input to the model into word vector representations, obtaining a word vector matrix containing word-level features. The feature extraction layer further extracts semantic features from the word vector matrix, obtaining a sentence vector matrix containing sentence-level features. The feature classification layer annotates the input text based on the sentence vector matrix, obtaining the identification results of entities present in the input text.

[0129] To facilitate understanding, we will first explain the word embedding layer.

[0130] In this embodiment, the word embedding layer is used to obtain word vector representations of unstructured data (i.e., text data), which are N-dimensional word vector matrices {x1,x2,x3,x4,…,x} containing semantic features. n ,}, where x i The feature vector representing the i-th word. Specifically, the word embedding layer may include a text input layer and a vector conversion layer. The text input layer is used to process unstructured data to obtain multiple tokens, and the vector conversion layer is used to extract the semantic features of the unstructured data based on the multiple tokens and generate a word vector matrix containing semantic features. It should be understood that the word embedding layer may also add more levels according to the actual business scenario to implement the above functions, and the present application does not make specific limitations.

[0131] Optionally, the processing process of the text input layer may include text preprocessing and tokenization operations. Among them, text preprocessing is used to clean and standardize unstructured data, and tokenization is used to perform tokenization processing on the unstructured data after text preprocessing to obtain multiple tokens. In specific implementation, text preprocessing may include data cleaning, removing noises such as special characters and garbled codes in the original text data, removing useless information such as modal particles, spaces, and punctuation marks, and reducing the complexity of subsequent processing; text preprocessing may also include removing stop words. Stop words refer to words that frequently appear in text data but do not contain useful information, such as "is", "of", "in", etc. Removing stop words can reduce the data dimension and improve the efficiency of subsequent processing. Text preprocessing may also include more content, such as normalization, standardization, etc. operations, and the present application does not make specific limitations. Tokenization operation refers to dividing text data into words or phrases with certain meanings, and a tokenization tool can be used to cut the text into a sequence of words, such as the jieba tokenization tool. By performing tokenization processing on text data to obtain multiple tokens, it helps the system understand the semantics of the text data, thereby better supporting subsequent entity recognition and naming operations.

[0132] Optionally, the vector conversion layer is used to generate a word vector matrix corresponding to the unstructured data (that is, the text data of the input embedding layer) according to the text features of multiple tokens. This word vector matrix is a vector representation containing the deep semantic features of multiple tokens and is the semantic feature at the word level. In specific implementation, the vector conversion layer may first determine multiple different types of embedding vectors corresponding to each token. Each type of embedding vector is used to represent a feature of the token. The embedding vector can be understood as a supplementary feature of the token. Input the multiple embedding vectors of the above multiple tokens into the feature extraction layer to obtain the deep semantic vector representation of the unstructured data, that is, the above word vector matrix.

[0133] Exemplarily, Figure 4 is a schematic diagram of the word embedding layer of an entity recognition and naming model provided by the present application, as Figure 4 As shown, the input data for the word embedding layer is unstructured data, i.e., text data, and the output is a word vector matrix. The word embedding layer includes a text input layer and a vector transformation layer. The text input layer is used to preprocess the text data and perform word segmentation to obtain multiple words (see the previous content for details). The vector transformation layer first determines various types of embedding vectors for the word segments, including word embeddings (TE) vectors, sentence embeddings (SE) vectors, and position embeddings (PE) vectors. Then, these three types of embedding vectors are input into the transformer structure to obtain the word vector matrix.

[0134] Among them, the TE vector is obtained based on a vocabulary that includes the mapping relationship between words and word vectors. The TE vector is a unique vector corresponding to the word segmentation. The SE vector is used to identify the sentence or paragraph where the word segmentation is located, and the PE vector is used to identify the position of the word segmentation in the sentence. Based on the above-mentioned multiple types of embedding vectors, a comprehensive embedding vector is generated. This comprehensive embedding vector contains the semantic features of the word segmentation, the features of the sentence where the word segmentation is located, and the position features of the word segmentation in the sentence. Inputting this comprehensive embedding vector into the transformer structure can enable the model to obtain richer information, better extract the semantic features of the word segmentation, and obtain a deep semantic vector representation.

[0135] Furthermore, the transformer structure comprises multiple transformer blocks. The output of each transformer block can be input into the next transformer block, allowing the model to process, learn, and represent the input data layer by layer, thereby better capturing the semantics and structure of the input data. Each transformer block includes a multi-head attention network and a feedforward neural network. The output of the multi-head attention network is input into the feedforward neural network. The multi-head attention network consists of multiple attention heads, each generating an attention weight matrix. These attention weights represent the importance of a position, guiding the subsequent feedforward neural network to prioritize certain positions when processing the data. Moreover, each attention head determines its attention weight matrix by calculating the similarity between that position and other positions. Therefore, each attention weight matrix also includes the relationship between the word segment and its adjacent segments, enabling the extraction of deep bidirectional semantic features of the word segment. Feedforward neural networks are used to perform non-linear transformations and feature extraction on word segmentation at each position, helping the model learn richer and more abstract feature representations, thereby better extracting the deep semantic features of text data. Combined with the attention weights output by the attention head, semantic features can be extracted with emphasis, so that the final output word vector matrix is ​​a word vector matrix containing deep features.

[0136] It should be understood that the above Figure 4 The embedding layer shown is for illustrative purposes only. The vector transformation layer can also be implemented based on pre-trained models such as BERT, RoBERTa, ALBERT, and GPT; this application does not impose any specific limitations. It should be noted that the word vector matrix generated by the embedding layer is a vector representation generated after deep semantic feature extraction from unstructured data. Inputting this word vector matrix into subsequent feature extraction layers can extract richer and more accurate semantic information. Based on this, entity recognition and naming can be performed, resulting in more accurate entity and association information and a more accurate knowledge graph.

[0137] Secondly, the feature extraction layer will be explained.

[0138] In this embodiment, the feature extraction layer is used to extract features from the word vector matrix generated by the word embedding layer to obtain a sentence vector matrix. This sentence vector representation is an N-dimensional sentence vector matrix {w1, w2, w3, w4, ..., w...} containing semantic features. n ,}, where w i This represents the feature vector of the i-th word. It should be understood that x is the feature vector extracted by the word embedding layer. i These are word-level features, which can effectively express the meaning of a word within its adjacent words. The w extracted by the feature extraction layer... i These are sentence-level features, representing the semantic and structural information of a sentence. They effectively express the meaning of a sentence within its context. This sentence vector matrix not only contains the semantic information of each word in the text data but also integrates the semantic information of the text data within its context. Such feature information is more conducive to entity recognition and naming. Specifically, the feature extraction layer may include a context feature extraction network and a multi-head attention network.

[0139] Optionally, the context feature extraction network is used to extract context features from the input word vector matrix, generating a multi-dimensional feature sequence containing context features. Specifically, the context feature extraction network may include a forward extraction network, a backward extraction network, and a fusion network. The forward extraction network processes the text data from the first word to the last word, the backward extraction network processes the text data from the last word to the first word, extracting a second context feature, and the fusion network generates a multi-dimensional feature sequence containing context features based on the first and second context features. Specifically, each word segment can have its information between the first word and the current word extracted by the forward extraction network, obtaining a first context feature; its information between the last word and the current word can be extracted by the backward extraction network, obtaining a second context feature; and its feature representation can be obtained based on the fusion network. The feature representations of multiple words in the text data are concatenated to obtain a multi-dimensional feature sequence.

[0140] In specific implementations, the context feature extraction network can be implemented based on various pre-trained models such as convolutional neural network (CNN), bi-directional long short-term memory (BiLSTM), transformer model, and gated recurrent unit (GRU), and this application does not impose any specific limitations.

[0141] Optionally, a multi-head attention network is used to determine the key positions in the multi-dimensional feature sequence and generate sentence-level features, namely a sentence vector matrix. This sentence vector matrix contains feature information extracted from the multi-head attention network and the context feature extraction network, so that in subsequent entity recognition and naming, not only can word-level features be obtained, but also text context features can be obtained, and important positions can be identified, classified and named, thus achieving better entity recognition and naming results.

[0142] For example, Figure 5 This is a schematic diagram of the feature extraction layer of an entity recognition naming model provided in this application, as shown below. Figure 5 As shown, the feature extraction layer may include a context feature extraction network and a multi-head attention network. The context feature extraction network generates a multi-dimensional feature sequence corresponding to the word vector matrix, and then the multi-dimensional feature sequence is input into the multi-head attention network to generate a sentence vector matrix.

[0143] The multi-head attention network can include multiple attention heads, each of which is used to process a portion of the multidimensional feature sequence. In a specific implementation, the multidimensional feature sequence can be divided equally according to the number of attention heads to obtain multiple segmented data, and each segmented data is assigned to an attention head for processing.

[0144] Furthermore, each attention head can first pass through a linear transformation layer to obtain a linear transformation sequence corresponding to the segmented data, which is used for subsequent attention calculations. This linear transformation is typically implemented by multiplying a weight matrix. Based on the linear transformation sequence, the attention score corresponding to the segmented data is calculated, usually through a single scaling dot product calculation. This attention score includes the score for each position in the segmented data, used to measure the importance of each element at each position. Then, a scaling operation is performed on the attention score to stabilize the subsequent softmax operation. Scaling ensures numerical stability in the subsequent softmax calculation, preventing gradient explosion or vanishing problems. Next, a softmax operation is performed based on the scaled attention score, mapping the attention score to the range of 0-1 to obtain the attention weights corresponding to the segmented data. This is done so that the model can focus on different positions in the input data in a probability distribution. Finally, the attention weights are multiplied by the segmented data to obtain the attention sequence of the segmented data. The purpose of this step is to apply the attention weights to the multidimensional feature sequence generated by the upper and lower feature extraction networks, emphasizing the important parts in the multidimensional feature sequence. Such attention training participates in subsequent entity recognition and naming, which enables the model to complete the task more effectively and improves the performance of entity recognition and naming.

[0145] Furthermore, each attention head can obtain the attention sequence corresponding to the segmented data according to the above process. The attention sequences of multiple attention heads can be input into the fully connected layer to obtain sentence-level feature vectors. In the fully connected layer, the attention sequences obtained by multiple attention heads can be concatenated or weighted to obtain richer sentence representations and obtain the sentence vector matrix corresponding to the word vector matrix.

[0146] It should be understood that Figure 5 The structure shown is for illustrative purposes only. In actual implementation, the feature extraction layer can be further expanded with more network layers according to business needs. This application does not impose any specific limitations.

[0147] Finally, the feature classification layer will be explained.

[0148] In this embodiment, the feature classification layer is used to label and name the entity parts in the text data based on the sentence vector matrix output by the feature extraction layer, thereby obtaining entity information of the text data. Specifically, the feature classification layer includes a labeling network and an optimization network. The labeling network generates a label prediction result for each word segment based on the sentence vector matrix, while the optimization network comprehensively considers the label prediction results for each word segment to generate the optimal label sequence, thus obtaining entity data in the text data. The label indicates whether a word segment is an entity. The label prediction result for each word includes the probability that the word belongs to each label. For example, if the prediction result for word segment 1 is (0.2, 0.3, 0.4), it means that word segment 1 has a probability of 0.2 belonging to label 1, a probability of 0.3 belonging to label 2, and a probability of 0.4 belonging to label 3. The above example is for illustration only and is not intended to limit the scope of this application.

[0149] Optionally, the annotation network can predict the label for each word segment based on a preset labeling method (e.g., BIO (begin, inside, outside) labeling method) to obtain the label prediction result corresponding to each word segment. In specific implementation, a sample set can be labeled using the BIO labeling method, and then the entity naming model can be trained using the sample set, enabling the annotation network to learn the ability to label text data using the BIO labeling method based on the sentence vector matrix. The labels generated by the BIO labeling method can include B tags, I tags, and O tags. A word segment labeled with a B tag indicates that it is the beginning of an entity; a word segment labeled with an I tag indicates that it is a non-beginning part of an entity, that is, the remaining part of the entity excluding the beginning part; and a word segment labeled with an O tag indicates that it is not an entity.

[0150] Furthermore, the tags also include the category to which the entity belongs. Based on the category in the tag, the entity can be named. For example, the B-per tag indicates the beginning of the entity, and the type is person; the I-per tag indicates the non-beginning part of the entity, and the type is person; the B-loc tag indicates the beginning of the entity, and the type is location; the I-loc tag indicates the non-beginning part of the entity, and the type is location. The above examples are for illustration only, and this application does not impose specific limitations. By expanding the tags, the entity part and the entity type can be clearly identified. For example, Table 1 is a tag example table provided by this application. In this example, the text data is: "Zhang San founded Company X in Hangzhou". Its tags can be as shown in Table 1 below. It should be understood that Table 1 is for illustration only, and this application does not impose specific limitations.

[0151] Table 1: Example of Labels

[0152] Word segmentation Zhang San exist Hangzhou Establishment Company X Label B-PER I-PER O B-LOC I-LOC O B-ORG I-ORG I-ORG

[0153] It should be noted that, in addition to using BIO tags to annotate the sample set, other annotation methods can also be used, such as IO tags (I indicates the inside of an entity, O indicates a non-entity), BIOES tags (B indicates the beginning of an entity, I indicates the inside of an entity, O indicates a non-entity, E indicates the end of an entity, and S indicates an entity composed of a single word), etc. This application does not impose specific limitations on these methods.

[0154] Optionally, the optimization network can comprehensively consider the label prediction results of each word segment, capture the global dependencies between the labels of each word segment, adjust the label prediction results of each word segment, and obtain the optimal label sequence corresponding to the text data. Specifically, it can consider the transition probabilities between labels of different words and the degree of correlation between labels, and use a dynamic programming algorithm to determine the optimal label sequence by setting an optimization objective. In a specific implementation, the optimization network can be based on a conditional random field (CRF).

[0155] For example, Figure 6 This is a schematic diagram of the feature classification layer of an entity recognition naming model provided in this application, such as... Figure 6 As shown, this feature classification layer includes an optimization network and a labeling network. After the sentence vector matrix output by the feature extraction layer is input into the labeling network, the labeling network can output the label prediction result for each word segmentation. For example, the label prediction result for word segment 1 is (1.3, 0.8, 0.5, 0.12, 0.2), indicating that the predicted label for word segment 1 is "B-per", which is the beginning of the entity, and the entity category is "person". Similarly, the label prediction results for other words can be obtained, such as... Figure 6 As shown, details will not be elaborated here. The annotation network inputs the predicted label results for each word segment into the optimization network, such as the CRF network. The CRF network, based on the labels of multiple word segments, comprehensively considers the semantics and contextual information of the entire sentence, adjusts the predicted label results for each word segment, and outputs the optimal label sequence. This optimal label sequence includes the label corresponding to each word, and entity information of the text data can be obtained based on these labels.

[0156] For example, such as Figure 6 As shown, based on Figures 4-6 The entity recognition and naming model constructed from multiple network layers shown can identify entity information in text data. This entity information can be used to construct entity relation triples for knowledge graphs. The system can also provide users with an entity query interface, such as displaying... Figure 6 The interface 610 shown presents the user with the entity information automatically labeled by the model. If the user believes that the entity information is incorrectly labeled, the user can manually correct the entity information. The corrected entity information can be used as a new sample set to further train the model and continuously optimize the performance of the entity recognition naming model. Figure 6 This application is for illustrative purposes only and does not constitute a specific limitation.

[0157] It should be understood that the label prediction results of each word segment obtained by the annotation network are based on local features. However, a word may represent different entities in different contexts. For example, ink can represent writing ink or knowledge. If the sentence is "I go to buy ink", ink represents an object. If the sentence is "I have ink in my belly", ink represents knowledge. Therefore, label prediction based solely on the annotation network may result in errors. By using an optimization network to analyze the label prediction results of each word segment in the entire sentence, considering the contextual information of the entire sentence, and understanding the meaning and structure of the sentence, the entity recognition results obtained are more accurate.

[0158] S350: Data Asset Management System 200 determines the relationship information between entities in unstructured data.

[0159] This step can be performed by Figure 1 The knowledge graph device 210 in the embodiment is implemented.

[0160] In one possible implementation, the entity recognition naming model can output entity information and relational information. Optionally, the feature classification layer can output entity information and relational information, and the labels used to train the entity naming model can be composite labels, which include not only the entity category but also the entity's relation, such as the label "person-organization-job title". The model is then trained using a sample set labeled with composite labels, so that the model can directly output entity information and relational information.

[0161] Optionally, the feature classification layer may only output entity information. The entity recognition naming model may also include a relation determination layer, which is essentially a classifier that classifies relations based on the input features. Specifically, the input data of the relation determination layer includes the entity information output by the feature classification layer and the sentence vector matrix output by the feature extraction layer. The output data of the relation determination layer includes the relationship information between entities. This relation determination layer can be implemented based on classifiers such as fully connected networks, convolutional neural networks (CNNs), and recursive neural networks (RNNs), and this application does not impose specific limitations.

[0162] It should be understood that the word embedding layer and feature extraction layer of the entity recognition naming model deeply mine the features in the text data and effectively capture the relationships between word segments. Therefore, based on the sentence vector matrix output by the feature extraction layer, a classifier can generate information about the relationships between entities. For example, the context feature extraction network in the feature extraction layer processes text data in both forward and backward directions, effectively capturing the positional information of word segments in the sentence and the relationships between word segments. The multi-head attention network can identify the dependencies between different positions in the sentence and assign different attention weights to word segments at different positions in the sentence, which can help the model better understand the semantic relationships between entities in the sentence. Such sentence vector features not only include the semantic information of the sentence but also the relationship features between word segments. Inputting such sentence vector features into the classifier can classify relationships and obtain information about the relationships between entities.

[0163] In another possible implementation, the entity recognition naming model can also generate only entity information. The data asset management system 200 identifies relationships between entities based on preset rules. These preset rules may include templates for fixed entities. The relationships between entities are determined by judging whether the text contains the content of these templates. For example, "X and Y cooperate" can be used as a template, and the corresponding relationship is a cooperation relationship. If the sentence is "Company A and Company B have reached a strategic cooperation agreement," the relationship between the two entities, Company A and Company B, can be determined to be a "cooperation" relationship based on this template. The above examples are for illustration only and are not intended to limit the scope of this application.

[0164] It should be understood that this application may also determine the relationship information between entities in other ways, and this application does not make specific limitations.

[0165] S360: Data Asset Management System 200 determines the second entity relationship data corresponding to unstructured data based on entity information and relationship information.

[0166] This step can be performed by Figure 1 The knowledge graph device 210 in the embodiment is implemented.

[0167] In practice, the second entity relation data is similar to the first entity relation data; specifically, it can be the entity relation triple (entity 1, relation, entity 2) required to build the knowledge graph. For distinction, entity relation triples generated from structured and semi-structured data are referred to as the first entity relation data, while entity relation triples generated from unstructured data are referred to as the second entity relation data.

[0168] S370: Data Asset Management System 200 establishes a knowledge graph based on first entity relationship data and second entity relationship data.

[0169] This step can be performed by Figure 1 The knowledge graph device 210 in the embodiment is implemented.

[0170] In specific implementation, the data asset management system 200 can establish node 1 and node 2 in the knowledge graph based on entity relationship triples (entity 1, relationship 1, entity 2), where node 1 is identified as entity 1, node 2 is identified as entity 2, and the edge between node 1 and node 2 represents relationship 1. This process continues, establishing multiple nodes and edges based on each entity relationship triple to complete the knowledge graph construction. Furthermore, attributes of the corresponding nodes can be defined based on the relevant information of each entity. For example, for a person entity, the corresponding nodes include attributes such as name and age; for a location entity, the corresponding nodes include attributes such as location name, address, and type. The relevant information used to define the attributes can be extracted from the business system 300 or determined based on the entity relationship triples; this application does not impose specific limitations. During the construction of the knowledge graph, data can be cleaned, integrated, and expanded according to requirements.

[0171] In its implementation, the data asset management system 200 can also use graph databases or graph processing frameworks to manage and query knowledge graphs, such as Neo4j and Apache Spark; this application does not impose specific limitations. The data asset management system 200 can build indexes based on graph databases or graph processing frameworks to improve data retrieval efficiency. Indexes can avoid traversing the entire graph to find data, and can quickly locate query nodes, relationships, or attributes. In its implementation, indexes can include entity-based indexes, relationship-based indexes, attribute-based indexes, hybrid indexes, etc.; this application does not impose specific limitations.

[0172] For example, indexes can be created for attributes such as a person's name and age to support querying person entities by name or age. Indexes can be created for "friend relationships" and "following relationships" to support querying relationships by relationship type. Full-text indexes can be created for article content to support keyword searches. Geospatial indexes can be created for the latitude and longitude coordinates of locations to support geolocation queries. Composite indexes can also be created as needed; for example, a composite index can be created for the combination of a person's name and age attributes to support querying person entities based on the combination of name and age. The above examples are for illustrative purposes only and are not intended to be specific limitations.

[0173] Optionally, entities in a knowledge graph can also be associated with their corresponding business data, allowing users to read the original business data based on the knowledge graph. Specifically, structured data is typically stored in a database, so a unique identifier for the entity (e.g., entity ID) can be used as a foreign key to associate the entity with table data in the business system. This way, when a user queries an entity, they can retrieve the corresponding data in the relevant business system based on the foreign key. Unstructured data is usually text data, which can be indexed using full-text indexing technology. The entity's identifier can be associated with the corresponding document in the business system. This way, when querying an entity, if the user needs the original text document, they can locate the text document in the business system through the index. Semi-structured data, such as XML and JSON files, is stored in the file system. In this case, the entity can be associated with its corresponding file path. When querying an entity, the associated path can be used to retrieve the semi-structured data. Thus, users can not only read data based on the knowledge graph but also perform information queries.

[0174] It should be understood that in the field of enterprise data asset management, data retrieval and information query needs are the primary needs of relevant business personnel. Data retrieval refers to retrieving data that meets certain conditions from raw business data based on user needs, such as extracting order records within a specific time period or extracting specific types of time records from log files. Information query refers to analyzing and inferring information from raw business data based on user needs to obtain the information required by the user. This information can be an analysis result or an inference result to meet the needs of daily business operations and decision-making, such as asset owner queries, organizational structure queries, personnel information queries, and equity structure queries. This type of information query does not require complete raw data; the system only needs to answer the question. Based on knowledge graphs, entities or relationships can be quickly located, supporting daily operational decisions. The data asset management method provided in this application can meet the above-mentioned data retrieval and information query needs, and the implementation process is simple and fast. Users only need to input a simple question, and the retrieval efficiency based on knowledge graphs is also very high, improving the user experience.

[0175] The above is based on Figures 3-6 The steps and procedures of the data asset management method provided in this application during the configuration phase are explained in detail below. Figure 7 The steps and procedures in the application phase will be explained.

[0176] Figure 7 This is a flowchart illustrating the steps of a data asset management method provided in this application during the application phase. This method can be applied to, for example... Figure 1 and Figure 2 In the data asset management system 200 shown, such as Figure 7 As shown, the method may include the following steps:

[0177] S710: Client 100 sends a query request to data asset management system 200.

[0178] This step can be performed by Figure 1 The question-and-answer device 220 in the embodiment is implemented.

[0179] In its implementation, a query request may include a question input by the user through client 100, which is used to read data and / or retrieve information. This question is in natural language text, such as "Give me the sales report for product A in Q1 2022" or "Which department does Zhang San work in?". It should be understood that traditional business systems require users to input database query language, such as SQL, when reading data. The data asset management method provided in this application allows users to input natural language when querying data assets, enabling non-technical users to easily read data and reducing their learning curve.

[0180] Data reading is used to retrieve raw business data from multiple business systems, such as financial statements and sales records. Information query refers to analyzing, summarizing, and inferring business data from multiple business systems to provide answers, such as asking which department Zhang San works in or whether client A has paid its prepayment for this year. This data is not directly stored in the database; it needs to be retrieved based on a knowledge graph, analyzed, and inferred to generate answers.

[0181] In practice, query requests may also carry other information, such as authentication information and format requirements. Authentication information is used to ensure that the user has permission to execute the query or to obtain the answer to the query. Authentication information may include username, password, token, etc., which are not specifically limited in this application. Format requirements include the format of the answer the user expects to obtain, such as a table or an XML file. The examples above are for illustration only and are not specifically limited in this application.

[0182] S720: Data Asset Management System 200 retrieves question entities.

[0183] This step can be performed by Figure 1 The question-and-answer device 220 in the embodiment is implemented.

[0184] In practice, the entity in a question refers to the entity contained within the question. An entity can be a single, referential transaction, object, or concept. Specifically, it can be a concrete transaction in the real world, such as a person, place, organization, or product, or it can be an abstract concept, such as an event, date, or number. For details, please refer to the description of entities mentioned above; it will not be repeated here.

[0185] In specific implementation, the data asset management system 200 can be based on Figures 3-6 The entity recognition naming model described in the example is used to determine the entities in the question. This ensures that the same entity recognition naming model is used when building the knowledge graph and when handling user queries, maintaining consistency in entity recognition standards during configuration and application phases, which helps reduce errors.

[0186] S730: Data Asset Management System 200 obtains question categories based on questions.

[0187] This step can be performed by Figure 1 The question-and-answer device 220 in the embodiment is implemented.

[0188] In its implementation, the data asset management system 200 can pre-maintain various question categories. For example, question categories may include data retrieval type and information query type. Data retrieval type refers to retrieving data that meets certain conditions from the original business data according to user needs. Information query type refers to analyzing and inferring from the original business data according to user needs to obtain the analysis results or inference results required by the user. Data retrieval type and information query type can also be further subdivided according to specific application scenarios, which are not specifically limited in this application.

[0189] For example, data retrieval types may include conditional queries, statistical queries, and relational queries. Conditional queries refer to retrieving data based on specific conditions. Statistical queries refer to retrieving certain data and completing statistics, such as calculating averages, sums, and other information. Relational queries refer to retrieving certain data and completing relational analysis, such as comparing the differences between two things. The above examples are for illustration purposes. Data retrieval types may include many more types, and this application does not make any specific limitations.

[0190] Information query types can include comprehensive analysis queries, predictive queries, and trend analysis queries. Comprehensive analysis queries refer to common analytical queries, such as analyzing the relationship between two attributes of an entity. Predictive queries refer to predicting or inferring from data, such as whether an attribute of an entity will change within a certain time period. Trend analysis queries refer to analyzing the changing trends of data, such as what the changing trend of an entity's attributes is within a certain time period. The above examples are for illustration only; information query types can include many more categories, which are not specifically limited in this application.

[0191] S740: Data Asset Management System 200 obtains question templates based on question categories, and determines graph query statements based on question entities and question templates.

[0192] This step can be performed by Figure 1 The question-and-answer device 220 in the embodiment is implemented.

[0193] In its implementation, the data asset management system 200 can pre-maintain multiple question templates. A question category can include various question templates, and each question template corresponds to a query statement template. The query statement template defines the basic structure and syntax of the query statement and contains variables that need to be filled in. Thus, after obtaining the question entity from the user-input question and determining its category, the system retrieves the multiple question templates included in that category, identifies the corresponding question template from these templates, and fills in the query statement template based on the question entity to obtain the graph query statement corresponding to the question. It should be understood that determining the question category based on the question entity before obtaining the question template allows for preliminary screening of question templates, improving the system's processing efficiency.

[0194] In practice, the question template can include keywords and fill areas. The system can match the question with the keywords in the question template to determine the corresponding question template. Keywords can be single characters or words, or sentence structures composed of multiple words. For example, the question template could be: "Which [entities] are within [time period] [conditions]?", where the keywords are "which... are within...", and the fill areas are [entities], [time period], and [conditions]. Based on the sentence structure, a corresponding query template can be configured. The variables in this query template are the aforementioned entities, time periods, and conditions. If the user's question is "In the past year, which products have sales exceeding $50 million?", this question includes the keyword "which... are within...". This question can then be matched with the question template to obtain the corresponding query template. By filling in the question and the question entities, a graph query statement can be obtained.

[0195] For example, for data retrieval types, a sample question template for a conditional query could be: "Which [entities] meet the [condition] within the [time period]?", a sample question template for a statistical query could be: "What is the average value of the [attribute] in the [entity]?", and a sample question template for a relational query could be: "What is the relationship between [entity 1] and [entity 2]?". For information query types, a sample question template for a comprehensive analysis query could be: "What is the relationship between [attribute 1] and [attribute 2] in the [entity]?", a sample question template for a predictive query could be: "How will the [attribute] of the [entity] change within the future [time period]?", and a sample question template for a trend analysis query could be: "What is the trend of the [attribute] of the [entity] over the past [time period]?". The above sample question templates are for illustrative purposes only and are not intended to impose specific limitations in this application.

[0196] It should be understood that the keywords in the question template and the variables that need to be filled in the query template are all obtained based on the question. In the S720, the data asset management system 200 can perform entity recognition on the question to obtain the question entity. If the keywords in the question template and the variables that need to be filled in the query template include content other than the question entity, such as time period, conditions, and attributes in the above example, the data asset management system 200 can also perform semantic analysis on the question to identify the keywords required by the question template and the variables required by the query template. The recognition process can also be implemented based on natural language processing technology, similar to the entity recognition and naming process described above, which will not be repeated here.

[0197] S750: The Data Asset Management System 200 uses graph query statements to search in the knowledge graph and obtain query results.

[0198] This step can be performed by Figure 1 The knowledge graph device 210 in the embodiment is implemented.

[0199] In its implementation, the data asset management system 200 uses graph databases or graph processing frameworks to manage and query knowledge graphs, such as Neo4j. It executes graph query statements based on the graph database engine or graph processing framework engine to obtain the corresponding query results. For data retrieval type queries, the query results can be business data, such as table data, XML files, or text data. For information query type queries, the query results can be entity attribute values, relationships, and other information. These query results require further processing to generate the answers needed by the user.

[0200] S760: Data Asset Management System 200 generates answers to questions based on query results.

[0201] This step can be performed by Figure 1 The question-and-answer device 220 in the embodiment is implemented.

[0202] Optionally, for data retrieval type questions, the query result is business data. In this case, based on the query request in S710 carrying the user's required format, the answer to the question can be generated based on the query result and the format requirements. For example, it can be converted into text, graphics, tables, etc., suitable for user reading.

[0203] Optionally, for questions of the information query type, the query results include information such as the attribute values ​​and relationships of entities. The data asset management system 200 can also maintain multiple answer templates. Each answer template can include keywords and fill areas. The fill areas are filled based on the query results to obtain the answer corresponding to the question. A question template can correspond to one or more answer templates, and an answer template can also correspond to one or more question templates. This application does not make specific limitations.

[0204] In practical implementation, the keywords in the answer template can be a single character, a single word, or a sentence structure composed of multiple characters and multiple words. For example, if the question template is: "Who is responsible for managing the [data entity] of [company entity]?", the answer template is: "The [data entity] of [company entity] is managed by [personal entity]". If the question is: "Who is responsible for managing the customer data of Company A?", the query statement template based on the question template can generate a graph query statement. Assuming that the nodes in the knowledge graph include company nodes, employee nodes, and data nodes, and there are relationships between the nodes, then the graph query statement can first match the company node named Company A, then find the data node managed by this company that is of the type customer data, and then find the personnel nodes associated with these two nodes, returning the company name, data type, and person in charge's name. By querying nodes and edges, the query results can be obtained. Based on the query results, the answer template is populated, and the answer is: "The customer data of Company A is managed by Zhang San". The above examples are for illustration only, and this application does not impose specific limitations.

[0205] It should be noted that the system can also be configured with specific dialogue templates to obtain supplementary information from users. If no matching question template is available, or if some areas in the answer template cannot be filled, a supplementary message can be sent to the user based on the dialogue template to obtain a new question. Then, the question template can be determined based on the old and new questions, or the answer template can be filled, and so on, continuously communicating with the user in conjunction with the dialogue templates. For example, if a user asks, "I want to see the sales report for product A," the system can ask, "Which year do you need to see?" If the user replies that they need the 2022 report, the system can obtain the graph query statement based on the question template and then obtain the answer. The above example is for illustration only and is not intended to be specific.

[0206] S770: Data asset management system 200 sends the answer to client 100.

[0207] This step can be performed by Figure 1 The question-and-answer device 220 in the embodiment is implemented.

[0208] For example, Figure 8 This is an example diagram of a client interface provided in this application. The client 100 can display to the user. Figure 8 The interface shown is used to receive user-input questions and display answers to them. It should be understood that... Figure 8 This is an exemplary interface, and this application does not limit it. For example... Figure 8 As shown, the interface may include an input area 810, a question display area 820, and an answer display area 830.

[0209] The input area 810 is used to obtain questions input by the user. The user can input questions in natural language or by voice input in this area. This application does not make any specific limitations.

[0210] The question display area 820 is used to display the questions entered by the user in the past.

[0211] Answer display area 830 is used to display the answers obtained by the system based on the user's input question, as shown in the example line. Figure 8 Examples of multiple sets of questions and answers are provided. For instance, if a user asks, "Who is responsible for managing the customer data of Company A?", the system can execute steps S710-S760 to first identify the entities in the question, then determine the question category based on the question entities, obtain multiple question templates corresponding to the question category, match the question entities with the multiple question templates to obtain the corresponding question template, and then, based on the query statement template corresponding to the question template, populate the query statement template based on the question entities to obtain the corresponding graph query statement. The graph query statement is then used to search the knowledge graph to obtain the query results. For example, first search for nodes of Company A, then search for customer data nodes associated with nodes of Company A, and then determine that the manager node associated with the customer data node is "Zhang San". Alternatively, if the attributes of the customer data node include "manager", the manager can also be determined as "Zhang San" based on the attributes; this application does not impose specific limitations. After obtaining the query results, the system can populate the answer template based on the query results, based on the answer template corresponding to the question template, to generate... Figure 8 The answer shown is: "Zhang San is responsible for managing Company A's customer data."

[0212] Similarly, if a user inputs the question, "What was the result of the most recent customer data quality assessment of Company A?", the system can generate the answer, following the process described in S710 to S760 above, "The most recent customer data quality assessment of Company A scored 85 points." The system's generation process will not be described again here.

[0213] Furthermore, both questions described above are information query questions. The data asset management system provided in this application can also read raw data from the business system and present it to the user, meeting the data reading needs under data asset management. For example, such as... Figure 8 As shown, users can enter the question: "Give me the most recent customer data quality assessment table for all customers." The system can execute steps S710 to S760 to retrieve the link to the quality assessment table from the business system. Users can directly click the link to download and view the assessment table, thus meeting their usage needs.

[0214] In summary, the data asset management method provided in this application establishes communication connections with multiple business systems and builds a knowledge graph based on business data from these systems. This allows information from different departments and data sources to be interconnected and queried. When a user sends a question to the data asset management system via a client, the system can search the knowledge graph based on the question entity contained in the question to obtain query results related to that entity. Then, it generates the answer to the question based on the query results. This eliminates the need for users to access multiple business systems in a chain to obtain answers; they only need to input a simple question to get the answer, improving the user experience. Furthermore, structured and semi-structured data in the business data can have entities extracted based on their fixed format characteristics to build a knowledge graph, while unstructured data can have entities identified based on its semantic features to build a knowledge graph. The resulting knowledge graph encompasses the knowledge features contained in structured, semi-structured, and unstructured data from multiple business systems, providing users with more comprehensive answers. This not only enables the reading of business data from multiple systems but also allows for complex information queries, improving the work efficiency of business personnel.

[0215] The data asset management method provided in this application has been described in detail above. The following section will combine... Figure 9 This application provides a description of the data asset management system provided. This data asset management system is... Figures 1-8 The data asset management system described in the document 200.

[0216] Figure 9 This is a schematic diagram of the structure of a data asset management system provided in this application, such as... Figure 9 As shown, the data asset management system 200 may include a question-answering device 220 and a knowledge graph device 210. The question-answering device 220 may include an acquisition unit 221, a retrieval unit 222, and a generation unit 223. The knowledge graph device 210 may include a graph building unit 211. It should be understood that the question-answering device 220 and the knowledge graph device 210 are exemplary divisions, and the data asset management system 200 may also not have such device divisions. This application does not impose any specific limitations.

[0217] The acquisition unit 221, retrieval unit 222, generation unit 223, and map building unit 211 can all be implemented in software or in hardware. For example, the implementation of the acquisition unit 221 will be described below. Similarly, the implementation of the retrieval unit 222, generation unit 223, and map building unit 211 can refer to the implementation of the acquisition unit 221.

[0218] As an example of a software functional unit, the acquisition unit 221 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the aforementioned computing instance may be one or more. For example, the acquisition unit 221 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same Availability Zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.

[0219] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same VPC or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0220] As an example of a hardware functional unit, the acquisition unit 221 may include at least one computing device, such as a server. Alternatively, the acquisition unit 221 may be implemented using a central processing unit (CPU), an ASIC, or a PLD. The PLD may be implemented using a CPLD, FPGA, GAL, DPU, NPU, SoC, offload card, accelerator card, or any combination thereof.

[0221] The multiple computing devices included in the acquisition unit 221 can be distributed in the same region or in different regions. Similarly, the multiple computing devices included in module A can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices included in the document comparison unit 2124 can be distributed in the same VPC or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, GALs, DPUs, NPUs, SoCs, offloading cards, and accelerator cards.

[0222] It should be noted that, in other embodiments, the steps implemented by the acquisition unit 221, retrieval unit 222, generation unit 223 and map building unit 211 can be specified as needed. The acquisition unit 221, retrieval unit 222, generation unit 223 and map building unit 211 respectively implement different steps in the data asset management method to realize all the functions of the data asset management system 200.

[0223] The functions of the acquisition unit 221, the retrieval unit 222, the generation unit 223, and the map building unit 211 are explained below.

[0224] The acquisition unit 221 is used to receive a query request input by the user, which includes a question expressed by the user in natural language. Specifically, it is used to implement... Figure 7 S710 and its optional steps in the embodiments.

[0225] The retrieval unit 222 is used to determine the question entities contained in the question and the question category of the question. The question category includes data reading type and information query type. The data reading type question is used to read business data from the multiple business systems, and the information query type question is used to obtain information query results. Specifically, it is used to implement... Figure 7 S720 to S750 and their optional steps in the embodiments.

[0226] Generation unit 223 is used to determine the answer to the question based on the query results. Specifically, it is used to implement... Figure 7 S760 to S770 and their optional steps in the embodiments.

[0227] In one possible implementation, the retrieval unit 222 is used to obtain multiple question templates corresponding to the question category, match the question with the multiple question templates, and determine the question template corresponding to the question. The retrieval unit 222 is also used to obtain a query statement template corresponding to the question based on the question template, wherein one question template corresponds to one query statement template. The retrieval unit 222 is used to populate the query statement template based on the question entity to obtain a graph query statement corresponding to the question. When the question category is a query type question, the retrieval unit 222 is used to search for business data corresponding to the question entity in the knowledge graph using the graph query statement to obtain the query result. Finally, when the question category is an information query type question, the retrieval unit 222 is used to search for related entities, relationships, and attributes of the question entity in the knowledge graph using the graph query statement to obtain the query result. Specifically, this is used to implement... Figure 7 The embodiments include steps S720 to S750 and their optional steps.

[0228] In one possible implementation, structured data includes table data stored in a fixed format, semi-structured data includes non-table data stored in a fixed format, and unstructured data includes text data with varying formats. A graph building unit 211 is used to acquire business data from multiple business systems. The graph building unit 211 is used to acquire first entity information and first relationship information based on the fixed format features of structured and semi-structured data. The graph building unit 211 is used to perform named entity recognition on entities in unstructured data based on semantic features to obtain second entity information and second relationship information. The graph building unit 211 is used to build nodes in the knowledge graph based on the first and second entity information, and to build edges between nodes in the knowledge graph based on the first and second relationship information, thus building the knowledge graph. Specifically, it is used for implementation... Figure 3 S310 to S370 and their optional steps in the embodiments.

[0229] In one possible implementation, the graph building unit 211 is used to determine multiple entities in the structured data based on the fields and field values ​​in the structured data, and obtain the first entity information of the structured data. The multiple entities in the structured data include fields and field values ​​with keywords, fields and field values ​​with attribute types of preset types, and fields and field values ​​contained in the entity template. The graph building unit 211 is used to determine the first relationship information between the multiple entities in the structured data based on the relationship between the fields and field values, and the relationship between the fields.

[0230] In one possible implementation, the graph building unit 211 is used to determine multiple entities of the semi-structured data based on the tags and attributes of the tags in the semi-structured data, obtain the first entity information of the semi-structured data, and determine the first relationship information between the multiple entities of the semi-structured data based on the relationship between tags and attributes and the nesting relationship between tags, wherein the tags in the semi-structured data are used to separate the hierarchical structure and relationships of the data; or, the graph building unit 211 is used to determine multiple entities of the semi-structured data based on the key-value pairs in the semi-structured data, obtain the first entity information of the semi-structured data, and determine the second relationship information between the multiple entities of the semi-structured data based on the correspondence between keys and values ​​and the nesting relationship between the first key and the second key, wherein the second key is a subkey of the first key, and the second key is an attribute or member of the object described by the first key.

[0231] In one possible implementation, the graph building unit 211 is used to perform named entity recognition on entities in unstructured data based on an entity recognition naming model to obtain second entity information and second relation information. The entity recognition naming model includes a word embedding layer, a feature extraction layer, and a feature classification layer. The word embedding layer is used to convert unstructured data into a word vector matrix with word-level features. The feature extraction layer is used to extract semantic features from the word vector matrix to obtain a sentence vector matrix with sentence-level features. The feature classification layer is used to classify the word segments in the unstructured data according to the sentence vector matrix to obtain second entity information and second relation information.

[0232] In one possible implementation, the word embedding layer includes a text input layer and a vector transformation layer. The text input layer is used to perform word segmentation on unstructured data to obtain multiple words. The vector transformation layer inputs the multiple embedding vectors of each word into the transformer structure to extract the word-level features of each word and obtain a word vector matrix. The multiple embedding vectors include word embedding vectors, sentence embedding vectors, and position embedding vectors.

[0233] In one possible implementation, the feature extraction layer includes a context feature extraction network and a multi-head attention network. The context feature extraction network is used to extract context features from unstructured data based on the word vector matrix, generating a multi-dimensional feature sequence containing context features. The multi-head attention network is used to determine the attention score of each element in the multi-dimensional feature sequence to obtain a sentence vector matrix, where the attention score is used to indicate the degree of influence of each element on entity recognition and naming.

[0234] In one possible implementation, the feature classification layer includes an annotation network and an optimization network. The annotation network is used to generate a label prediction result for each word based on the sentence vector matrix. The optimization network is used to adjust the label of each word in the unstructured data based on the label prediction result for each word, combined with sentence semantics and contextual information, to obtain second entity information and second relation information in the unstructured data. The label is used to indicate whether the word is an entity and the category to which the entity belongs.

[0235] In one possible implementation, retrieval unit 222 is used to determine the question category of the question based on the question entity; retrieval unit 222 is used to obtain multiple question templates corresponding to the question category, and obtain the question template corresponding to the question from the multiple question templates; retrieval unit 222 is used to obtain the query statement template corresponding to the question based on the question template corresponding to the question, wherein one question template corresponds to one query statement template; retrieval unit 222 is used to fill the query statement template based on the question entity to obtain the graph query statement corresponding to the question; retrieval unit 222 is used to perform retrieval in the knowledge graph using the graph query statement to obtain the query results.

[0236] In summary, the data asset management system provided in this application establishes communication connections with multiple business systems and builds a knowledge graph based on business data from these systems. This allows information from different departments and data sources to be interconnected and queried. When a user sends a question to the data asset management system via a client, the system can search the knowledge graph based on the question entity contained in the question to obtain query results related to that entity. Then, it generates the answer to the question based on the query results. This eliminates the need for users to chain through multiple business systems to obtain answers; they only need to input a simple question to get the answer, improving the user experience. Furthermore, structured and semi-structured data in the business data can have entities extracted based on their fixed format characteristics to build a knowledge graph, while unstructured data can have entities identified based on its semantic features to build a knowledge graph. The resulting knowledge graph encompasses the knowledge features contained in structured, semi-structured, and unstructured data from multiple business systems, providing users with more comprehensive answers. This not only enables the reading of business data from multiple systems but also allows for complex information queries, improving the work efficiency of business personnel.

[0237] The data asset management method and data asset management system provided in this application have been described in detail above. The following section will combine... Figures 10-12 The computing device provided in this application will be explained.

[0238] Figure 10 This is a schematic diagram of the structure of a computing device provided in this application, such as... Figure 10 As shown, the computing device 1000 includes a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. The processor 1004, memory 1006, and communication interface 1008 communicate with each other via the bus 1002. The computing device 1000 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 1000. The computing device 1000 can be any of the aforementioned... Figures 1-9 The data asset management system 200 in this embodiment.

[0239] Bus 1002 can be a Peripheral Component Interconnect Express (PCIe) bus, an Extended Industry Standard Architecture (EISA) bus, a Unified Bus (Ubus or UB), a Compute Express Link (CXL) bus, a Cache Coherent Interconnect for Accelerators (CCIX) bus, etc. The Unified Bus is also known as the Lingqu Bus. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 10 The bus 1002 is represented by only one line, but this does not mean that there is only one bus or one type of bus. The bus 1002 may include a path for transmitting information between various components of the computing device 1000 (e.g., memory 1006, processor 1004, communication interface 1008). The unified bus may also be referred to as the Lingqu bus.

[0240] The processor 1004 may include any one or more of the following computing devices: central processing unit (CPU), graphics processing unit (GPU), microprocessor (MP) or digital signal processor (DSP), ASIC, FPGA, CPLD, NPU, SoC, offload card, accelerator card, etc.

[0241] The memory 1006 may include volatile memory, such as random access memory (RAM). The processor 1004 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD). Furthermore, the memory 1006 may also be implemented using storage class memory (SCM), phase change memory (PCM), or other types of storage media.

[0242] It is worth noting that the same type of storage medium can be configured in the same computing device to realize the function of memory 1006, or two or more types of storage media can be configured to realize the function of memory 1006. This application does not limit this.

[0243] The memory 1006 stores executable program code, and the processor 1004 executes the executable program code to implement the functions of the aforementioned data asset management system 200, including... Figure 9 The acquisition unit, retrieval unit, and generation unit shown herein implement the application stage steps of the data asset management method provided in this application. That is, the memory 1006 stores instructions for executing the data asset management method.

[0244] Alternatively, the memory 1006 stores executable code, which the processor 1004 executes to implement the functions of the acquisition unit, retrieval unit, generation unit, and map building unit, thereby realizing the application stage and configuration stage steps of the data asset management method provided in this application. That is, the memory 1006 stores instructions for executing the data asset management method.

[0245] The communication interface 1008 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 1000 and other devices or communication networks.

[0246] It is worth noting that computing device 1000 can correspond to, for example... Figures 3 to 9 The method described herein refers to the operational steps performed by the corresponding subject, and can be corresponding to... Figure 9 The above and other operations and / or functions of each unit in the data asset management system 200 shown are respectively for the purpose of implementing Figures 3 to 9 For the sake of brevity, the corresponding processes of each method in the code will not be elaborated here.

[0247] As one possible implementation, the computing device 1000 may also include a chip system, which includes a processor and a power supply circuit. The power supply circuit supplies power to the processor, and the processor executes the operation steps corresponding to the data asset management method. For simplicity, further details are omitted here. The processor can be implemented using a CPU, or it can be implemented using computing devices or AI chips such as GPUs, DPUs, NPUs, XPUs, SoCs, offloading cards, or accelerator cards.

[0248] As one possible implementation, the computing device 1000 may include various types of processors 1004, meaning the computing device 1000 is a heterogeneous device. For example, the computing device 1000 may include a CPU and a GPU, and at least one of the processors 1004 may execute the operation steps corresponding to the data asset management method. For the sake of brevity, further details will not be elaborated here.

[0249] It is worth noting that the processor in a chip system can correspond to, for example... Figures 3 to 9 The method described herein refers to the operational steps performed by the corresponding subject, and can be corresponding to... Figure 9 The above and other operations and / or functions of each unit in the data asset management system 200 shown are respectively for the purpose of implementing Figures 3 to 9 For the sake of brevity, the corresponding processes of each method in the code will not be elaborated here.

[0250] This application also provides a cluster of computing devices. For example... Figure 11 As shown, Figure 11 This is an example diagram of a computing device cluster provided in this application, which includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0251] like Figure 11 As shown, the computing device cluster includes at least one computing device 1000. The memory 1006 of one or more computing devices 1000 in the computing device cluster may store the same instructions for executing data asset management methods.

[0252] In some possible implementations, the memory 1006 of one or more computing devices 1000 in the computing device cluster may also store partial instructions for executing the data asset management method. In other words, a combination of one or more computing devices 1000 can jointly execute instructions for executing the data asset management method.

[0253] It should be noted that the memory 1006 in different computing devices 1000 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the data asset management system 200. That is, the instructions stored in the memory 1006 of different computing devices 1000 can implement the functions of one or more modules within the data asset management system 200.

[0254] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 12 One possible implementation is shown. For example... Figure 12 As shown, Figure 12 This is a schematic diagram of another computing device cluster structure provided in this application. Two computing devices, 1000A and 1000B, are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this possible implementation, the memory 1006 in computing device 1000A stores instructions for executing the map building unit. Simultaneously, the memory 1006 in computing device 1000B stores instructions for executing the acquisition unit, retrieval unit, and generation unit.

[0255] Figure 12 The connection method between the computing device clusters shown can be based on the fact that the data asset management method provided in this application has a configuration stage and an application stage. The configuration stage requires the establishment of a knowledge graph, which requires a certain amount of storage space. The application stage requires retrieval based on the knowledge graph. Therefore, it is considered to distribute the unit modules on different computing devices. The function of the graph establishment unit in the configuration stage is handed over to computing device 1000A, and the functions of the acquisition unit, retrieval unit and generation unit in the application stage are handed over to computing device 1000B.

[0256] It should be understood that Figure 12 The functions of computing device 1000A shown can also be performed by multiple computing devices 1000. Similarly, the functions of computing device 1000B can also be performed by multiple computing devices 1000.

[0257] In some possible implementations, the memory 1006 of one or more computing devices 1000 in the computing device cluster may also store partial instructions for executing the data asset management method. In other words, a combination of one or more computing devices 1000 can jointly execute instructions for executing the data asset management method.

[0258] It should be noted that the memory 1006 in different computing devices 1000 within the computing device cluster can store different instructions for executing some functions of the data asset management system. That is, the instructions stored in the memory 1006 of different computing devices 1000 can implement the functions of one or more devices within the data asset management system.

[0259] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to execute a data asset management method.

[0260] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform a data asset management method, or instruct the computing device to perform a data asset management method.

[0261] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of this application.< / product> < / order> < / product> < / order> < / person> < / person> < / person>

Claims

1. A data asset management method, characterized in that, The method is applied to a data asset management system, which establishes communication connections with multiple business systems. The method includes: Receive a query request input by the user, the query request including a question expressed by the user in natural language; The question entity contained in the question and the question category of the question are determined. The question category includes data reading type and information query type. The data reading type question is used to read business data in the multiple business systems, and the information query type question is used to obtain information query results. Based on the question category and the question entity, a search is performed in the knowledge graph to obtain query results. The knowledge graph in the data asset management system is built based on the business data, which includes structured data, semi-structured data, and unstructured data. The answer to the question is determined based on the query results.

2. The method according to claim 1, characterized in that, Based on the question category and the question entity, a search is performed in the knowledge graph to obtain the following query results: Obtain multiple question templates corresponding to the question category, match the question with the multiple question templates, and determine the question template corresponding to the question; Based on the question template corresponding to the question, obtain the query statement template corresponding to the question, wherein one question template corresponds to one query statement template; The query statement template is populated based on the question entity to obtain the graph query statement corresponding to the question; When the question type is a query type question, the graph query statement is used to search for the business data corresponding to the question entity in the knowledge graph to obtain the query result; When the question type is an information query type, the graph query statement is used to search for the related entities, relationships, and attributes of the question entity in the knowledge graph to obtain the query results.

3. The method according to claim 1 or 2, characterized in that, The structured data includes table data stored in a fixed format, the semi-structured data includes non-table data stored in a fixed format, and the unstructured data includes text data with an irregular format. Before receiving the user-input query request, the method further includes: The business data is obtained from the multiple business systems; Based on the fixed format characteristics of the structured data and the semi-structured data, first entity information and first relationship information are obtained; Based on the semantic features of the unstructured data, named entity recognition is performed on the entities in the unstructured data to obtain second entity information and second relationship information; Nodes are established in the knowledge graph based on the first entity information and the second entity information, and edges are established between the nodes based on the first relationship information and the second relationship information, thereby establishing the knowledge graph.

4. The method according to claim 3, characterized in that, The acquisition of the first entity information and the first relationship information based on the fixed format characteristics of the structured data includes: Based on the fields and field values ​​in the structured data, multiple entities of the structured data are determined, and the first entity information of the structured data is obtained. The multiple entities of the structured data include fields and field values ​​with keywords, fields and field values ​​with attribute types of preset types, and fields and field values ​​contained in the entity template. Based on the relationship between the fields and field values, and the relationship between the fields, the first relationship information between multiple entities in the structured data is determined.

5. The method according to claim 3 or 4, characterized in that, The acquisition of the first entity information and the first relationship information based on the fixed format characteristics of the semi-structured data includes: Based on the tags and attributes in the semi-structured data, multiple entities in the semi-structured data are identified to obtain the first entity information of the semi-structured data. Based on the relationship between the tags and attributes and the nesting relationship between the tags, the first relationship information between the multiple entities in the semi-structured data is determined, wherein the tags in the semi-structured data are used to separate the hierarchical structure and relationships of the data; or, Based on the key-value pairs in the semi-structured data, multiple entities of the semi-structured data are determined, and the first entity information of the semi-structured data is obtained. Based on the correspondence between keys and values ​​and the nesting relationship between the first key and the second key, the second relationship information between the multiple entities of the semi-structured data is determined, wherein the second key is a subkey of the first key, and the second key is an attribute or member of the object described by the first key.

6. The method according to any one of claims 2 to 5, characterized in that, The step of performing named entity recognition on entities in the unstructured data based on the semantic features of the unstructured data to obtain second entity information and second relation information includes: Named entity recognition is performed on entities in the unstructured data based on an entity recognition naming model to obtain the second entity information and the second relation information. The entity recognition naming model includes a word embedding layer, a feature extraction layer, and a feature classification layer. The word embedding layer is used to convert the unstructured data into a word vector matrix with word-level features. The feature extraction layer is used to extract semantic features from the word vector matrix to obtain a sentence vector matrix with sentence-level features. The feature classification layer is used to classify the word segments in the unstructured data according to the sentence vector matrix to obtain the second entity information and the second relation information.

7. A data asset management system, characterized in that, The data asset management system establishes communication connections with multiple business systems, and the data asset management system includes: The acquisition unit is used to receive a query request input by a user, the query request including a question expressed by the user in natural language; The retrieval unit is used to determine the question entity contained in the question, and to perform a retrieval in the knowledge graph based on the question entity to obtain query results. The knowledge graph in the data asset management system is built based on business data from the multiple business systems. The business data includes structured data, semi-structured data, and unstructured data. When the question is a data reading type question, the query results include the business data corresponding to the question entity. When the question is an information query type question, the query results include the associated entities, relationships, and attributes of the question entity. A generation unit is used to determine the answer to the question based on the query results.

8. A computing device, characterized in that, The computing device includes a processor and memory; The processor is configured to execute instructions stored in the memory to cause the computing device to perform the operational steps of the method as described in any one of claims 1 to 6.

9. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the operational steps of the method as described in any one of claims 1 to 6.

10. A computer program product containing instructions, characterized in that, When the instruction is executed by a computing device or a cluster of computing devices, it causes the computing device or cluster of computing devices to perform the operational steps of the method as described in any one of claims 1 to 6.