Metadata management method and computing device

By obtaining and converting metadata in management nodes, and dynamically integrating metadata relationships using graph analysis algorithms and machine learning models, the problem of inefficient metadata management across systems is solved, and efficient and secure metadata management is achieved.

CN120336434APending Publication Date: 2025-07-18XFUSION DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510241947.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing data management systems are difficult to implement cross-system metadata management and cannot dynamically adapt to changes in metadata relationships, resulting in low metadata management efficiency.

Method used

The first standard metadata is obtained through the management node, the relationship network is determined according to the graph analysis algorithm model, the relationship relationship between each standard metadata is dynamically integrated, the metadata type is predicted using machine learning model and the metadata of different formats is converted into a unified format through mapping rules, and the relationship network and query results are displayed using terminal devices.

Benefits of technology

Cross-system metadata management is realized, the efficiency and accuracy of metadata management is improved, data consistency and security are ensured, and user experience is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336434A_ABST
    Figure CN120336434A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a metadata management method and computing equipment, belongs to the field of computing, and is used for improving metadata management efficiency. The method is applied to a metadata management system, the metadata system comprises a management node and a plurality of data source nodes, and semantics and / or formats of metadata corresponding to the data source nodes are different. The method is executed by a management node and comprises the following steps: acquiring first standard metadata; the first standard metadata is determined according to metadata corresponding to the first data source node; determining a relational network according to the first standard metadata; wherein the relational network is used for indicating an association relationship between the first standard metadata and other standard metadata based on different dimensions; the incidence relation represents the tightness degree of mutual connection between the standard metadata; the other standard metadata is any number of standard metadata acquired from the plurality of data source nodes except the first data source node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computing, and particularly to a method for metadata management and a computing device. Background Art

[0002] In the modern data management field, metadata, as the "data" describing data, plays a crucial role in summarizing, locating, and managing data. With the explosive growth of the data volume and data sources, the management of metadata has become increasingly complex. Especially in a distributed environment, there are significant differences in the data standards, formats, or structures of different systems, which further increases the difficulty of metadata management.

[0003] In the related art, the widely used data management systems or data governance platforms only provide basic metadata management functions and manage metadata based on static mappings. However, the above methods are not only difficult to implement cross-system metadata management, but also unable to dynamically adapt to changes in metadata relationships and meet the requirements of dynamic management of complex metadata, resulting in low metadata management efficiency. Summary of the Invention

[0004] The embodiments of the present application provide a method for metadata management and a computing device, which are used to meet the requirements of dynamic management of complex metadata and improve the efficiency of metadata management. The technical solution is as follows:

[0005] In a first aspect, a method for metadata management is provided. The method is applied to a metadata management system, which includes a management node and multiple data source nodes, and the semantics and / or formats of the metadata corresponding to the multiple data source nodes are different; the method is executed by the management node and includes: obtaining first standard metadata; the first standard metadata is in a standard format recognizable by the metadata management system; the first standard metadata is determined according to the metadata corresponding to the first data source node; the first data source node is any one of the multiple data source nodes; according to the first standard metadata, determining a relationship network; wherein, the relationship network is used to indicate the association relationship between the first standard metadata and other standard metadata based on different dimensions; the association relationship represents the degree of tight connection between each standard metadata; the other standard metadata is any standard metadata obtained from multiple data source nodes other than the first data source node.

[0006] As can be seen from the above, the management node determines the relationship network according to the first standard metadata corresponding to the first data source node, which can realize the dynamic integration of the association relationships between each standard metadata, meet the dynamic management requirements of metadata in a distributed environment, and thus improve the efficiency of metadata management.

[0007] In a possible implementation, the relationship network includes a first relationship network, which is used to indicate the association relationship based on the first dimension between the first standard metadata and other standard metadata; the first dimension is any one of different dimensions; the first dimension includes a data classification dimension, a content dimension, a time dimension, or a source dimension; determining the relationship network according to the first standard metadata includes: determining the association relationship based on the first dimension between the first standard metadata and other standard metadata based on a graph analysis algorithm model; determining the first relationship network according to the association relationship based on the first dimension between the first standard metadata and other standard metadata.

[0008] As can be seen from the above, the graph analysis algorithm model can process data of complex and multi-dimensional relationship networks, making it more convenient to add or delete standard metadata, and can improve the operation efficiency of management nodes. The graph analysis algorithm model can also reflect the association relationship between standard metadata in real time and dynamically update the relationship network, thereby improving the efficiency of metadata management.

[0009] In a possible implementation, the metadata management system further includes a terminal device, which is used to display the relationship network and / or query results; the query results include the target standard metadata determined by the metadata management system according to the query request sent by the receiving terminal; the target standard metadata is used to indicate any number of standard metadata in the metadata management system.

[0010] As can be seen from the above, by using the visualization tool of the terminal device to display the multi-dimensional association relationship of standard metadata, it is ensured that users can efficiently obtain the required metadata. It improves the convenience and practicality of the metadata management system and enhances the user experience.

[0011] In a possible implementation, obtain the metadata corresponding to the first data source node; the metadata corresponding to the first data source node is obtained by extracting the data from the first data source node; determine the first standard metadata according to the metadata corresponding to the first data source node and the mapping rule; the mapping rule is used to indicate the mapping relationship between the fields of the metadata and the fields of the standard metadata.

[0012] As can be seen from the above, since the data in each data source node is of different types and / or structures, the semantics and / or formats of the metadata of the data source node extracted by the management node are not unified. The management node converts the metadata corresponding to the first data source node into the first standard metadata in a unified format through the mapping rule, enabling the metadata management system to recognize the metadata, which is beneficial to the subsequent update of the relationship network according to the metadata.

[0013] In a possible implementation manner, based on a machine learning model, predict the type of metadata corresponding to the first data source node, and determine a target mapping rule; the target mapping rule is applicable to the type of metadata corresponding to the first data source node.

[0014] As can be seen from the above, there is no need for the management node to determine the type of metadata corresponding to the first data source node after extracting the metadata corresponding to the data in the first data source node; the management node then determines the target mapping rule according to the type of metadata corresponding to the first data source node. Therefore, not only is the process of obtaining the first standard metadata shortened, but also the accuracy of obtaining the first standard data can be improved.

[0015] In a possible implementation manner, generate a first mapping log; the first mapping log is used to record the mapping process from the metadata corresponding to the first data source node to the first standard metadata; according to the first mapping log, determine data inconsistency and / or mapping anomalies between the metadata corresponding to the first data source node and the first standard metadata.

[0016] As can be seen from the above, the management node's recording of the mapping log enables the metadata management system to track the mapping process of the metadata corresponding to each data source node according to the mapping log, identify data inconsistencies or mapping anomalies, thereby ensuring the consistency and standardization of multi-source metadata.

[0017] In a possible implementation manner, when the first data source node includes structured data, based on the database system of the first data source node, extract the metadata corresponding to the first data source node; when the first data source node includes unstructured data, based on a natural language processing (NLP) model and a computer vision (CV) model, extract the metadata corresponding to the first data source node.

[0018] As can be seen from the above, since the data in each data source node is of different structures, adopting different metadata extraction methods for structured and unstructured data can ensure the accuracy of metadata extraction and improve the efficiency of obtaining the metadata corresponding to the data source node.

[0019] In a possible implementation manner, preprocess the data in the first data source node; the preprocessing includes removing duplicate data, filling in missing values, or consistency checking.

[0020] As can be seen from the above, the management node's preprocessing of the data in the first data source node can ensure the accuracy of extracting the metadata corresponding to the first data source node.

[0021] In a possible implementation manner, according to the type of the first standard metadata, determine the permission information of the first standard metadata; the permission information is used to indicate the users who are authorized to access the first standard metadata.

[0022] As can be seen from the above, the management node determines the permission information of the first standard metadata, which can ensure the reasonable distribution and strict management of the access permissions of different users, protect the metadata resources in the metadata management system from being accessed by unauthorized users, and ensure the security of the metadata management system.

[0023] In a second aspect, a metadata management device is provided. In the embodiments of the present application, the functional modules of the metadata management device can be divided according to the method provided in the first aspect above. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. Exemplarily, in the embodiments of the present application, the metadata management device can be divided into an acquisition module and a determination module according to functions. The descriptions of the possible technical solutions and beneficial effects executed by each of the above-divided functional modules can refer to the technical solutions provided in the first aspect above or its corresponding possible implementation manners, and will not be elaborated here.

[0024] In a third aspect, embodiments of the present application provide a computing device, which includes a processor and a memory. Computer instructions are stored in the memory, and the computer instructions are loaded and executed by the processor to enable the computing device to implement the metadata management method as described in the first aspect above.

[0025] In a fourth aspect, embodiments of the present application provide a terminal device, which includes a display screen and an input device. The display screen is used to display a first interface, and the first interface includes a relationship network and / or a query result; the input device is used to input a query request.

[0026] In a fifth aspect, embodiments of the present application provide a metadata management system, which includes a management device, a data source node device, and a terminal device; the management device is used to execute the metadata management method as described in the first aspect above; the data source node is used to produce and / or store data; the terminal device is used to display the first interface.

[0027] In a sixth aspect, embodiments of the present application provide a computer-readable storage medium, in which at least one computer program is stored, and the computer program is loaded and executed by a processor to implement the metadata management method as described in the first aspect above.

[0028] In a seventh aspect, an embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of the server reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the server executes the component deployment method of the application provided in various alternative implementations of the first aspect above.

[0029] For the specific descriptions of the second aspect to the seventh aspect and their various implementations in the embodiments of the present application, reference may be made to the detailed descriptions in the first aspect and its various implementations; and, for the beneficial effects of the second aspect to the seventh aspect and their various implementations, reference may be made to the analysis of the beneficial effects in the first aspect and its various implementations, which will not be elaborated here.

[0030] These aspects or other aspects of the embodiments of the present application will be more clearly understood in the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 FIG. 1 shows a schematic architecture diagram of a metadata management system 100 provided by an embodiment of the present application;

[0032] Figure 2 FIG. 2 shows a schematic hardware structure diagram of a computing device 200 provided by an embodiment of the present application;

[0033] Figure 3 FIG. 3 shows a schematic diagram of a metadata management method process provided by an embodiment of the present application;

[0034] Figure 4 FIG. 4 shows a schematic diagram of a first interface 300 provided by an embodiment of the present application;

[0035] Figure 5 FIG. 5 shows a schematic diagram of a second interface 310 provided by an embodiment of the present application;

[0036] Figure 6 FIG. 6 shows a schematic diagram of a data preprocessing method process provided by an embodiment of the present application;

[0037] Figure 7 FIG. 7 shows a schematic software architecture diagram of a metadata management system 100 provided by an embodiment of the present application;

[0038] Figure 8 FIG. 8 shows a schematic diagram of the structure of a metadata management device 400 provided by an embodiment of the present application;

[0039] Figure 9 FIG. 9 shows a schematic diagram of a computing device cluster provided by an embodiment of the present application;

[0040] Figure 10 The figure shows a schematic diagram of a possible implementation of a network connection provided by an embodiment of the present application. Detailed implementation manners

[0041] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. Among them, in the description of the present application, unless otherwise specified, " / " means that the objects associated before and after are in an "or" relationship. For example, A / B may represent A or B; "and / or" in the present application is only a description of the association relationship of the associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. These three situations, where A and B may be singular or plural. Also, in the description of the present application, unless otherwise specified, "a plurality of" means two or more than two. "At least one (item)" or its similar expression below refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c may represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c may be single or multiple. In addition, in order to clearly describe the technical solutions of the embodiments of the present application, in the embodiments of the present application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and roles.

[0042] Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and "first", "second", etc. do not necessarily mean different. At the same time, in some embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific way for easy understanding.

[0043] In addition, the device architecture and business scenarios described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art know that with the evolution of the device architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0044] First, an exemplary introduction to the application scenarios of the embodiments of the present application will be given.

[0045] Metadata is information that describes the characteristics of data, including descriptive information about the content, format, structure, source, or usage of the data. Metadata plays a crucial role in numerous fields. Due to the emergence of a vast amount of cross-system data, metadata management has become increasingly complex. Especially in a distributed environment, there are significant differences in the data standards, formats, or structures of different systems, further increasing the difficulty of metadata management.

[0046] Currently, data management systems and data governance platforms with multiple data sources that are widely used all require metadata management functions. Additionally, the needs and values of metadata management are also reflected in aspects such as data understanding and discovery, data governance and improvement, data governance and sharing, and data lifecycle management.

[0047] The embodiment of the present application provides a metadata management method, which is applied to a metadata management system. The metadata system includes a management node and multiple data source nodes, and the semantics and / or formats of the metadata corresponding to the multiple data source nodes are different; this method is executed by the management node and includes: obtaining first standard metadata; the first standard metadata is in a standard format recognizable by the metadata management system; the first standard metadata is determined based on the metadata corresponding to the first data source node; the first data source node is any one of the multiple data source nodes; determining a relationship network according to the first standard metadata; wherein, the relationship network is used to indicate the association relationship between the first standard metadata and other standard metadata under different dimensions; the association relationship represents the degree of tight connection between each standard metadata; the other standard metadata is any number of standard metadata obtained from the multiple data source nodes except the first data source node.

[0048] In the above method, the management node can obtain the first standard metadata from the first data source node, and determine the relationship network of different dimensions between the first standard metadata and other standard metadata according to the obtained first standard metadata, realizing the dynamic integration and update of the association relationship between each standard metadata, and meeting the dynamic management requirements of complex metadata in a distributed environment. Thus, multi-source and cross-system metadata management is achieved, improving the efficiency of metadata management.

[0049] Secondly, an exemplary introduction to the system architecture of the embodiment of the present application is provided.

[0050] Figure 1 Shows a schematic architecture diagram of a metadata management system 100 provided by an embodiment of the present application. As Figure 1 shown, the metadata management system 100 includes a management node 110 and multiple data source nodes 120; the management node 110 and the multiple data source nodes 120 can communicate with each other through a network.

[0051] Among them, the management node 110 can be a computing device. For example, the management node 110 can be a standard general server, specifically, it can be a blade server, a high-density server, a rack server, or a high-performance server, etc.

[0052] The management node 110 is used to obtain the first standard metadata, and determine a relationship network according to the first standard metadata. Among them, the first standard metadata is in a standard format recognizable by the metadata management system. The first standard metadata is determined according to the metadata corresponding to the first data source node; the first data source node is any one of multiple data source nodes. The relationship network is used to indicate the association relationship between the first standard metadata and other standard metadata based on different dimensions; the association relationship represents the tightness of connection between each standard metadata; the other standard metadata is any standard metadata obtained from multiple data source nodes other than the first data source node.

[0053] The semantics and / or formats of the metadata corresponding to each data source node 120 are different. That is to say, the data source node 120 is used to store and / or generate different types of data. The data source node 120 can specifically be an industrial terminal device or software system for storing data, or an industrial terminal device or software system for generating data. For example, the industrial terminal device includes a foldable electronic device, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an Ultra-Mobile Personal Computer (UMPC), a netbook, a cellular phone, a Personal Digital Assistant (PDA), an Augmented Reality (AR) device, a Virtual Reality (VR) device, an Artificial Intelligence (AI) device, a wearable device, a vehicle-mounted device, a sensor device, and other electronic devices with communication capabilities. The software system includes software systems for data management or enterprise management such as an enterprise resource planning (ERP) system, a customer relationship management (CRM) system, and a supply chain management (SCM) system. The specific type of the data source node 120 in the embodiments of the present application is not limited.

[0054] Optionally, the metadata management system 100 further includes a terminal device, which can communicate with the management node 110 in the metadata management system 100 through a network. The terminal device is used to display the relationship network and / or query results; the query results include target standard metadata determined by the metadata management system according to a query request sent by the terminal device.

[0055] Figure 2 FIG. shows a schematic hardware structure diagram of a computing device 200 provided by an embodiment of the present application. As Figure 2 shown, the computing device 200 includes a processor 210, a memory 220, a communication interface 230, and a bus 240. Among them, the processor 210, the memory 220, and the communication interface 230 communicate through the bus 240. It should be understood that the present application does not limit the number of processors and memories in the computing device 200. The processor 210 may specifically be a central processing unit (CPU).

[0056] The processor 210 is used to obtain first standard metadata, and the first standard metadata is a standard format recognized by the metadata management system. Among them, the first standard metadata is a standard format recognizable by the metadata management system; the first standard metadata is determined according to the metadata corresponding to the first data source node; the first data source node is any one of multiple data source nodes.

[0057] The processor 210 is used to determine a relationship network according to the first standard metadata; wherein, the relationship network is used to indicate the association relationship between the first standard metadata and other standard metadata under different dimensions; the association relationship represents the tightness of connection between each standard metadata; the other standard metadata is any number of standard metadata obtained from multiple data source nodes other than the first data source node.

[0058] The memory 220 stores executable program code, and the processor 210 executes the executable program code to implement the above metadata management method.

[0059] The communication interface 230 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 200 and other devices or communication networks.

[0060] It should be noted that the system architecture and application scenarios described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0061] For ease of understanding, the following provides an exemplary introduction to the metadata management method provided in this application with reference to the accompanying drawings. This metadata management method is executed by Figure 1 the management node 110 shown.

[0062] Figure 3 FIG. shows a schematic diagram of a metadata management method flow provided in an embodiment of this application. This metadata management method is executed by a management node and includes the following steps:

[0063] S101, the management node obtains first standard metadata.

[0064] In the embodiment of this application, the first standard metadata is in a standard format recognizable by the metadata management system; the first standard metadata is determined according to the metadata corresponding to the first data source node; the first data source node is any one of multiple data source nodes.

[0065] In a possible implementation method, the management node obtains the metadata corresponding to the first data source node; the management node determines the first standard metadata according to the metadata corresponding to the first data source node and the mapping rule;

[0066] wherein, the metadata corresponding to the first data source node is obtained by extracting the data obtained from the first data source node; the mapping rule is used to indicate the mapping relationship between the fields of the metadata and the fields of the standard metadata.

[0067] Exemplarily, first, relevant technical personnel construct a mapping rule library and a standard template for defining the conversion method of different metadata to standard metadata. Among them, the standard template includes each field in the standard metadata. This mapping rule library supports manual configuration and dynamic update. The static configuration of the mapping rule library can pre-define the mapping rule and the standard template by relevant technical personnel. The management node can optimize the mapping rule in real time by analyzing and learning the format, structure, and mutual relationship of the metadata through a machine learning model.

[0068] For example, the mapping rule includes that the "customer number" field or "Customer_ID" is mapped to the "customer ID" field in the standard template. If the first original metadata includes the "customer number" field; the management node determines that the first standard metadata includes the "customer ID" field according to the mapping rule and the first original metadata.

[0069] As can be seen from the above, since the data in each data source node may be data of different types, structures, or semantics, the semantics and / or format of the metadata corresponding to the first data source extracted by the management node may be inconsistent. The management node converts the metadata corresponding to the first data source into the first standard metadata in a unified format through the mapping rule, so that the metadata management system can recognize the metadata and perform further management.

[0070] In a possible implementation manner, the management node determines a target mapping rule based on predicting the type of metadata corresponding to the first data source node by a machine learning model. Among them, the target mapping rule is applicable to the type of metadata corresponding to the first data source node.

[0071] Exemplarily, when the management node obtains the data in the first data source node, the management node predicts the type of metadata corresponding to the first data source node through a machine learning model, and automatically selects the most suitable target mapping rule. Among them, the type of metadata corresponding to the first data source node includes descriptive metadata, shared metadata or dynamic metadata.

[0072] For example, taking the case where the machine learning model predicts that the metadata corresponding to the first data source node is business metadata describing the commodity classification rule as an example, the target mapping rule may include that the "raw material" field and the "material" field have a mapping relationship with the "material" field in the standard template.

[0073] As can be seen from the above, by predicting the type of metadata corresponding to the first data source node through a machine learning model to determine the target mapping rule, it is not necessary for the management node to determine the type of metadata corresponding to the first data source node after extracting the metadata corresponding to the first data source node; then the management node determines the target mapping rule according to the type of metadata corresponding to the first data source node. Therefore, not only the process of obtaining the first standard metadata is shortened, but also the accuracy of obtaining the first standard data can be improved.

[0074] In a possible implementation manner, the management node generates a first mapping log; the management node determines that there is data inconsistency or mapping anomaly between the metadata corresponding to the first data source node and the first standard metadata according to the first mapping log.

[0075] Among them, the first mapping log is used to record the mapping process from the metadata corresponding to the first data source node to the first standard metadata.

[0076] Exemplarily, in the process of the management node determining the first standard metadata according to the metadata corresponding to the first data source node and the mapping rule, by recording the first mapping log, the mapping process is monitored and traced. The management node can track the mapping process of the metadata corresponding to the first data source node, identify data inconsistency or mapping anomaly, and generate corresponding warning notifications, so as to ensure the consistency and standardization of multi-source metadata.

[0077] For example, the first mapping log may include the fields before mapping and the fields after mapping. The first mapping log may include the fields "customer number" and "customer ID". Based on the first mapping log and the mapping rules, the management node can determine whether the current field mapping complies with the mapping rules, so as to monitor and trace the mapping process. If the field before mapping in the first mapping log is the "role number" field and there is no mapping from the "role number" field to the "customer ID" field in the mapping rules, the management node determines that the mapping process of this field is abnormal.

[0078] In a possible implementation manner, when the first data source node includes structured data, the management node extracts the metadata corresponding to the first data source node based on the database system of the first data source node;

[0079] When the first data source node includes unstructured data, the management node extracts the metadata corresponding to the first data source node based on natural language processing (NLP) models and computer vision (CV) models.

[0080] Among them, structured data refers to data with a fixed format and organizational form that can be neatly arranged in a database table. For example, an employee information table. Unstructured data has no predefined format and is difficult to process with traditional databases. For example, images, audio and video, text, etc.

[0081] For example, when the first data source node includes structured data, the data of the first data source node can be stored in a relational database (MySQL) management system, and the database defines the table structure, field names, and quantity types. The management node can directly extract the metadata corresponding to the first data source node through views or data dictionaries.

[0082] Optionally, the management node is integrated with the database through a metadata governance tool (apache atlas) to automatically obtain the metadata corresponding to the data source node.

[0083] When the first data source node includes structured data, the management node extracts keywords, categories, timestamps, etc. based on natural language processing (NLP) models and computer vision (CV) models, and converts this information into a structured data format, so as to extract the metadata corresponding to the first data source node.

[0084] S102. The management node determines the relationship network according to the first standard metadata.

[0085] Among them, the relationship network is used to indicate the association relationship between the first standard metadata and other standard metadata based on different dimensions; the association relationship represents the tightness of the connection between each standard metadata; the other standard metadata is any standard metadata obtained from multiple data source nodes other than the first data source node.

[0086] In the embodiment of the present application, the management node obtains each standard metadata in multiple data source nodes, and the management node creates a relationship network between each standard metadata based on different dimensions.

[0087] In a possible implementation manner, if the relationship network includes a first relationship network, the first relationship network is used to indicate the association relationship between the first standard metadata and other standard metadata based on the first dimension; the first dimension is any one of different dimensions; the first dimension includes a data classification dimension, a content dimension, a time dimension, or a source dimension; determining the relationship network according to the first standard metadata includes: determining the association relationship between the first standard metadata and other standard metadata based on the first dimension based on a graph analysis algorithm model; determining the first relationship network according to the association relationship between the first standard metadata and other standard metadata based on the first dimension.

[0088] Among them, the graph analysis algorithm is used to process graph data composed of vertices and edges. The first dimension includes a data classification dimension, a content dimension, a time dimension, and a source dimension.

[0089] Exemplarily, the management node determines the association relationship between the first standard metadata and other standard metadata based on the first dimension, and constructs the first relationship network. Specifically, the management node can determine the association relationship between each standard metadata based on the relevance of the standard metadata content. The management node can use a graph data structure to store the association relationship between the standard metadata. Among them, the graph data structure is a complex non-linear data structure used to represent objects and the relationships between objects, and is composed of a vertex set and an edge set. The vertices of the graph data structure represent the standard metadata, the edges represent the association relationship between the standard metadata, and the weight of the edge represents the association strength.

[0090] For example, the management node determines a first relationship network based on the content dimension between the first standard metadata and other standard metadata according to the first standard metadata. Specifically, the management node determines the similarity between the content of the first standard metadata and each other standard metadata through keyword similarity or vectorization techniques (such as the word embedding tool Word2Vec, bidirectional encoder representations from transformers (BERT)), and determines the association relationship based on the data content between the first standard metadata and each other standard metadata. The management node determines the first relationship network through a graph analysis algorithm model.

[0091] For another example, the management node can update the first relationship network according to the first standard metadata; the method for the management node to establish a connection between the first standard metadata and the metadata in the existing relationship network is the same as the above method, and will not be elaborated here.

[0092] In a possible implementation manner, the metadata management system further includes a terminal device, which is used to display the relationship network and / or query results; the query results include the target standard metadata determined by the metadata management system according to the query request sent by the receiving terminal; the target standard metadata is used to indicate any standard metadata in the metadata management system.

[0093] Among them, the query request is used to indicate obtaining at least one standard metadata in the relationship network.

[0094] Exemplarily, the management node receives the query request sent by the terminal device, the management node parses the query request, and determines the target standard metadata. The management node sends the target standard metadata and / or the target relationship network to the terminal device. The management node uses a graph database as the storage structure and displays the target relationship network through a visualization tool.

[0095] Among them, the target relationship network is used to indicate the association relationship between the target standard metadata.

[0096] Optionally, the management node uses multi-dimensional indexing and keyword matching techniques to support multi-condition queries of the terminal device. At the same time, the management node adopts a distributed query architecture and a caching mechanism to optimize the query speed to improve the user experience.

[0097] For example, taking the metadata management system including a book management system as an example, Figure 4 shows a schematic diagram of a first interface 300 provided by an embodiment of the present application. As Figure 4As shown, the terminal device displays a first interface 300 through a display screen. The first interface 300 can be an information query interface of a book management system. The first interface 300 includes a search box and a query button. The preset prompt of the search box is "Please enter the query condition", and the preset prompt of the search button is "Query". The user enters a query condition through the search box and clicks the query button to generate a query request. The terminal device obtains the query request input by the user through the first interface 300 and sends the query request to the management node in the metadata management system to query the standard metadata related to the basic information of Book 01. The management node determines the target standard metadata according to the identification information of Book 01 in the query request. The management node determines the target relationship network according to the target standard metadata and sends the target relationship network to the terminal device, and the terminal device displays the target relationship network through the display screen. Figure 5 The figure shows a schematic diagram of a second interface 310 provided by an embodiment of the present application. As Figure 5 shown, the second interface 310 displays the target relationship network. The target relationship network includes the book name, author, publisher, borrower, borrowing date, and creation time. Each vertex of the target relationship network represents each target standard metadata, and the edge represents the association relationship between the standard metadata. The weight of the edge is presented by the length of the edge. Figure 5 It can be understood from this that the longer the side length, the weaker the association relationship, and the shorter the side length, the stronger the association relationship. As can be seen from the figure, the book name is strongly associated with the author, publisher, and borrower; the borrower is strongly associated with the borrowing date; the author is strongly associated with the creation time; the book name is weakly associated with the borrowing date and creation time. Figure 5 It can clearly show the relationship between the standard metadata related to the basic information of Book 01.

[0098] It should be noted that the weight of the edge can also be represented by the depth of the edge color and the thickness of the line. Figure 5 The presentation of the weight of the edge by the length of the edge in this is only an example and does not constitute a specific limitation.

[0099] In a possible implementation manner, the management node determines the permission information of the first standard metadata according to the type of the first standard metadata. Among them, the permission information is used to indicate the users who have the right to access the first standard metadata.

[0100] Exemplarily, the management node adopts a role-based access control (RBAC) method to dynamically adjust user permissions. The management node determines the permission information of the first standard metadata according to the permission mapping relationship and the type of the first standard metadata. Among them, the permission mapping relationship is used to indicate the mapping relationship between the role information and the standard metadata type in the metadata management system.

[0101] For example, the permission mapping relationship includes that there is a mapping relationship between technical metadata and maintenance roles, a mapping relationship between business metadata and sales roles, and a mapping relationship between operational metadata and administrator roles. If the management node determines that the first standard metadata is operational metadata, it can determine that the permission information of the first standard metadata includes the identification information of the administrator role.

[0102] As can be seen from the above, the management node determines the permission information of the first standard metadata, which can ensure the reasonable allocation and strict management of the access permissions of different users, protect the metadata resources in the metadata management system from being accessed by unauthorized users, and ensure the security of the metadata management system.

[0103] In a possible implementation manner, the management node collects and saves operation logs; among them, the operation logs include the user operation records and standard metadata access records of the receiving terminal in the metadata management system.

[0104] Exemplarily, the management node adopts an immutable operation log recording method (for example, blockchain technology or data signature technology), combines the time stamp and the identification information of the user, and collects the operation and access data of the user in real time. And stores it in the operation log file, generates an audit report regularly, provides a comprehensive operation record, and provides security guarantee for the metadata management system.

[0105] The above text has described the metadata management method provided by the present application as a whole. Before the management node obtains the first standard metadata in step S101, the management node also needs to preprocess the data in the data source node and extract the data in the preprocessed data source node to obtain the metadata corresponding to the data source node. Among them, the data source node includes structured data, semi-structured data or unstructured data. Specifically, it may include tables, numerical values, texts, pictures, audios, videos, etc. Next, in combination with Figure 7 The steps of the above data preprocessing will be described in detail.

[0106] In a possible implementation manner, the management node preprocesses the data in the data source node, where the preprocessing includes removing duplicate data, filling in missing values or consistency checking.

[0107] For example, Figure 6 shows a schematic flow chart of a data preprocessing method provided by an embodiment of the present application. As Figure 6 shown, it includes the following steps:

[0108] S201, the management node obtains the data of multiple data source nodes.

[0109] Specifically, the management node collects and obtains data from multiple data source nodes, including sensors, device interfaces, user inputs, etc. Among them, the management node supports multiple data access protocols, such as the hyper text transfer protocol (HTTP), the message queuing telemetry transport (MQTT), the WebSocket, a full-duplex communication protocol based on TCP, etc., which can ensure that the management node connects to different types of terminal devices and data sources. The management node realizes the unified access to different data formats through the adapter mode, ensuring the flexibility and compatibility of data collection.

[0110] Optionally, through the data caching mechanism, the management node can ensure the orderly reception of data and avoid packet loss or data delay in the case of high-concurrency data collection. The management node can also monitor duplicate or incorrect raw data in real time. For example, it can exclude duplicate raw data through unique identifiers or check codes, improving the quality of the collected raw data.

[0111] S202. The management node screens and filters the data from the data source nodes.

[0112] Specifically, the management node screens the data from the data source nodes according to predefined rules (such as value range, data format, timestamp, etc.), removing obviously unqualified data. For example, the data generated by a temperature sensor is numerical data. If the data obtained by the management node from the temperature sensor is in text format, then the current text-format data is obviously unqualified. In addition, the management node can also use methods such as moving average and low-pass filters to process the noise of the data continuously collected from the data source nodes, ensuring the smoothness and reliability of the raw data.

[0113] Optionally, for the scenario of high-frequency raw data collection, the management node adopts batch processing technology to improve the screening efficiency, thereby reducing the system processing delay.

[0114] S203. The management node preprocesses the data from the data source nodes.

[0115] Among them, the preprocessing includes removing duplicate data, filling in missing values, or consistency checking.

[0116] The metadata management method will be described by way of example in combination with specific usage scenarios below.

[0117] Large enterprise A faces problems of multiple data sources, data formats, and data interoperability among multiple systems in its daily operations. The enterprise decides to deploy the metadata management system 100 provided by the embodiments of this application. The metadata management system 100 includes a management node 110, and the data source nodes 120 are specifically an ERP system, an SCM system, and a CRM system. The management node 110 obtains each original metadata in the ERP system, the SCM system, and the CRM system through an application programming interface (API).

[0118] Figure 7 FIG. shows a schematic software architecture diagram of a metadata management system 100 provided by an embodiment of this application. As Figure 7 shown, the metadata management system 100 includes a data collection layer, a mapping layer, a relationship network layer, and an interaction layer.

[0119] Among them, the data collection layer is used to collect metadata corresponding to multiple data source nodes; the data source nodes may include an ERP system, an SCM system, and a CRM system. Before obtaining the metadata corresponding to each data source node, the data collection layer preprocesses the data of each system. The data collection layer extracts the metadata of multiple data source nodes according to the preprocessed data.

[0120] The mapping layer is used to determine the first standard metadata according to the metadata corresponding to the first data source node and the mapping rule. For example, fields such as "Customer_ID" in the ERP system and "customer number" in the CRM system are marked as a unified "customer ID" field through standardization conversion. The specific method is the same as the method described in S101 above Figure 3 and will not be elaborated here.

[0121] The relationship network layer is used to determine the relationship network from the first standard metadata. The relationship network layer is also used to generate a relationship mapping between systems, analyze the association relationships of "customer information" in different dimensions in the ERP, CRM, and SCM systems, and generate an integrated data relationship diagram for information such as "customer orders" and "supply chain records" through the relationship network.

[0122] The scheduling layer is used to collect and save operation logs and mapping logs. The operation logs include user operation records and standard metadata access records of the receiving terminal in the metadata management system. The mapping logs are used to record the mapping process of metadata corresponding to multiple data sources to standard metadata; the mapping logs include the first mapping log. The scheduling layer is used to monitor and trace the mapping process according to the first mapping log. This enables the metadata management system to track the mapping process of metadata corresponding to each data source node to standard metadata, identify data inconsistencies or mapping anomalies, and generate corresponding warning notifications, thereby ensuring the consistency and standardization of multi-source metadata. The scheduling layer is also used to regularly generate audit reports based on the operation logs, providing comprehensive operation records and ensuring the security of the metadata management system.

[0123] The interaction layer is used to provide an interaction interface to view the associated data of customer objects in different business systems, supporting further data mining and insights. The interaction layer is also used for permission control to ensure that only authorized users can view or edit metadata content.

[0124] As can be seen from the above, the metadata management system 100 has completed cross-system metadata standardization mapping, enabling the metadata fields of different systems to communicate with each other and improving the efficiency of data integration. Users can also intuitively view the relationship network diagram of metadata in different dimensions, thereby understanding the source and flow of data and enhancing the transparency of data governance. Therefore, the metadata management system 100 not only realizes cross-system metadata management but also can dynamically adapt to changes in metadata relationships, thereby improving the metadata management efficiency of Enterprise A.

[0125] In summary, the embodiment of the present application provides a metadata management method. This method is applied to a metadata management system, which includes a management node and multiple data source nodes, and the semantics and / or formats of the metadata corresponding to the multiple data source nodes are different; this method is executed by the management node. The method includes: the management node obtains the first standard metadata; the management node updates the relationship network according to the first standard metadata; wherein, the first standard metadata is the standard format recognized by the metadata management system; the first standard metadata is determined according to the metadata corresponding to the first data source node; the first data source node is any one of the multiple data source nodes; the relationship network is used to indicate the association relationship between the first standard metadata and other standard metadata based on different dimensions; the association relationship represents the degree of tight connection between each standard metadata; the other standard metadata is any number of standard metadata obtained from multiple data source nodes other than the first data source node.

[0126] In the above method, the management node can obtain the first standard metadata from the first data source node. According to the obtained first standard metadata, the relationship network of different dimensions between the first standard metadata and other standard metadata is determined, realizing the dynamic integration and update of the association relationships between the standard metadata of different systems, and realizing the metadata management of multiple sources and across systems, thereby improving the metadata management efficiency.

[0127] The above mainly introduces the solution of the embodiment of the present application from the perspective of the method. It can be understood that, in order for the metadata management device to implement the above functions, it includes at least one of the corresponding hardware structures and software modules for executing each function. Those skilled in the art should easily realize that, combined with the units and algorithm steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0128] The embodiments of the present application can divide the functional units of the metadata management device according to the above method examples. For example, each functional unit can be divided corresponding to each function, or two or more functions can be integrated into one processing unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. It should be noted that the division of units in the embodiments of the present application is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0129] Exemplarily, Figure 8 shows a schematic diagram of the structure of a metadata management device 400 provided by an embodiment of the present application. As Figure 8 shown, the metadata management device 400 can be applied to Figure 1 the management node 110 shown, and the metadata management device 400 includes:

[0130] A collection module 410, configured to obtain the first standard metadata; the first standard metadata is the standard format recognized by the metadata management system; the first standard metadata is determined according to the metadata corresponding to the first data source node; the first data source node is any one of multiple data source nodes.

[0131] An update module 420 is configured to determine a relationship network according to first standard metadata, where the relationship network is used to indicate the association relationship between the first standard metadata and other standard metadata under different dimensions; the association relationship represents the tightness of the connection between each standard metadata; the other standard metadata is any standard metadata obtained from multiple data source nodes except the first data source node. In a possible implementation, the acquisition module 410 is further configured to obtain the metadata corresponding to the first data source node; the metadata corresponding to the first data source node is obtained by extracting the data obtained from the first data source node; according to the metadata corresponding to the first data source node and the mapping rule, determine the first standard metadata; the mapping rule is used to indicate the mapping relationship between the fields of the metadata and the fields of the standard metadata.

[0132] In a possible implementation, the acquisition module 410 is further configured to, when the first data source node includes structured data, extract the metadata corresponding to the first data source node based on the database system of the first data source node;

[0133] When the first data source node includes unstructured data, extract the metadata corresponding to the first data source node based on a natural language processing (NLP) model and a computer vision (CV) model.

[0134] In a possible implementation, the update module 420 is further configured to determine the association relationship between the first standard metadata and other standard metadata based on the first dimension according to a graph analysis algorithm model; determine the first relationship network according to the association relationship between the first standard metadata and other standard metadata based on the first dimension.

[0135] In a possible implementation, the metadata management device 400 further includes a judgment module, configured to determine a target mapping rule based on predicting the type of the metadata corresponding to the first data source node by a machine learning model; the target mapping rule is applicable to the type of the metadata corresponding to the first data source node.

[0136] In a possible implementation, the judgment module is further configured to generate a first mapping log; the first mapping log is used to record the mapping process from the metadata corresponding to the first data source node to the first standard metadata; determine the data inconsistency and / or mapping abnormality between the metadata corresponding to the first data source node and the first standard metadata according to the first mapping log.

[0137] In a possible implementation, the judgment module is further configured to determine the permission information of the first standard metadata according to the type of the first standard metadata; the permission information is used to indicate the users who are authorized to access the first standard metadata.

[0138] In a possible implementation manner, the metadata management device 400 further includes a preprocessing module for preprocessing the data in the first data source node; the preprocessing includes removing duplicate data, filling in missing values, or performing consistency checks.

[0139] Optionally, the component deployment device 400 of the foregoing application can be applied to Figure 2 the computing device 200 shown in the figure.

[0140] For the specific description of the foregoing optional manner, reference can be made to the foregoing method embodiments, which will not be elaborated herein. In addition, the explanations and descriptions of the beneficial effects of any of the foregoing metadata management devices 400 can all refer to the Figure 3 corresponding method embodiments and will not be elaborated.

[0141] Among them, both the acquisition module 410 and the update module 420 can be implemented by software or by hardware. Exemplarily, next, taking the acquisition module 410 as an example, the implementation manner of the acquisition module 410 will be introduced. Similarly, the implementation manner of the update module 420 can refer to the implementation manner of the acquisition module 410.

[0142] As an example of a software functional unit, the acquisition module 410 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the foregoing computing instance may be one or more. For example, the acquisition module 410 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running this code may be distributed in the same region, or may be distributed in different regions. Further, the multiple hosts / virtual machines / containers for running this code may be distributed in the same availability zone (AZ), or may be distributed in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. Among them, generally one region may include multiple AZs.

[0143] Similarly, the multiple hosts / virtual machines / containers for running this code may be distributed in the same virtual private cloud (VPC), or may be distributed in multiple VPCs. Among them, generally one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is realized through the communication gateway.

[0144] As an example of a hardware functional unit, the acquisition module 410 may include at least one computing device, such as a server. Alternatively, the acquisition module 410 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0145] The multiple computing devices included in the acquisition module 410 may be distributed in the same region or in different regions. The multiple computing devices included in the acquisition module 410 may be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the acquisition module 410 may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, the multiple computing devices may be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0146] It should be noted that in other embodiments, the acquisition module 410 may be used to execute any step in the metadata management method, and the update module 420 may be used to execute any step in the metadata management method. The steps implemented by the acquisition module 410 and the update module 420 can be specified as needed, and all functions of the metadata management apparatus 400 are realized by implementing different steps in the metadata management method through the acquisition module 410 and the update module 420 respectively.

[0147] As an example, in combination with Figure 2 , some or all of the functions of the acquisition module 410 and the update module 420 of the metadata management apparatus 400 may be implemented by the computing device 200 in Figure 2 .

[0148] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. For example, it is a server, specifically a central server, an edge server, or a local server in a local data center. In some embodiments, the server may also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0149] Figure 9The figure shows a schematic diagram of a computing device cluster provided by an embodiment of the present application. As Figure 9 shown, the computing device cluster includes at least one computing device 200. Instructions for executing a data management method that are the same may be stored in the memory 220 of one or more of the computing devices 200 in the computing device cluster.

[0150] In some possible implementation manners, partial instructions for executing a metadata management method may also be separately stored in the memory 220 of one or more of the computing devices 200 in the computing device cluster. In other words, a combination of one or more computing devices 200 may jointly execute the instructions for executing the metadata management method.

[0151] It should be noted that different computing devices 200 in the computing device cluster may store different instructions, which are respectively used to implement partial functions of the metadata management device 400. That is, the instructions stored in the memory 220 of different computing devices 200 may implement the functions of one or more of the acquisition module 410 and the update module 420.

[0152] In some possible implementation manners, one or more computing devices in the computing device cluster may be connected through a network. Among them, the network may be a wide area network or a local area network, etc. Figure 10 The figure shows a schematic diagram of a possible implementation manner of a network connection provided by an embodiment of the present application. As Figure 10 shown, a first computing device 200A and a second computing device 200B are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation manner, instructions for executing the function of the acquisition module 410 are stored in the memory 220 of the first computing device 200A. At the same time, instructions for executing the function of the update module 420 are stored in the memory 220 of the second computing device 200B.

[0153] Figure 10 The connection manner between the computing device clusters shown may be considered in view of the needs of the data management method provided by the present application (such as storing a large amount of data), so it is considered to hand over the implemented function to the second computing device 220B for execution.

[0154] It should be understood that Figure 10 the function of the first computing device 200A shown may also be completed by multiple computing devices 200. Similarly, the function of the second computing device 200B may also be completed by multiple computing devices 200.

[0155] The present application embodiment also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster may be similarly referred to Figure 9 and Figure 10The connection method of the shown computing device cluster. Different from this, in the memory 220 of one or more computing devices 200 in the computing device cluster, there may be stored the same instructions for executing the metadata management method.

[0156] In some possible implementation manners, in the memory 220 of one or more computing devices 200 in the computing device cluster, there may also be separately stored partial instructions for executing the metadata management method. In other words, the combination of one or more computing devices 200 can jointly execute the instructions for executing the metadata management method.

[0157] It should be noted that the memories 220 in different computing devices 200 in the computing device cluster can store different instructions for executing partial functions of the metadata management apparatus 400. That is to say, the instructions stored in the memories 220 in different computing devices 200 can implement the functions of the metadata management apparatus 400.

[0158] The embodiment of the present application also provides a computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are run on a computer, the computer is caused to execute the operations of any one implementation manner and its various feasible implementation manners corresponding to the metadata management method.

[0159] The embodiment of the present application also provides a computer program product including instructions. When the computer program product is run on a computer, the computer is caused to execute the operations of any one implementation manner and its various feasible implementation manners corresponding to the data limit management method.

[0160] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraint conditions of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0161] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0162] The embodiment of the present application also provides a chip system, including: a processor, the processor is coupled to a memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the chip system implements the method in any one of the foregoing method embodiments.

[0163] Optionally, there may be one or more processors in the chip system. The processor may be implemented by hardware or by software. When implemented by hardware, the processor may be a logic circuit, an integrated circuit, etc. When implemented by software, the processor may be a general-purpose processor that implements its functions by reading software codes stored in a memory.

[0164] Optionally, there may also be one or more memories in the chip system. The memory may be integrated with the processor or may be separately provided from the processor. The embodiments of the present application do not limit this. Exemplarily, the memory may be a non-transitory processor, such as a read-only memory (ROM). It may be integrated with the processor on the same chip or may be separately provided on different chips. The embodiments of the present application do not specifically limit the type of the memory and the setting manner of the memory and the processor.

[0165] Exemplarily, the chip system may be a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on chip (SoC), a central processing unit (CPU), a network processor (NP), a digital signal processing circuit (DSP), a micro controller unit (MCU), a programmable logic device (PLD), or other integrated chips.

[0166] The electronic device, computer storage medium, or computer program product provided in the present application is all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.

[0167] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional module is used as an example. In actual applications, the above functions may be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0168] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0169] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, which may be located in one place or distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0170] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0171] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that makes a contribution, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks or optical discs and other various media that can store program codes.

[0172] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A metadata management method, characterized in that, The method is applied to a metadata management system, which includes a management node and multiple data source nodes, and the semantics and / or formats of the metadata corresponding to the multiple data source nodes are different; The method is executed by the management node and includes: Obtain first standard metadata; the first standard metadata is in a standard format recognizable by the metadata management system; the first standard metadata is determined according to the metadata corresponding to the first data source node; the first data source node is any one of the multiple data source nodes; Determine a relationship network according to the first standard metadata; Among them, the relationship network is used to indicate the association relationship between the first standard metadata and other standard metadata based on different dimensions; the association relationship represents the tightness of connection between each standard metadata; the other standard metadata is any number of standard metadata obtained from the multiple data source nodes except the first data source node.

2. The method according to claim 1, wherein The relationship network includes a first relationship network, which is used to indicate the association relationship between the first standard metadata and the other standard metadata based on the first dimension; the first dimension is any one of the different dimensions; the first dimension includes a data classification dimension, a content dimension, a time dimension or a source dimension; The determining the relationship network according to the first standard metadata includes: Based on a graph analysis algorithm model, determine the association relationship between the first standard metadata and the other standard metadata based on the first dimension; Determine the first relationship network according to the association relationship between the first standard metadata and the other standard metadata based on the first dimension.

3. The method according to claim 1 or 2, characterized in that, The metadata management system further includes a terminal device, which is used to display the relationship network and / or query results; the query results include target standard metadata determined by the metadata management system according to a query request sent by the terminal device; The target standard metadata is used to indicate any number of standard metadata in the metadata management system.

4. The method according to any one of claims 1 to 3, characterized in that, The obtaining the first standard metadata includes: Obtain the metadata corresponding to the first data source node; the metadata corresponding to the first data source node is obtained by extracting the data obtained from the first data source node; Determine the first standard metadata according to the metadata corresponding to the first data source node and a mapping rule; the mapping rule is used to indicate the mapping relationship between the fields of the metadata and the fields of the standard metadata.

5. The method according to claim 4, characterized in that, Before determining the first standard metadata according to the metadata corresponding to the first data source node and the mapping rule, the method further includes: Predict the type of the metadata corresponding to the first data source node based on a machine learning model, and determine a target mapping rule; the target mapping rule is applicable to the type of the metadata corresponding to the first data source node.

6. The method according to claim 4 or 5, characterized in that, The method further includes: Generate a first mapping log; the first mapping log is used to record the mapping process from the metadata corresponding to the first data source node to the first standard metadata. Determine data inconsistency and / or mapping anomalies between the metadata corresponding to the first data source node and the first standard metadata according to the first mapping log.

7. The method according to any one of claims 4 to 6, characterized in that, The obtaining of the metadata corresponding to the first data source node includes When the first data source node includes structured data, extracting the metadata corresponding to the first data source node based on the database system of the first data source node; When the first data source node includes unstructured data, extracting the metadata corresponding to the first data source node based on a natural language processing (NLP) model and a computer vision (CV) model.

8. The method according to any one of claims 4 to 7, characterized in that Before obtaining the metadata corresponding to the first data source node, the method further includes: Preprocessing the data in the first data source node; the preprocessing includes removing duplicate data, filling in missing values, or performing consistency checks.

9. The method according to any one of claims 1-8, characterized in that, The method further includes: Determine the permission information of the first standard metadata according to the type of the first standard metadata; the permission information is used to indicate the users who have the right to access the first standard metadata.

10. A computing device, characterized in that, The computing device includes: a processor and a memory for storing instructions executable by the processor; The processor is configured to execute the instructions, so that the computing device executes the metadata management method according to any one of claims 1-9.