Metadata tag generation method and device, equipment, medium and program product

By extracting metadata from multiple data sources and generating metadata tags using a pre-built metadata model and lineage resolution method, the problem of low accuracy and poor timeliness of metadata tags in existing technologies is solved, thus achieving efficient and accurate metadata management.

CN120929448APending Publication Date: 2025-11-11CHINA MOBILE FINANCIAL TECHNOLOGY CO LTD +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511057867.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies suffer from low accuracy and poor timeliness in metadata tagging, making it difficult to effectively manage and optimize data models.

Method used

Metadata is extracted from multiple data sources, standardized using a pre-built metadata model, and bloodline relationships are generated through various bloodline resolution methods. Metadata tags are generated using a pre-trained model to ensure the accuracy and timeliness of the tags.

Benefits of technology

It improves the accuracy and timeliness of metadata tagging and management, simplifies the complexity of data management, provides a unified data access layer, and enhances data lineage coverage and management efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929448A_ABST
    Figure CN120929448A_ABST
Patent Text Reader

Abstract

The invention provides a metadata tag generation method and device, equipment, a medium and a program product, and the method comprises the steps: extracting metadata from a plurality of data sources, and obtaining to-be-processed metadata; performing standardization processing on the to-be-processed metadata by using a pre-constructed metadata model to generate a first data model; the first data model is subjected to blood relationship analysis through multiple blood relationship analysis modes, the blood relationship of the first data model is obtained, and the blood relationship is used for indicating circulation information of data in the first data model; generating blood relationship data of the first data model according to the blood relationship of the first data model; and inputting the blood relationship data of the first data model into a pre-trained first model to obtain a metadata tag corresponding to the first data model output by the first model. By adopting the embodiment of the invention, the accuracy of generating the metadata tag is relatively high, and the generation efficiency is relatively high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data technology, and in particular to a method, apparatus, device, medium, and program product for generating metadata tags. Background Technology

[0002] With the rapid development of big data technology, the widespread adoption of cloud computing, and the continuous expansion of artificial intelligence application scenarios, the value of data is increasingly evident, leading to a surge in demand for data models from various organizations. The proliferation of open-source tools and frameworks has further simplified the creation of data models. This increased demand and improved efficiency in data model creation have jointly fueled the explosive growth of data models. To better leverage the value of data and enhance the efficiency of data model usage, organizations need to better manage their data models. Establishing metadata application tags for data models is fundamental to ensuring that models are manageable, understandable, assessable, and optimizable.

[0003] However, existing technologies that use data extraction, transformation, and loading (ETL) tools to collect or log capture data, along with manual maintenance of relevant metadata (model ownership, scope of use, definition, business application, value assessment, etc.) to create metadata tags suffer from low accuracy of metadata tags and poor timeliness in metadata tag management. Summary of the Invention

[0004] The purpose of this invention is to provide a method, apparatus, device, medium, and program product for generating metadata tags, in order to solve the problems of low accuracy of metadata tags and poor timeliness of metadata tag management.

[0005] To address the aforementioned technical problems, embodiments of the present invention provide a method for generating metadata tags, comprising:

[0006] Metadata is extracted from multiple data sources to obtain the metadata to be processed;

[0007] The metadata to be processed is standardized using a pre-built metadata model to generate a first data model;

[0008] The first data model is analyzed using multiple lineage analysis methods to obtain the lineage relationship of the first data model, wherein the lineage relationship is used to indicate the flow information of data in the first data model;

[0009] Generate bloodline data for the first data model based on the bloodline relationships in the first data model;

[0010] The blood relationship data of the first data model is input into the pre-trained first model to obtain the metadata tag corresponding to the first data model output by the first model.

[0011] Optionally, the step of standardizing the metadata to be processed using a pre-built metadata model to generate a first data model includes:

[0012] According to the pre-set mapping rules, the metadata to be processed is mapped to a pre-built metadata model for standardization processing to generate a first data model;

[0013] If a data definition conflict is detected in the metadata to be processed, the data definition of the metadata to be processed is processed according to a pre-set conflict resolution strategy, and the first data model is updated.

[0014] Optionally, the lineage relationship includes multiple nodes, edges, and data flow rules between the nodes. The nodes include a first node corresponding to the first data model and a second node corresponding to at least one second data model that has a data flow relationship with the first data model. The edges are used to connect the first node and the second node that have a data flow relationship and to indicate the direction of data flow between the connected first node and the second node. The edges also carry the source channel and lineage resolution method of the lineage relationship.

[0015] Optionally, the step of using multiple lineage analysis methods to perform lineage analysis on the first data model to obtain the lineage relationship of the first data model includes:

[0016] The first data model is subjected to kinship analysis using multiple kinship analysis methods to obtain kinship analysis results corresponding to the first data model, wherein the kinship analysis results include at least one candidate kinship relationship;

[0017] Based on the bloodline analysis method, source channel, and data processing frequency corresponding to the candidate bloodline relationship, the at least one candidate bloodline relationship is screened according to the first priority order among multiple bloodline analysis methods, the second priority order among multiple source channels, and the third priority order among data processing frequencies to determine the bloodline relationship of the first data model.

[0018] Optionally, the blood relationship may also include the date of formation of the blood relationship;

[0019] The method further includes:

[0020] Based on the date of bloodline formation, the bloodline relationships that meet preset conditions are pruned and reconstructed, wherein the preset conditions include the time interval between the date of bloodline formation and the current date exceeding a first duration.

[0021] Optionally, after performing kinship analysis on the first data model using multiple kinship analysis methods to obtain the kinship relationships of the first data model, the method further includes:

[0022] In the case where at least one of the second nodes is a temporary node, the first lineage relationship of the temporary node is obtained, wherein the first lineage relationship includes at least one third node that has a data flow relationship with the temporary node and the source channel of the first lineage relationship;

[0023] If the source of the blood relationship of the first node is the same as the source of the first blood relationship, when the temporary node disappears, the blood relationship of the first node is updated by connecting the first node and the third node through an edge.

[0024] Optionally, the bloodline relationship further includes at least one of the following: the processing task information of the bloodline relationship carried by the edge, the quality alarm information of the node, and the data usage frequency information of the leaf nodes in the node;

[0025] The method further includes at least one of the following:

[0026] The processing task information of the blood relationship is obtained through the application programming interface (API).

[0027] The quality result data corresponding to the first data model is obtained through the data quality control platform, and quality alarm information corresponding to each node is generated.

[0028] The frequency of data usage of the leaf nodes is obtained through the application-side operation monitoring interface.

[0029] Optionally, the generation of kinship data for the first data model based on the kinship relationship of the first data model includes:

[0030] Input the bloodline relationships from the first data model into the graph database to generate a bloodline database;

[0031] The identification information of the first data model is input into the lineage database to obtain the initial lineage relationship data corresponding to the first data model output by the lineage database. The initial lineage relationship data includes at least one of the following: statistical indicator information of the first data model, basic indicator information of downstream applications, and atomic metadata tags.

[0032] The initial bloodline data is standardized and normalized to obtain bloodline data.

[0033] Optionally, the step of inputting the identification information of the first data model into the kinship database to obtain the initial kinship data corresponding to the first data model output by the kinship database includes at least one of the following:

[0034] The identification information of the first data model is input into the bloodline bank, and the statistical index information of the first data model output by the bloodline bank is obtained by using the application programming interface (API) call method.

[0035] The identification information of the first data model is input into the lineage database, and the basic indicator information and atomic metadata tags of the downstream application corresponding to the first data model output by the lineage database are obtained using the depth-first search (DFS) algorithm.

[0036] Optionally, the step of inputting the kinship data of the first data model into the pre-trained first model to obtain the metadata tag corresponding to the first data model output by the first model includes:

[0037] The blood relationship data of the first data model is input into the decision tree model in the pre-trained first model to obtain the target blood relationship data output by the decision tree model, wherein the decision tree model is used to filter the blood relationship data according to the importance of the data features;

[0038] The target bloodline data of the first data model is input into the random forest model in the pre-trained first model to obtain the metadata tags corresponding to the first data model output by the random forest model.

[0039] Optionally, after obtaining the metadata tag corresponding to the first data model output by the first model, the method further includes:

[0040] When an abnormal output of the first model is detected, a corresponding alarm message is generated and a corresponding recovery strategy is executed, wherein the recovery strategy is used to resolve the abnormal output of the first model;

[0041] Upon receiving feedback from the user regarding the metadata tags output by the first model, the first model is optimized based on the feedback.

[0042] This invention also provides a metadata tag generation apparatus, comprising:

[0043] The first extraction module is used to extract metadata from multiple data sources to obtain metadata to be processed.

[0044] The first processing module is used to standardize the metadata to be processed using a pre-built metadata model to generate a first data model.

[0045] The first parsing module is used to perform lineage analysis on the first data model using multiple lineage analysis methods to obtain the lineage relationship of the first data model, wherein the lineage relationship is used to indicate the flow information of data in the first data model;

[0046] The first generation module is used to generate blood relationship data of the first data model based on the blood relationship of the first data model;

[0047] The second processing module is used to input the blood relationship data of the first data model into the pre-trained first model to obtain the metadata tag corresponding to the first data model output by the first model.

[0048] This invention also provides a network device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the metadata tag generation method as described in any of the preceding embodiments.

[0049] This invention also provides a readable storage medium, comprising: a program stored on the readable storage medium, wherein when the program is executed by a processor, it implements the steps of the metadata tag generation method as described in any of the preceding claims.

[0050] This invention also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the metadata tag generation method as described in any of the preceding claims.

[0051] At least one of the above technical solutions of the present invention has the following beneficial effects:

[0052] In the above scheme, firstly, metadata to be processed is extracted from multiple data sources. Then, a pre-built metadata model is used to standardize the metadata to generate a first data model, ensuring that the structure and data definition of metadata from different data sources are the same, which facilitates subsequent data processing. Secondly, multiple lineage resolution methods are used to resolve the lineage of the first data model to obtain the lineage relationship of the first data model. Based on the lineage relationship of the first data model, lineage relationship data of the first data model is generated. The lineage relationships of multiple lineage resolution methods are integrated to improve the lineage coverage and accuracy of the data, thereby obtaining the lineage relationship data of the first data model. Finally, the lineage relationship data is input into a pre-trained first model to obtain the metadata tags corresponding to the first data model output by the first model, improving the accuracy of the metadata tags. The tags are also automatically generated by the first model to improve the timeliness of the metadata tags. Attached Figure Description

[0053] Figure 1 This is a flowchart illustrating the method for generating metadata tags according to an embodiment of the present invention;

[0054] Figure 2 This is a schematic diagram of the structure of the metadata tag generation device according to an embodiment of the present invention. Detailed Implementation

[0055] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0056] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0057] like Figure 1 As shown, this embodiment of the invention provides a method for generating metadata tags, including:

[0058] Step S101: Extract metadata from multiple data sources to obtain metadata to be processed;

[0059] In step S101, the data source includes, but is not limited to, relational databases (e.g., Structured Query Language, SQL), non-relational databases (e.g., Not Only SQL, NoSQL), data lakes, cloud storage, file systems, and Application Programming Interface (API) services. Before step S101, staff maintain the data source list in the online system. Specific operations include: configuring data source connection parameters, such as database connection strings, API access endpoints, and file paths; data source access authorization information, such as open authorization protocols (OAuth), API keys, usernames, passwords, and other authentication methods (stored in encrypted form); and the identity document (ID) of users who can access the data source, ensuring that only authorized users and the system can access the data. After configuration, the system performs connection tests to ensure successful access to the data source and its business information (e.g., the business system corresponding to the data source, the business system level, and the department to which it belongs).

[0060] When extracting metadata from multiple data sources, a dedicated connector is configured for each data source on the data weaving platform. The connector extracts the corresponding metadata from the data source, which includes at least one of the following: data schema, data table structure, data field definitions, data indexes, and data relationships. The connector handles communication with the data source, data extraction, and transformation. The connector automatically handles details such as connection protocols and authentication mechanisms for different data sources, connecting all data sources in the data weaving platform's list to the weaving network.

[0061] Step S102: Standardize the metadata to be processed using a pre-built metadata model to generate a first data model;

[0062] In step S102, based on the subsequent requirements for building the metadata tagging system and industry standards, a unified metadata model is designed and constructed before metadata access. This metadata model determines the structure of the metadata, standardizes metadata naming conventions (such as field naming conventions), and defines metadata data types, ensuring that fields with the same meaning use consistent names across different data sources. The metadata model is then used to standardize the metadata extracted from the data source, generating a first data model. Each set of metadata to be processed corresponds to one first data model, which includes the corresponding data model (table) for the metadata to be processed, as well as the corresponding field and attribute information.

[0063] Optionally, embodiments of the present invention further include: after step S102, storing the standardized metadata, i.e., the first data model, in a centralized metadata repository; ensuring that any update, deletion, or change of metadata can be recorded and traced back by setting the version control of metadata to allow tracking of all changes to metadata, thus guaranteeing the maintainability of metadata; and ensuring that only authorized users can access or modify metadata by configuring metadata access permissions and access mechanisms. Specifically: 1) Recording change history: Whenever metadata changes (such as modifying fields, adding new tables, etc.), the system generates a new version and records all change details of that version, including change time, person making the change, and content of the change; 2) Version rollback: If there is a problem with the metadata change of a certain version, it can be rolled back to a previous version based on the records; 3) Parallel development: Allowing different teams to modify metadata simultaneously, avoiding metadata corruption caused by multiple change conflicts.

[0064] Optionally, the embodiments of the present invention further include: after step S102, using the inspection tools of the data weaving platform, automatically detecting the consistency of metadata in the first data model, such as field names, data types, indexes and foreign key relationships, automatically generating an audit report on metadata consistency checks, and repairing inconsistent metadata based on the discrepancy audit report to ensure that the metadata of all data sources is consistent with the unified metadata model.

[0065] Step S103: Perform lineage analysis on the first data model using multiple lineage analysis methods to obtain the lineage relationship of the first data model, wherein the lineage relationship is used to indicate the flow information of data in the first data model;

[0066] In step S103, the lineage resolution methods include, but are not limited to: lineage resolution tools, scheduling platform extraction, transformation, and loading (ETL) task parsing, execution engine execution plan parsing, and distributed file system (Hadoop Distributed File System, HDFS) audit log parsing. By using multiple lineage resolution methods to perform lineage resolution on the first data model, data from multiple channels is obtained, thereby generating the lineage relationship of the first data model.

[0067] Step S104: Generate bloodline data for the first data model based on the bloodline relationship of the first data model;

[0068] In step S104, the bloodline data includes, but is not limited to, at least one of the following: statistical indicator information of the first data model, basic indicator information of downstream applications, and atomic metadata tags.

[0069] Step S105: Input the blood relationship data of the first data model into the pre-trained first model to obtain the metadata tag corresponding to the first data model output by the first model.

[0070] In step S105, the metadata tags include at least one of the following: business source tags, data quality tags, business application tags, and comprehensive value tags. Specifically: business source tags include: Tag 1: Business source tag (classification of the main business systems from which the data model originates), used to help understand its business functions, such as financial management, customer relationship management (CRM), e-commerce, etc.; Tag 2: Data collection frequency tag, such as real-time, daily, hourly, weekly, where the frequency of data updates affects the timeliness of the data. Data quality tags include: Tag 3: Data model completion timeliness rate, data model quality alarm frequency; Tag 4: Metadata quality indicators, such as metadata information completeness rate. Business application tags include: Tag 5: Data access frequency and / or dataset access frequency, such as the number of data accesses per month or week, used to reflect data usage; Tag 6: Data impact on business tags, specifically including impact on business level and impact on business category (e.g., company core key performance indicator (KPI) reports, departmental operation reports, online product systems, not currently in use, etc.); and comprehensive value tags are the data model value level comprehensively evaluated based on data quality and data application, such as high value, medium value, and low value.

[0071] The first model includes a decision tree model and a random forest model. The kinship data of the first data model is input into the pre-trained first model to obtain the metadata tags corresponding to the first data model output by the first model.

[0072] In this embodiment of the invention, metadata to be processed is first extracted from multiple data sources. Then, a pre-built metadata model is used to standardize the metadata, generating a first data model. This ensures that the structure and data definitions of metadata from different data sources are the same, facilitating subsequent data processing. This embodiment integrates data from different platforms and formats, providing a unified data access layer. This allows users to access, query, and manage metadata tags from all data sources in a single view, simplifying the complexity of data management. Secondly, multiple lineage resolution methods are used to resolve the lineage of the first data model, obtaining its lineage relationships. Based on these lineage relationships, lineage relationship data is generated for the first data model. By fusing lineage relationships from multiple lineage resolution methods, the lineage coverage and accuracy of the data are improved, resulting in the lineage relationship data for the first data model. Finally, the lineage relationship data is input into a pre-trained first model to obtain the metadata tags corresponding to the first data model output by the first model. This improves the accuracy of the metadata tags, and the automatic generation of tags by the first model enhances the timeliness of the metadata tags.

[0073] Optionally, the step of standardizing the metadata to be processed using a pre-built metadata model to generate a first data model includes:

[0074] According to the pre-set mapping rules, the metadata to be processed is mapped to a pre-built metadata model for standardization processing to generate a first data model;

[0075] If a data definition conflict is detected in the metadata to be processed, the data definition of the metadata to be processed is processed according to a pre-set conflict resolution strategy, and the first data model is updated.

[0076] In this embodiment of the invention, after extracting metadata from the data source, the automated tools of the data weaving platform are used to detect data definition conflicts between multiple pending metadata and data definition conflicts between the pending metadata and metadata in the metadata repository. For example, if data with the same name field corresponds to different data types, it is determined that the corresponding metadata has a data definition conflict; if fields with the same meaning in different systems have different names, it is determined that the corresponding metadata has a data definition conflict.

[0077] According to pre-set mapping rules, the metadata to be processed is mapped to a pre-built metadata model for data transformation and standardization to generate a first data model. For metadata to be processed that has data definition conflicts, the data definitions of the metadata to be processed are processed according to a pre-set conflict resolution strategy, and the first data model is updated. The conflict resolution strategy includes at least one of the following: determining the data definitions of metadata in data sources with higher priority as a unified standard based on the priority of the data sources; determining the data definitions of metadata in data sources selected based on specific rules as a unified standard; resolving data definition conflicts between metadata through methods such as metadata merging, transformation, and renaming.

[0078] Optionally, the lineage relationship includes multiple nodes, edges, and data flow rules between the nodes. The nodes include a first node corresponding to the first data model and a second node corresponding to at least one second data model that has a data flow relationship with the first data model. The edges are used to connect the first node and the second node that have a data flow relationship and to indicate the direction of data flow between the connected first node and the second node. The edges also carry the source channel and lineage resolution method of the lineage relationship.

[0079] In this embodiment of the invention, the nodes include a first node and a second node. The first node includes a data model (table) of the metadata to be processed in the first data model being processed, along with corresponding field and attribute information. The second node includes a data model (table) of metadata in at least one second data model that has a data flow relationship with the first data model, along with corresponding field and attribute information. The sources of the lineage relationships carried by the edges include, but are not limited to, formal scheduling tasks, data complement tasks, and test tasks. The lineage resolution methods include, but are not limited to, lineage resolution tools, scheduling platform ETL task resolution, execution engine execution plan resolution, and HDFS audit log resolution. The data flow rules include cleaning and processing rules during data flow.

[0080] Optionally, the step of using multiple lineage analysis methods to perform lineage analysis on the first data model to obtain the lineage relationship of the first data model includes:

[0081] The first data model is subjected to kinship analysis using multiple kinship analysis methods to obtain kinship analysis results corresponding to the first data model, wherein the kinship analysis results include at least one candidate kinship relationship;

[0082] Based on the bloodline analysis method, source channel, and data processing frequency corresponding to the candidate bloodline relationship, the at least one candidate bloodline relationship is screened according to the first priority order among multiple bloodline analysis methods, the second priority order among multiple source channels, and the third priority order among data processing frequencies to determine the bloodline relationship of the first data model.

[0083] In this embodiment of the invention, different lineage resolution methods are used to perform lineage resolution on the first data model to obtain multiple different lineage relationships of the first data model. Specifically: 1) Lineage resolution tool: Lineage analysis tool (Apache Atlas) is used to perform lineage resolution, automatically capturing and tracking the flow and change process of data models (tables) and corresponding field and attribute information in the first data model between the system and the metadata repository, and constructing lineage relationship 1; 2) ETL task resolution of scheduling platform: The field mapping relationship information in the ETL scheduling task is resolved (for example, data field colA in the source system is mapped to data field colA1 in the data warehouse acquisition layer by removing null values), to obtain lineage relationship 2; 3) Execution engine execution plan resolution: The SQL execution plan in environments including Hive, Spark, Presto, etc. is resolved, to obtain lineage relationship 3; 4) HDFS audit log resolution: The HDFS audit log is resolved, to obtain lineage relationship 4.

[0084] Based on a predefined first priority order among lineage resolution methods, a second priority order among multiple source channels, and a third priority order among data processing frequencies, a three-dimensional rule base is created for lineage resolution methods, source channels, and data processing frequencies. The first priority order among lineage resolution methods is: lineage resolution tool > scheduling platform ETL task parsing > execution engine execution plan parsing > HDFS audit log parsing. The second priority order among source channels is: formal scheduling tasks > data supplementation tasks > test tasks. The third priority order among data processing frequencies is: periodic data processing > temporary data processing. Selecting the final lineage relationship from candidate lineage relationships generated by multiple lineage resolution methods using the three-dimensional rule base improves lineage coverage. It should be noted that while the highest priority candidate lineage relationship is determined as the final lineage relationship based on the three-dimensional rule base, the importance order among lineage resolution methods, source channels, and data processing frequencies is not limited in this invention; it is sufficient to meet actual needs.

[0085] The following specific example illustrates the process of resolving blood relations:

[0086] Perform lineage analysis on data model A (Mobile Customer Call Charges Order Summary Table A): 1) Lineage analysis tool: Lineage relationship 1 is an empty set; 2) ETL task analysis by the scheduling platform: Lineage relationship 2 is an empty set; 3) Execution engine execution plan analysis: Lineage relationship 3 has two candidate lineage relationships, i.e., two flow relationships: In candidate lineage relationship 1, the parent node of data model A is data models B, C, and D, and the source channel is the formal scheduling task; In candidate lineage relationship 2, the parent node of data model A is data models C and D, and the source channel is the data complement task; 4) HDFS audit log analysis: Lineage relationship 3 has one candidate lineage relationship, i.e., one flow relationship: In candidate lineage relationship 3, the parent node of data model A is... Data models B and C originate from scheduling tasks. The data processing frequency of the above candidate lineage relationships is periodic. The selection is based on a three-dimensional rule base set according to the priority of lineage resolution method, source channel, and data processing frequency. Candidate lineage relationship 1 in lineage relationship 3 is prioritized as the final lineage relationship of data model A. Specifically, this includes: nodes (parent nodes: data models B, C, D, child nodes: data model A), edges (B, C, D flow to A, lineage resolution method is SQL execution plan parsing, source channel is formal scheduling task), and data flow rules (data model C filters data with order status of canceled as the main table, and obtains customer basic information from data model B and address information from data model D through left join).

[0087] Optionally, the blood relationship may also include the date of formation of the blood relationship;

[0088] The method further includes:

[0089] Based on the date of bloodline formation, the bloodline relationships that meet preset conditions are pruned and reconstructed, wherein the preset conditions include the time interval between the date of bloodline formation and the current date exceeding a first duration.

[0090] In this embodiment of the invention, after step S102, the standardized first data model is stored in a metadata warehouse. After step S103, the kinship relationship corresponding to the first data model is also stored in the metadata warehouse. Furthermore, the data in the first data model is updated according to the data processing frequency. When the data in the first data model is updated, the kinship relationship corresponding to the first data model is also updated accordingly. Each update generates a corresponding kinship formation date.

[0091] If the time interval between the date of formation of a blood relationship and the current date exceeds a first duration (e.g., 3 months), it indicates that the blood relationship is an invalid historical blood relationship, and the blood relationship is pruned and reconstructed.

[0092] Optionally, after performing kinship analysis on the first data model using multiple kinship analysis methods to obtain the kinship relationships of the first data model, the method further includes:

[0093] In the case where at least one of the second nodes is a temporary node, the first lineage relationship of the temporary node is obtained, wherein the first lineage relationship includes at least one third node that has a data flow relationship with the temporary node and the source channel of the first lineage relationship;

[0094] If the source of the blood relationship of the first node is the same as the source of the first blood relationship, when the temporary node disappears, the blood relationship of the first node is updated by connecting the first node and the third node through an edge.

[0095] In this embodiment of the invention, if a temporary node exists in the lineage relationship, the lineage relationship will have a breakpoint when the temporary node disappears. This is for example, in scenarios where a temporary table is created and deleted simultaneously using the `with as` method or in a script. To solve this problem and further improve lineage coverage, when a temporary node disappears, a connection is made between the first node and a third node that has a data flow relationship with the temporary node. A specific embodiment is provided below to illustrate this:

[0096] Taking data model A as an example, the lineage of data model A specifically includes: nodes (parent nodes: data models B, C, D; child nodes: data model A), edges (flow from B, C, D to A, lineage resolution method is SQL execution plan resolution, source channel is formal scheduling task), and data flow rules. Following the above lineage resolution method, the lineage relationships corresponding to data models B, C, and D are obtained, and then the parent nodes corresponding to data models B, C, and D are obtained. For example, the parent node of data model B is data model B1, the parent node of data model C is data model C1, and the parent nodes of data model D are data models D1 and D2. If data model D is a temporary node, and the temporary node corresponding to data model D disappears, if the source channel in the lineage relationship of data model A is the same as the source channel in the lineage relationship of data model D, then edges connect data model A and data model D1, and data model A and data model D2.

[0097] Optionally, the bloodline relationship further includes at least one of the following: the processing task information of the bloodline relationship carried by the edge, the quality alarm information of the node, and the data usage frequency information of the leaf nodes in the node;

[0098] The method further includes at least one of the following:

[0099] The processing task information of the blood relationship is obtained through the application programming interface (API).

[0100] The quality result data corresponding to the first data model is obtained through the data quality control platform, and quality alarm information corresponding to each node is generated.

[0101] The frequency of data usage of the leaf nodes is obtained through the application-side operation monitoring interface.

[0102] In this embodiment of the invention, after step S103, the information of nodes and edges in the blood relationship is further enriched, specifically including the following three methods:

[0103] 1. Integrate the scheduling task API interface to obtain the processing task information of the bloodline relationship, and add the processing task information to the information carried by the edge in the bloodline relationship. The processing task information includes, but is not limited to, the person in charge information and the execution cycle.

[0104] 2. First, access the quality result data that needs to be inspected from the data quality control platform. This quality result data includes, but is not limited to, the number of alarms, alarm types, response times, and severity. Then, clean and standardize the quality result data. Specific operations include: noise removal, handling missing values, and standardizing data formats. Finally, based on the cleaned and standardized quality result data, generate quality alarm information corresponding to each node. This quality alarm information includes, for example, the average delay time over the past month, the frequency of quality alarms over the past month, and the frequency of the most recent severe alarm.

[0105] 3. Integrate the application-side operation monitoring interface to obtain and supplement information such as the data usage frequency of leaf nodes (e.g., online systems, reports) in the lineage relationship.

[0106] Optionally, the generation of kinship data for the first data model based on the kinship relationship of the first data model includes:

[0107] Input the bloodline relationships from the first data model into the graph database to generate a bloodline database;

[0108] The identification information of the first data model is input into the lineage database to obtain the initial lineage relationship data corresponding to the first data model output by the lineage database. The initial lineage relationship data includes at least one of the following: statistical indicator information of the first data model, basic indicator information of downstream applications, and atomic metadata tags.

[0109] The initial bloodline data is standardized and normalized to obtain bloodline data.

[0110] In this embodiment of the invention, the lineage relationship of the first data model is input into a graph database to generate a lineage database. The graph database includes, but is not limited to, graph databases (Neo4j), multimodal databases (ArangoDB), or managed graph databases (Amazon Neptune). In practical applications, the graph database into which the lineage relationship of the first data model is input can be the original graph database or a graph database that includes the lineage relationship of metadata in a metadata repository.

[0111] After generating the lineage database, the identification information of the first data model is input into the lineage database for querying to obtain the initial lineage relationship data corresponding to the first data model output by the lineage database. The initial lineage relationship data includes, but is not limited to, at least one of the following: statistical indicator information, basic indicator information of downstream applications, and atomic metadata tags. The statistical indicator information includes, but is not limited to, business domain, subject domain, metadata integrity, model collection count, model processing frequency, model recent processing time, data completion timeliness, and data accuracy. The basic indicator information of downstream applications includes, but is not limited to, platform downstream call count, platform downstream access count, number of downstream users, number of terminal application systems, and downstream system level. The atomic metadata tags include, but are not limited to, downstream application type (e.g., report, core KPI analysis dashboard, application system), downstream application name, and data business flow description (flow from the order system and marketing system to the core KPI analysis dashboard).

[0112] Optionally, the step of inputting the identification information of the first data model into the kinship database to obtain the initial kinship data corresponding to the first data model output by the kinship database includes at least one of the following:

[0113] The identification information of the first data model is input into the bloodline bank, and the statistical index information of the first data model output by the bloodline bank is obtained by using the application programming interface (API) call method.

[0114] The identification information of the first data model is input into the lineage database, and the basic indicator information and atomic metadata tags of the downstream application corresponding to the first data model output by the lineage database are obtained using the depth-first search (DFS) algorithm.

[0115] In this embodiment of the invention, querying the initial kinship data corresponding to the first data model in the kinship database involves the following two parts: First, inputting the identification information of the first data model into the kinship database and obtaining the statistical indicator information of the first data model output by the kinship database using an API call method; Second, inputting the identification information of the first data model into the kinship database and outputting the basic indicator information and atomic metadata tags of the downstream applications of the first data model using a depth-first search (DFS) algorithm.

[0116] Optionally, the step of inputting the kinship data of the first data model into the pre-trained first model to obtain the metadata tag corresponding to the first data model output by the first model includes:

[0117] The blood relationship data of the first data model is input into the decision tree model in the pre-trained first model to obtain the target blood relationship data output by the decision tree model, wherein the decision tree model is used to filter the blood relationship data according to the importance of the data features;

[0118] The target bloodline data of the first data model is input into the random forest model in the pre-trained first model to obtain the metadata tags corresponding to the first data model output by the random forest model.

[0119] In this embodiment of the invention, the first model includes a decision tree model and a random forest model. The decision tree model is used to evaluate the importance of each data feature in the kinship data, and the random forest model is used to automatically generate metadata tags corresponding to the first data model. In application, firstly, the kinship data of the first data model is input into the pre-trained decision tree model to obtain the importance of each data feature in the kinship data. Then, based on the importance, the data in the kinship data is filtered to obtain target kinship data. Finally, the target kinship data is input into the pre-trained random forest model to obtain the metadata tags corresponding to the first data model output by the random forest model.

[0120] Optionally, after obtaining the metadata tag corresponding to the first data model output by the first model, the method further includes:

[0121] When an abnormal output of the first model is detected, a corresponding alarm message is generated and a corresponding recovery strategy is executed, wherein the recovery strategy is used to resolve the abnormal output of the first model;

[0122] Upon receiving feedback from the user regarding the metadata tags output by the first model, the first model is optimized based on the feedback.

[0123] In this embodiment of the invention, an anomaly handling and alarm mechanism and a feedback mechanism are also established. After generating metadata tags corresponding to metadata using the first model, the generation of metadata tags is optimized through the above mechanisms, as follows:

[0124] 1) Anomaly Handling and Alarm Mechanism: When an anomaly is detected in the output of the first model (e.g., inaccurate metadata tags, excessively long response time), a corresponding alarm message is generated to alert the first model to the output anomaly. Based on the anomaly, a corresponding recovery strategy is automatically executed, including automatically restarting relevant services or rolling back to a safe state. For example, when inaccurate metadata tags are detected in the output of the first model, specific tag information is obtained through rule validation and automatic data content verification (e.g., whether the metadata tags conform to predefined standards, whether key attributes are missing, whether the data type is consistent with the actual data, comparison of metadata tags with the reference model, etc.), and then an alarm message is generated based on the tag information.

[0125] 2) Feedback mechanism: The data weaving platform provides a user feedback interface, allowing users to evaluate, correct, or supplement the automatically generated metadata tags. The manually processed metadata tags are used as new training data to further train or fine-tune the first model, optimize the generation of metadata tags, identify the weak links of the model based on user feedback, and prioritize the optimization of the tag generation logic in the weak areas, so that the metadata tags generated by the first model are more accurate.

[0126] Optionally, the method further includes:

[0127] Metadata is extracted from multiple data sources to obtain multiple sets of initial metadata;

[0128] This is used to standardize the first metadata using a pre-built metadata model to generate a second data model, wherein a set of first metadata corresponds to a second data model;

[0129] For each second data model, multiple lineage resolution methods are used to perform lineage resolution on the second data model to obtain the lineage relationship of the second data model, wherein the lineage relationship is used to indicate the data flow information in the second data model;

[0130] Based on the blood relationship, generate blood relationship data corresponding to each of the second data models;

[0131] A first dataset is constructed based on the identification information corresponding to multiple second data models, pre-labeled metadata tags, and the blood relationship data;

[0132] The data in the first dataset is input into a pre-trained decision tree model for filtering to obtain a second dataset, wherein the decision tree model is used to filter the data in the first dataset according to the importance of the data features;

[0133] The random forest model was trained using the second dataset.

[0134] In this embodiment of the invention, the training process of the random forest model in the first model is described: It should be noted that the specific methods for extracting multiple sets of first metadata from the data source, standardizing the first metadata, and obtaining the lineage relationship and lineage relationship data of each second data model are the same as the methods used in the above-mentioned metadata tag generation process, and will not be elaborated here.

[0135] First, after acquiring the lineage data corresponding to each of the second data models, the lineage data is standardized and normalized to form standardized first lineage data. Then, a first dataset is constructed with the second data models as the granularity. The first dataset includes: identification information corresponding to multiple second data models, pre-labeled metadata tags, and the lineage data. The identification information includes, but is not limited to, at least one of the following: the data model's identity document (ID) and the data model name. The metadata tags are labeled by data warehouse experts and business personnel based on business knowledge and historical data. The comprehensive value tag in the metadata tags is high-level, medium-level, or low-level. Second, a pre-trained decision tree model is used to evaluate the importance of each data feature in the first dataset, and selection is performed based on importance to obtain the second dataset. This improves the model's generalization ability and avoids overfitting. Finally, the second dataset is used to train the random forest model. The specific training process is as follows:

[0136] The first step is to randomly divide the second dataset into a training set and a test set according to the first preset ratio;

[0137] The second step is to determine the key hyperparameters of the random forest model, such as the number of trees (n_estimators), the maximum tree depth (max_depth), and the minimum number of sample splits (min_samples_split), and then train the random forest model using the training set.

[0138] The third step is to use the test set to predict the random forest model. Based on the accuracy, recall, F1 score and other metrics of the label classification results and the labeled metadata results output by the test set, the random forest model is evaluated on unseen data. Based on the evaluation results, the parameters are tuned, and the final random forest model is output.

[0139] In summary, the metadata tag generation method provided in this invention allows data access across multiple data sources, platforms, and environments during metadata collection, including databases, data lakes, data warehouses, and cloud storage, providing a rich data foundation for tag generation. By using a rule engine and employing methods such as lineage reconstruction and lineage pruning to integrate data lineage relationships from multiple channels and parsing methods, the method improves data lineage coverage and accuracy. Furthermore, by utilizing pre-trained decision tree and random forest models, the method automatically generates metadata tags, improving the efficiency of metadata tag generation.

[0140] like Figure 2 As shown, this embodiment of the invention also provides a metadata tag generation apparatus, including:

[0141] The first extraction module 201 is used to extract metadata from multiple data sources to obtain metadata to be processed;

[0142] The first processing module 202 is used to standardize the metadata to be processed using a pre-built metadata model to generate a first data model;

[0143] The first parsing module 203 is used to perform lineage parsing on the first data model using multiple lineage parsing methods to obtain the lineage relationship of the first data model, wherein the lineage relationship is used to indicate the flow information of data in the first data model;

[0144] The first generation module 204 is used to generate blood relationship data of the first data model based on the blood relationship of the first data model.

[0145] The second processing module 205 is used to input the blood relationship data of the first data model into the pre-trained first model to obtain the metadata tag corresponding to the first data model output by the first model.

[0146] Optionally, the first processing module 202 includes:

[0147] The first processing submodule is used to map the metadata to be processed to a pre-built metadata model for standardization processing according to a pre-set mapping rule, and generate a first data model.

[0148] The second processing submodule is used to process the data definition of the metadata to be processed according to a pre-set conflict resolution strategy when a data definition conflict is detected in the metadata to be processed, and update the first data model.

[0149] Optionally, the lineage relationship in the first parsing module 203 includes multiple nodes, edges, and data flow rules between the nodes. The nodes include a first node corresponding to the first data model and a second node corresponding to at least one second data model that has a data flow relationship with the first data model. The edges are used to connect the first node and the second node that have a data flow relationship and to indicate the data flow direction between the connected first node and the second node. The edges also carry the source channel and lineage parsing method of the lineage relationship.

[0150] Optionally, the first parsing module 203 includes:

[0151] The first parsing submodule is used to perform kinship analysis on the first data model using multiple kinship analysis methods to obtain the kinship analysis result corresponding to the first data model, wherein the kinship analysis result includes at least one candidate kinship relationship;

[0152] The first screening submodule is used to screen at least one candidate blood relationship based on the blood relationship analysis method, source channel and data processing frequency corresponding to the candidate blood relationship, according to the first priority order among multiple blood relationship analysis methods, the second priority order among multiple source channels and the third priority order among data processing frequencies, to determine the blood relationship of the first data model.

[0153] Optionally, the blood relationship in the first parsing module 203 may also include the date of blood formation;

[0154] The device further includes:

[0155] The first pruning module is used to prune and reconstruct the blood relationship that meets preset conditions based on the blood relationship formation date, wherein the preset conditions include the time interval between the blood relationship formation date and the current date exceeding a first duration.

[0156] Optionally, the device further includes:

[0157] The first acquisition module is used to acquire the first lineage relationship of the temporary node when at least one of the second nodes is a temporary node. The first lineage relationship includes at least one third node that has a data flow relationship with the temporary node and the source channel of the first lineage relationship.

[0158] The first update module is used to update the bloodline relationship of the first node by connecting the first node and the third node through an edge when the temporary node disappears, provided that the source channel of the bloodline relationship of the first node is the same as the source channel of the bloodline relationship.

[0159] Optionally, the bloodline relationship in the first parsing module 203 may further include at least one of the following: the processing task information of the bloodline relationship carried by the edge, the quality alarm information of the node, and the data usage frequency information of the leaf nodes in the node;

[0160] The device further includes at least one of the following:

[0161] The second acquisition module is used to acquire the processing task information of the blood relationship through the application programming interface (API).

[0162] The second generation module is used to obtain the quality result data corresponding to the first data model through the data quality control platform and generate quality alarm information corresponding to each node.

[0163] The third acquisition module is used to acquire the data usage frequency information of the leaf node through the application-side operation monitoring interface.

[0164] Optionally, the first generation module 204 includes:

[0165] The first generation submodule is used to input the bloodline relationship of the first data model into the graph database to generate the bloodline database;

[0166] The third processing submodule is used to input the identification information of the first data model into the bloodline database to obtain the initial bloodline relationship data corresponding to the first data model output by the bloodline database. The initial bloodline relationship data includes at least one of the following: statistical indicator information of the first data model, basic indicator information of downstream applications, and atomic metadata tags.

[0167] The fourth processing submodule is used to standardize and normalize the initial blood relationship data to obtain blood relationship data.

[0168] Optionally, the third processing submodule includes at least one of the following:

[0169] The first processing unit is used to input the identification information of the first data model into the bloodline bank and obtain the statistical indicator information of the first data model output by the bloodline bank by using the application programming interface (API) call method.

[0170] The second processing unit is used to input the identification information of the first data model into the lineage database, and use the depth-first search (DFS) algorithm to obtain the basic indicator information and atomic metadata tags of the downstream application corresponding to the first data model output by the lineage database.

[0171] Optionally, the second processing module 205 includes:

[0172] The fifth processing submodule is used to input the blood relationship data of the first data model into the decision tree model in the pre-trained first model to obtain the target blood relationship data output by the decision tree model, wherein the decision tree model is used to filter the blood relationship data according to the importance of the data features;

[0173] The sixth processing submodule is used to input the target blood relationship data of the first data model into the random forest model in the pre-trained first model, and obtain the metadata tags corresponding to the first data model output by the random forest model.

[0174] Optionally, the device further includes:

[0175] The third processing module is used to generate corresponding alarm information and execute corresponding recovery strategies when an output anomaly of the first model is detected, wherein the recovery strategies are used to resolve the output anomaly of the first model.

[0176] The first optimization module is used to optimize the first model based on the feedback information received from the user regarding the metadata tags output by the first model.

[0177] It should be noted that the embodiments of this device are devices corresponding to the embodiments of the above methods. All implementations in the embodiments of the above methods are applicable to the embodiments of this device and can achieve the same technical effect.

[0178] This invention also provides a network device, including: a processor, a memory, and a program stored in the memory and executable on the processor. When the program is executed by the processor, it implements the metadata tag generation method as described in any of the preceding claims and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0179] This invention also provides a readable storage medium, comprising: a program stored on the readable storage medium, wherein when the program is executed by a processor, it implements the steps of the metadata tag generation method described in any of the preceding claims, and achieves the same technical effect; to avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0180] This invention also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the metadata tag generation method as described in any of the preceding claims, and achieve the same technical effect. To avoid repetition, these will not be repeated here.

[0181] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0182] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for generating metadata tags, characterized in that, include: Metadata is extracted from multiple data sources to obtain the metadata to be processed; The metadata to be processed is standardized using a pre-built metadata model to generate a first data model; The first data model is analyzed using multiple lineage analysis methods to obtain the lineage relationship of the first data model, wherein the lineage relationship is used to indicate the flow information of data in the first data model; Generate bloodline data for the first data model based on the bloodline relationships in the first data model; The blood relationship data of the first data model is input into the pre-trained first model to obtain the metadata tag corresponding to the first data model output by the first model.

2. The method for generating metadata tags according to claim 1, characterized in that, The step of standardizing the metadata to be processed using a pre-built metadata model to generate a first data model includes: According to the pre-set mapping rules, the metadata to be processed is mapped to a pre-built metadata model for standardization processing to generate a first data model; If a data definition conflict is detected in the metadata to be processed, the data definition of the metadata to be processed is processed according to a pre-set conflict resolution strategy, and the first data model is updated.

3. The method for generating metadata tags according to claim 1, characterized in that, The lineage relationship includes multiple nodes, edges, and data flow rules between the nodes. The nodes include a first node corresponding to the first data model and a second node corresponding to at least one second data model that has a data flow relationship with the first data model. The edges are used to connect the first node and the second node that have a data flow relationship and to indicate the direction of data flow between the connected first node and the second node. The edges also carry the source channel and lineage resolution method of the lineage relationship.

4. The method for generating metadata tags according to claim 1, characterized in that, The step of using multiple kinship analysis methods to perform kinship analysis on the first data model to obtain the kinship relationships of the first data model includes: The first data model is subjected to kinship analysis using multiple kinship analysis methods to obtain kinship analysis results corresponding to the first data model, wherein the kinship analysis results include at least one candidate kinship relationship; Based on the bloodline analysis method, source channel, and data processing frequency corresponding to the candidate bloodline relationship, the at least one candidate bloodline relationship is screened according to the first priority order among multiple bloodline analysis methods, the second priority order among multiple source channels, and the third priority order among data processing frequencies to determine the bloodline relationship of the first data model.

5. The method for generating metadata tags according to claim 1, characterized in that, The blood relationship also includes the date of formation of the blood relationship; The method further includes: Based on the date of bloodline formation, the bloodline relationships that meet preset conditions are pruned and reconstructed, wherein the preset conditions include the time interval between the date of bloodline formation and the current date exceeding a first duration.

6. The method for generating metadata tags according to claim 3, characterized in that, After performing kinship analysis on the first data model using multiple kinship analysis methods to obtain the kinship relationships in the first data model, the method further includes: In the case where at least one of the second nodes is a temporary node, the first lineage relationship of the temporary node is obtained, wherein the first lineage relationship includes at least one third node that has a data flow relationship with the temporary node and the source channel of the first lineage relationship; If the source of the blood relationship of the first node is the same as the source of the first blood relationship, when the temporary node disappears, the blood relationship of the first node is updated by connecting the first node and the third node through an edge.

7. The method for generating metadata tags according to claim 3, characterized in that, The bloodline relationship also includes at least one of the following: the processing task information of the bloodline relationship carried by the edge, the quality alarm information of the node, and the data usage frequency information of the leaf nodes in the node; The method further includes at least one of the following: The processing task information of the blood relationship is obtained through the application programming interface (API). The quality result data corresponding to the first data model is obtained through the data quality control platform, and quality alarm information corresponding to each node is generated. The frequency of data usage of the leaf nodes is obtained through the application-side operation monitoring interface.

8. The method for generating metadata tags according to claim 1, characterized in that, The step of generating kinship data for the first data model based on the kinship relationships in the first data model includes: Input the bloodline relationships from the first data model into the graph database to generate a bloodline database; The identification information of the first data model is input into the lineage database to obtain the initial lineage relationship data corresponding to the first data model output by the lineage database. The initial lineage relationship data includes at least one of the following: statistical indicator information of the first data model, basic indicator information of downstream applications, and atomic metadata tags. The initial bloodline data is standardized and normalized to obtain bloodline data.

9. The method for generating metadata tags according to claim 8, characterized in that, The step of inputting the identification information of the first data model into the kinship database to obtain the initial kinship data corresponding to the first data model output by the kinship database includes at least one of the following: The identification information of the first data model is input into the bloodline bank, and the statistical index information of the first data model output by the bloodline bank is obtained by using the application programming interface (API) call method. The identification information of the first data model is input into the lineage database, and the basic indicator information and atomic metadata tags of the downstream application corresponding to the first data model output by the lineage database are obtained using the depth-first search (DFS) algorithm.

10. The method for generating metadata tags according to claim 1, characterized in that, The step of inputting the kinship data of the first data model into the pre-trained first model to obtain the metadata tag corresponding to the first data model output by the first model includes: The blood relationship data of the first data model is input into the decision tree model in the pre-trained first model to obtain the target blood relationship data output by the decision tree model, wherein the decision tree model is used to filter the blood relationship data according to the importance of the data features; The target bloodline data of the first data model is input into the random forest model in the pre-trained first model to obtain the metadata tags corresponding to the first data model output by the random forest model.

11. The method for generating metadata tags according to claim 1, characterized in that, After obtaining the metadata tags corresponding to the first data model output by the first model, the method further includes: When an abnormal output of the first model is detected, a corresponding alarm message is generated and a corresponding recovery strategy is executed, wherein the recovery strategy is used to resolve the abnormal output of the first model; Upon receiving feedback from the user regarding the metadata tags output by the first model, the first model is optimized based on the feedback.

12. A metadata tag generation apparatus, characterized in that, include: The first extraction module is used to extract metadata from multiple data sources to obtain metadata to be processed. The first processing module is used to standardize the metadata to be processed using a pre-built metadata model to generate a first data model. The first parsing module is used to perform lineage analysis on the first data model using multiple lineage analysis methods to obtain the lineage relationship of the first data model, wherein the lineage relationship is used to indicate the flow information of data in the first data model; The first generation module is used to generate blood relationship data of the first data model based on the blood relationship of the first data model; The second processing module is used to input the blood relationship data of the first data model into the pre-trained first model to obtain the metadata tag corresponding to the first data model output by the first model.

13. A network device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, the program, when executed by the processor, implementing the method for generating metadata tags as described in any one of claims 1 to 11.

14. A readable storage medium, characterized in that, include: The readable storage medium stores a program that, when executed by a processor, implements the steps of the metadata tag generation method as described in any one of claims 1 to 11.

15. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the method for generating metadata tags as described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Method and device for constructing metadata tag library

    CN113360496A

  • Electronic license construction method and device

    CN117093831A

  • Method, system and equipment for monitoring abnormal calling of API gateway of big data platform

    CN119583140A

  • Data blood relationship analysis method and device for ETL system and electronic equipment

    CN119646073A

  • Data blood relationship analysis visualization method and device, equipment and medium

    CN120386893A