Metadata Management Method, System, Device, and Storage Medium

By extracting and generating metadata attribute information, classifying and storing, and calculating abnormal warnings using indicator rules, the problem of difficulty in metadata governance in the financial industry is solved, and efficient management and quality improvement of metadata are achieved.

CN119669166BActive Publication Date: 2025-06-24EVERGROWING BANK CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510180989.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-24
Estimated Expiration
2045-02-19

AI Technical Summary

Technical Problem

The financial industry relies on the professional quality of developers in controlling existing metadata, which makes it difficult to manage metadata and does not adapt to the complexity of metadata.

Method used

By extracting the characteristic information of the metadata, generating attribute information, and assigning key prefixes to the metadata based on these attributes, for classification storage. Use preset indicator rules to calculate the indicator value of metadata and generate an abnormal warning when the indicator value exceeds the threshold.

Benefits of technology

It realizes efficient management of metadata, and through classified storage and metric calculation, the controllability and quality of metadata are improved and the dependence on developers is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119669166B_ABST
    Figure CN119669166B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data processing, and specifically provides a metadata management method, system, device, and storage medium, including: extracting feature information of metadata and generating attribute information based on the feature information; allocating a key prefix for the metadata based on the attribute information, and saving the metadata in a pre-allocated logical range of the key prefix in the form of key-value pairs; calculating index values of the attribute information of the metadata in each logical range respectively by using a preset index rule; and confirming that the index value of the logical range exceeds a set index threshold, and generating a metadata anomaly warning. The present invention realizes efficient management of metadata.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and particularly relates to a metadata management method, system, device and storage medium. Background Art

[0002] Metadata, also known as intermediate data or relay data, is data about data, mainly information describing the properties of data, used to support functions such as indicating storage locations, historical data, resource search, file records, etc. Metadata is a kind of electronic directory. To achieve the purpose of compiling a directory, it is necessary to describe and collect the content or characteristics of the data, so as to assist in data retrieval.

[0003] Currently, the control of existing metadata in the financial industry is still in a passive cycle that basically relies on the professional qualities of developers. That is, the current governance of metadata severely depends on the rules set by developers. However, due to the complexity and diversity of metadata, it may not necessarily fit the original development rules, so the difficulty of metadata governance has encountered a bottleneck. Summary of the Invention

[0004] In view of the above deficiencies of the prior art, the present invention provides a metadata management method, system, device and storage medium to solve the above technical problems.

[0005] In a first aspect, the present invention provides a metadata management method, including:

[0006] Extracting characteristic information of metadata and generating attribute information based on the characteristic information;

[0007] Allocating a key prefix for the metadata based on the attribute information and saving the metadata in a pre-allocated logical range of the key prefix in the form of key-value pairs;

[0008] Calculating index values of the attribute information of the metadata in each logical range respectively using a preset index rule;

[0009] Confirming that the index value of the logical range exceeds a set index threshold and generating a metadata anomaly warning.

[0010] In an optional embodiment, extracting characteristic information of metadata and generating attribute information based on the characteristic information includes:

[0011] Extract feature information from the ledger and metadata directory using an entity naming model. The feature information includes attribution source information, technical information, and resource information. The attribution source information includes the attribution application system, business management identity information, and technical management identity information. The technical information includes database information, schema information, table-level information, field-level information, and index information. The resource information includes data resource attribution identity information, data resource status, and attribution development platform information.

[0012] Set the attribution source information as the management attribute of the metadata, set the technical information as the technical attribute of the metadata, and set the resource information as the resource archiving attribute of the metadata.

[0013] According to the technical information and pre-set technical standard parameters, generate standard comparison information based on whether the technical information conforms to the corresponding technical standard parameters, and set the standard comparison information as the data comparison attribute of the metadata. The technical standard parameters include the mapped standard name, standard data type, standard data length, standard data precision, and standard code item corresponding to the combination of the table name and field name.

[0014] In an optional implementation, allocate a key prefix for the metadata based on the attribute information, and save the metadata in the pre-allocated logical range of the key prefix in the form of key-value pairs, including:

[0015] Determine the data type of the metadata according to the attribute information.

[0016] Allocate a key prefix for the metadata according to the correspondence between the data type and the key prefix and the determined data type.

[0017] Pre-allocate a logical range for each key prefix, and store the data type, key prefix, and logical range address of the metadata in a configuration file.

[0018] Serialize the metadata and store the processed metadata in the corresponding logical range.

[0019] In an optional implementation, calculate the index values of the attribute information of the metadata in each logical range using preset index rules, including:

[0020] Set the weights of each parameter item of the data comparison attribute, and calculate the weighted sum to obtain the data item-level standard compliance.

[0021] Perform weighted summation according to the preset data item weights and the item-level standard compliance to obtain the table-level standard compliance.

[0022] Perform weighted summation according to the preset table weights and the table-level standard compliance to obtain the application-level standard compliance.

[0023] In an alternative embodiment, when the metric value of the confirmation logic interval exceeds the set metric threshold, a metadata anomaly warning is generated, including:

[0024] Set a threshold combination for each logic interval, and the threshold combination includes a data item - level standard compliance threshold, a table - level standard compliance threshold, and an application - level standard compliance threshold;

[0025] Compare the data item - level standard compliance, table - level standard compliance, and application - level standard compliance within the logic interval with the corresponding thresholds. If it is lower than the corresponding threshold, generate the corresponding anomaly warning;

[0026] Use the One - Class SVM model to analyze the boundary values of each logic interval and determine the abnormal metadata based on the boundary values.

[0027] In a second aspect, the present invention provides a metadata management system, including:

[0028] An attribute generation module, configured to extract the feature information of the metadata and generate attribute information based on the feature information;

[0029] A data storage module, configured to allocate a key prefix for the metadata based on the attribute information and store the metadata in the pre - allocated logic interval of the key prefix in the form of key - value pairs;

[0030] A metric calculation module, configured to calculate the metric values of the attribute information of the metadata in each logic interval respectively by using preset metric rules;

[0031] An anomaly warning module, configured to confirm that the metric value of the logic interval exceeds the set metric threshold and generate a metadata anomaly warning.

[0032] In an alternative embodiment, the attribute generation module includes:

[0033] A feature extraction unit, configured to extract feature information from the ledger and metadata directory by using an entity naming model. The feature information includes attribution source information, technical information, and resource information; the attribution source information includes attribution application system, business management identity information, and technical management identity information; the technical information includes database information, schema information, table - level information, field - level information, and index information; the resource information includes data resource attribution identity information, data resource status, and attribution development platform information;

[0034] A first setting unit, configured to set the attribution source information as the management attribute of the metadata, set the technical information as the technical attribute of the metadata, and set the resource information as the resource filing attribute of the metadata;

[0035] A second setting unit, configured to generate standard comparison information according to whether the technical information conforms to corresponding technical standard parameters based on the technical information and preset technical standard parameters, and set the standard comparison information as the data comparison attribute of the metadata; the technical standard parameters include mapping standard names, standard data types, standard data lengths, standard data precisions, and standard code items corresponding to combinations of table names and field names.

[0036] In an optional embodiment, the data storage module includes:

[0037] A data classification unit, configured to determine the data type of the metadata according to the attribute information;

[0038] An identifier allocation unit, configured to allocate a key prefix for the metadata according to the correspondence between the data type and the key prefix and the determined data type;

[0039] A configuration storage unit, configured to pre-allocate a logical range for each key prefix, and store the data type, key prefix, and logical range address of the metadata in a configuration file;

[0040] An encoding storage unit, configured to perform serialization processing on the metadata and store the processed metadata in a corresponding logical range.

[0041] In a third aspect, a device is provided, including:

[0042] A memory, configured to store a metadata management program;

[0043] A processor, configured to implement the steps of the metadata management method provided in the first aspect when executing the metadata management program.

[0044] In a fourth aspect, a computer-readable storage medium is provided, on which a metadata management program is stored, and when the metadata management program is executed by a processor, the steps of the metadata management method provided in the first aspect are implemented.

[0045] The beneficial effects of the present invention are that the metadata management method, system, device, and storage medium provided by the present invention realize efficient management of metadata by generating attributes for the metadata and storing them classified based on the attributes, and simultaneously calculating metrics and detecting anomalies for the metadata in different storage areas.

[0046] In addition, the design principle of the present invention is reliable, the structure is simple, and it has a very wide application prospect. Description of the Drawings

[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0048] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention.

[0049] Figure 2 It is a schematic diagram of a scenario of the method according to an embodiment of the present invention.

[0050] Figure 3 It is a refined schematic diagram of a scenario of the method according to an embodiment of the present invention.

[0051] Figure 4 It is a schematic block diagram of the system according to an embodiment of the present invention.

[0052] Figure 5 It is a schematic structural diagram of a device provided by the embodiment of the present invention. Detailed implementation manners

[0053] In order to enable those skilled in the art of the present technology to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the description of the present invention in this specification are only for the purpose of describing specific embodiments, and are not intended to limit the present invention.

[0055] The metadata management method provided by the embodiment of the present invention is executed by a computer device. Correspondingly, the metadata management system runs in the computer device.

[0056] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention. Among them, Figure 1 The execution subject can be a metadata management system. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted.

[0057] Such as Figure 1 shown, the method includes:

[0058] S1. Extract the feature information of the metadata and generate attribute information based on the feature information;

[0059] S2. Assign a key prefix to the metadata based on the attribute information and save the metadata in the pre-allocated logical range of the key prefix in the form of key-value pairs;

[0060] S3. Calculate the index values of the attribute information of the metadata in each logical range respectively using the preset index rules;

[0061] S4. Confirm that the index value of the logical range exceeds the set index threshold and generate an abnormal warning for the metadata.

[0062] In an embodiment of the present invention, based on step S1, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation scheme.

[0063] S101. Use the entity naming model to extract feature information from the ledger and metadata directory. The feature information includes attribution source information, technical information, and resource information. The attribution source information includes the attribution application system, business management identity information, and technical management identity information. The technical information includes database information, schema information, table-level information, field-level information, and index information. The resource information includes data resource attribution identity information, data resource status, and attribution development platform information.

[0064] The entity naming model is an intelligent model constructed based on advanced natural language processing and machine learning technologies. It can accurately identify and extract various entity information in text data. In the data processing process, the ledger and metadata directory are used as input data sources to give full play to the powerful information extraction ability of the entity naming model.

[0065] Through the operation of the entity naming model, multi-dimensional feature information is successfully extracted from the ledger and metadata directory. These feature information cover important aspects such as attribution source information, technical information, and resource information. Each type of information provides a unique perspective and key clues for in-depth understanding of the data.

[0066] First, look at the attribution source information. It is the core element for understanding the data source and ownership. The attribution source information is further divided into the attribution application system, business management identity information, and technical management identity information. The attribution application system clarifies the specific application program or system from which the data originates, which is of great significance for tracing the generation background and business process of the data. For example, a batch of data may originate from the enterprise's customer relationship management system (CRM). By clarifying the attribution application system, it is possible to clearly understand the close connection between these data and customer management business.

[0067] Business management identity information involves the relevant information of personnel at the business level responsible for data management and decision-making. Such information includes the names, positions, and departments of business leaders, who play a crucial role in business planning for data, usage decisions, and coordination with other business processes. Knowing business management identity information helps in better communication and collaboration with relevant business personnel during data analysis and application, ensuring that data usage aligns with business needs and strategic directions.

[0068] Technical management identity information focuses on department information related to managing and maintaining data from a technical perspective. They are responsible for technical tasks such as data storage, processing, and security, which are decisive for data quality and usability. Identifying technical management identity information enables quick identification of relevant responsible persons in case of data technical problems, facilitating timely issue resolution and ensuring normal data flow and usage.

[0069] Next is technical information, which deeply reveals the composition and characteristics of data at the technical level. Technical information includes multiple levels such as database information, schema information, table-level information, field-level information, and index information.

[0070] Database information covers key aspects such as the type of database (e.g., relational database, non-relational database), version, and storage location. Different types of databases have their own unique advantages and application scenarios. Understanding database information helps in selecting appropriate database management tools and technical means based on data characteristics and business requirements. For example, for highly structured data with strict transaction processing requirements, a relational database might be a better choice; while for handling massive amounts of unstructured data, a non-relational database has more advantages.

[0071] Schema information describes the organization and structure of data in the database, including the architecture design of the database and the relationships between tables. Clear schema information provides important guidance for data query, analysis, and maintenance, enabling accurate understanding of the logical relationships between data and avoiding errors during data operations.

[0072] Table-level information provides detailed information for each table in the database, such as the table name, description, creation time, and modification time. This information helps in quickly understanding the purpose and lifecycle of each table, and in judging the timeliness and usability of data. For example, by checking the creation time and modification time of a table, the update frequency of data can be evaluated, and thus it can be determined whether the data needs to be updated or maintained in a timely manner.

[0073] Field-level information delves into each field in the table, detailing attributes such as the field name, data type, length, and whether it can be null. Field-level information is crucial for data entry, querying, and analysis. It determines the data format and value range, ensuring data accuracy and consistency. For example, when entering data, corresponding validation rules can be set according to the field attributes to prevent the entry of incorrect data.

[0074] Index information is a data structure established to improve data query efficiency. Through the entity naming model, the index information extracted can show which fields in the database have indexes, as well as the index type and usage. A reasonable index design can significantly enhance the data query speed and optimize the database performance.

[0075] Finally, there is the resource information, which comprehensively reflects the situation of data resources in terms of ownership, status, and development platform, etc. Resource information includes data resource ownership identity information, data resource status, and information on the affiliated development platform.

[0076] Data resource ownership identity information clarifies the owner or owning team of the data resource, which is of great significance for data permission management and responsibility division. Only by clarifying the ownership identity of the data can the legal use and protection of the data be ensured, and the risks of data leakage and abuse be avoided.

[0077] Data resource status describes the current status of the data resource, such as whether it is available, whether it is being updated, whether there are errors, etc. Timely understanding of the data resource status helps to reasonably arrange data usage and processing tasks, and avoid business interruptions or errors caused by abnormal data status. For example, if the data resource is being updated, relevant data query operations can be suspended and resumed after the update is completed to obtain the latest and accurate data.

[0078] Information on the affiliated development platform indicates the development platform or technical framework on which the data resource depends. Different development platforms have different characteristics and functions. Understanding the information on the affiliated development platform helps to better understand the development background and technical environment of the data, and provides an important reference basis for subsequent data migration, upgrade, or integration.

[0079] In summary, by using the entity naming model to extract multi-dimensional characteristic information covering source information, technical information, and resource information from the ledger and metadata catalog, a comprehensive and detailed data portrait can be constructed, providing a solid foundation for data management, analysis, and application, and helping enterprises make more informed decisions in the digital age.

[0080] S102. Set the source information as the management attribute of the metadata, set the technical information as the technical attribute of the metadata, and set the resource information as the resource archiving attribute of the metadata.

[0081] The following are some property examples:

[0082] Management properties (example):

[0083] Attributed application system: Private customer information management;

[0084] Business management department: Retail Finance Department;

[0085] Technical management department: R & D Department 1.

[0086] Technical properties (example):

[0087] Database information: GoldenDb;

[0088] Schema information: NCIM;

[0089] Table-level information (partial):

[0090] English table name: TBPC1090;

[0091] Chinese table name: Customer financial information;

[0092] Creation time: November 12, 2021;

[0093] Field-level information:

[0094] English field name: CST_ID;

[0095] Chinese field name: Customer number;

[0096] Field type: Fixed-length character type;

[0097] Field length: 18;

[0098] Is null allowed: No;

[0099] Is primary key: Yes;

[0100] Index information:

[0101] Index field.

[0102] Data resource archiving information (example):

[0103] Data resource attribution department: Retail Finance Department;

[0104] Data resource status: Enabled;

[0105] Attributed development platform information:

[0106] Business management department information: Retail Finance Department;

[0107] Technical management department information: R & D Department 1;

[0108] Data management department information: Data governance team.

[0109] S103. According to the technical information and pre-set technical standard parameters, generate standard comparison information based on whether the technical information conforms to the corresponding technical standard parameters, and set the standard comparison information as the data comparison attribute of the metadata; the technical standard parameters include mapping standard names, standard data types, standard data lengths, standard data precisions, and standard code items corresponding to the combination of table names and field names.

[0110] For example, data comparison attribute (example):

[0111] Table name: TBPC1090;

[0112] Field name: CST_ID;

[0113] Mapping standard name: Customer number;

[0114] Data type comparison result: Conform;

[0115] Data length comparison result: Conform;

[0116] Data precision comparison result: None;

[0117] Code item comparison result: None.

[0118] In an embodiment of the present invention, based on step S2, the following will give a possible embodiment to non-restrictively elaborate on its specific implementation.

[0119] S201. Determine the data type of the metadata according to the attribute information.

[0120] In the data processing process, after obtaining the attribute information of the metadata, the primary task is to accurately determine the data type of the metadata. The attribute information contains various characteristic descriptions of the metadata, and these descriptions are the key basis for determining the data type.

[0121] For example, for the metadata describing customer information, its attribute information may include customer name, age, contact information, purchase records, etc. From these attributes, different data types can be analyzed. The customer name belongs to the text type because it is descriptive information composed of characters; the age is usually a numerical type used to represent a specific quantity; the contact information may be a phone number or an email address. Although it is in text form, it has a specific format and use, and can also be regarded as a special text type; while the purchase records may involve different data type combinations such as dates and amounts.

[0122] In actual operation, a series of detailed judgment rules will be formulated. If the attribute information contains characteristics related to numerical values such as clear numerical ranges and measurement units, then this metadata is very likely to belong to the numerical type, such as sales amount, inventory quantity, etc. For attributes that contain character sequences and have no obvious numerical characteristics, if they are used to describe things, events or concepts, they are usually judged as text types, such as product names, project descriptions, etc.

[0123] In addition, for some information with specific formats and meanings, such as information in date-time format (such as "2023 - 10 - 05 14:30:00"), it will be determined as the date-time type; while for information such as "yes" or "no", "true" or "false" representing logical judgments, they are classified as the boolean type. Through these detailed analysis and judgment rules, the data type of the metadata can be accurately determined according to the attribute information.

[0124] S202. According to the correspondence between the data type and the key prefix and the determined data type, assign a key prefix to the metadata.

[0125] After determining the data type of the metadata, the next step is to assign a suitable key prefix to the metadata according to the correspondence between the data type and the key prefix. This correspondence is preset to classify and identify different types of metadata through the key prefix, so as to manage and retrieve data more efficiently later.

[0126] For example, set the key prefix for numerical type data as "num_", the key prefix for text type data as "txt_", the key prefix for date-time type data as "dt_", the key prefix for boolean type data as "bool_", etc.

[0127] Suppose it is determined through step S201 that the data type of a certain piece of metadata is text type, then according to the above correspondence, "txt_" will be assigned as the key prefix for this piece of metadata. If it is a numerical type metadata representing the customer's age, the "num_" key prefix will be assigned.

[0128] This correspondence is not fixed, and enterprises can adjust it flexibly according to their own data management needs and business logics. For example, in a specific business scenario, for numerical type metadata related to financial data, a separate key prefix of "fin_num_" may be set to better distinguish and manage financial-related data. By reasonably setting the correspondence between the data type and the key prefix and strictly assigning the key prefix to the metadata according to this relationship, the classification of data can be ensured to be clear and orderly.

[0129] S203. Allocate a logical range for each key prefix in advance, and store the data type, key prefix, and logical range address of the metadata in a configuration file.

[0130] To store and manage metadata more effectively, it is necessary to allocate a specific logical range for each key prefix in advance. A logical range is a specific area divided in the data storage system for storing metadata with the same key prefix.

[0131] For example, for the key prefix "num_", the logical range [1000 - 1999] may be allocated; for the key prefix "txt_", the logical range [2000 - 2999] is allocated; the logical range corresponding to the key prefix "dt_" may be [5000 - 3999], etc. The division of these logical ranges is considered based on various factors such as the estimation of data volume, storage performance optimization, and convenience of data management.

[0132] After allocating the logical range for each key prefix, it is necessary to store the data type, key prefix, and the corresponding logical range address of the metadata in a configuration file. The configuration file is an important data management tool that records various key information in the data processing process, facilitating the system to quickly obtain and use it in subsequent operations.

[0133] The format of the configuration file can adopt common text formats such as JSON or XML. For example:

[0134] { "metadata_config": [ { "data_type": "numeric type", "key_prefix": "num_", "logical_interval": [1000, 1999]}, { "data_type": "text type", "key_prefix": "txt_", "logical_interval": [2000, 2999]}, { "data_type": "date and time type", "key_prefix": "dt_", "logical_interval": [5000, 3999]} ]}。

[0135] By storing this information in the configuration file, when the system needs to query or process a certain piece of metadata, it can quickly locate the data type, key prefix, and the stored logical range address to which the metadata belongs, greatly improving the efficiency and accuracy of data processing.

[0136] S204. Serialize the metadata and store the processed metadata in the corresponding logical range.

[0137] After completing the previous steps, in order to store the metadata into the corresponding logical range, it is necessary to serialize the metadata first. Serialization is the process of converting a data object into a format that can be transmitted over a network or stored in a file. It converts complex data structures into byte sequences for easy storage and transmission.

[0138] For example, for a metadata object containing multiple attributes, such as customer information metadata (including name, age, contact information, etc.), it exists in a specific object structure in memory. But when storing it on a physical medium (such as a hard disk), it needs to be converted into a format suitable for storage. Common serialization formats include JSON, XML, ProtocolBuffers, etc.

[0139] Suppose JSON format is selected for serialization processing. For a customer information metadata object:

[0140] customer = { "name": "Zhang San", "age": 30, "contact": "138xxxxxxxx"}.

[0141] After serialization processing, it will be converted into the following JSON string:

[0142] {"name": "Zhang San", "age": 30, "contact": "138xxxxxxxx"}.

[0143] After completing the serialization processing, the system will find the corresponding logical range according to the previously determined key prefix and store the processed metadata into that logical range. For example, if the data type of this customer information metadata is text type and the key prefix is “txt_”, then the system will store the serialized metadata into the logical range [2000 - 2999] pre-allocated for the “txt_” key prefix.

[0144] In an embodiment of the present invention, based on step S3, a possible embodiment will be given below to non - restrictively elaborate on its specific implementation.

[0145] This step generates metrics in a batch processing manner. And multiple threads are created, and multiple threads perform synchronous processing on the metadata of multiple logical regions.

[0146] S301. Set the weights of each parameter item of the data benchmarking attribute, and calculate the weighted sum to obtain the data item - level standard compliance.

[0147] The composite index is aggregated from multiple basic indexes such as whether the data types match, whether the data lengths match, whether the data precisions match, and the code item matching degree. For each of these basic indexes, the values of the basic indexes in different library tables in the system need to be considered, and according to whether the data item is required, whether it is an index, whether it is a partition key, etc., the weight ratios in the entire data table are different. Additionally, different weights in the system are set respectively according to the importance degree, data volume, etc. of each table in the application system, and finally, the data standard compliance of the application system is calculated and formed through aggregation.

[0148] Let whether the data types match be with a weight of Let whether the data lengths match be with a weight of Let whether the data precisions match be with a weight of Let the code item matching degree be with a weight of Then the value of each basic index is with a weight value of Then the calculated data item-level data standard compliance is .

[0149] S302. Perform weighted summation according to the preset data item weights and item-level standard compliance to obtain the table-level standard compliance.

[0150] The weights of ordinary data items are formulated with standard weights respectively according to different requirements of the management attributes, business attributes, and technical attributes of the data items. Management attributes such as the affiliated line, the business R & D and operation teams to which it belongs, etc., business attributes such as the data model theme to which it belongs, etc., and technical attributes such as whether the data item is required, an index, or a partition key, etc. The obtained weight of the data item field is Z i Then the calculated table-level data standard compliance of this table is .

[0151] S303. Perform weighted summation according to the preset table weights and the table-level standard compliance to obtain the application-level standard compliance.

[0152] Let the table weight of an ordinary table in the system be Q i If the table is a parameter table, business table, or transaction table, different fixed values are increased. If the table is a statistical table, intermediate table, or flow table, etc., different fixed values are decreased. Then the weight of the table in the application system is Q i Then the calculated application-level data standard compliance of this application system is .

[0153] In an embodiment of the present invention, based on step S4, a possible embodiment will be given below to non - restrictively elaborate on its specific implementation scheme.

[0154] S401. Set a threshold combination for each logical interval, where the threshold combination includes a data item - level standard compliance threshold, a table - level standard compliance threshold, and an application - level standard compliance threshold.

[0155] In the complex system of data management and quality monitoring, in order to accurately identify and process potentially abnormal data, it is necessary to set detailed thresholds for each logical interval. As the basic unit of data storage and management, a logical interval carries metadata of different types and uses.

[0156] For each logical interval, construct a comprehensive threshold combination that covers the standard compliance thresholds at three different levels: data item - level, table - level, and application - level.

[0157] Data item - level standard compliance threshold: A data item is the most basic component of data, such as each field value in a database table. The data item - level standard compliance threshold is used to measure the matching degree of a single data item with a pre - set standard. This standard can be the data format, value range, data type, etc. For example, for a data item representing age, it is set that its value range should be between 0 and 120, and the data type is an integer. Then, the data item - level standard compliance threshold is used to determine whether the age data item is within this reasonable range and whether it meets the integer type requirement. If the degree to which the data item meets these standards is lower than the set threshold, it may mean that there are potential problems with the data. The specific threshold setting will be determined based on the importance of the data, business requirements, and analysis of historical data. For critical business data, a relatively high threshold may be set to ensure data accuracy and reliability; for some auxiliary data, the threshold can be relatively loose.

[0158] Table - level standard compliance threshold: The table - level standard compliance threshold focuses on the overall quality of the entire data table. It not only considers the situation of individual data items but also involves aspects such as the integrity, consistency of the data within the table, and the logical relationships between data. For example, in a customer information table, the unique identifier field (such as customer ID) for each customer should be unique without duplicate values; there may be certain association relationships between various fields, such as the date of birth and age fields should match each other. The table - level standard compliance threshold is used to evaluate the performance of the entire table in these aspects. Multiple factors will be considered to set this threshold, such as the missing rate of data in the table, the deviation degree of data consistency, etc. If the table - level standard compliance is lower than the set threshold, it indicates that there may be significant problems with the data quality of the table and further in - depth analysis and processing are required.

[0159] Application-level standard compliance threshold: The application-level standard compliance threshold is set from a higher level, that is, from the perspective of data in actual applications. It takes into account the availability and effectiveness of data for the entire application system and its support for business processes. For example, in an e-commerce order processing system, the timely update and accuracy of order data are crucial for business processes such as inventory management and delivery arrangements. The application-level standard compliance threshold is used to measure the performance of data in meeting these business requirements. This may involve factors such as the data update frequency and the degree of match between data and business rules. If the application-level standard compliance is below the threshold, it may have a serious impact on the normal operation of the entire application system and the smooth progress of business.

[0160] By carefully setting the threshold combinations at these three levels for each logical interval, it is possible to comprehensively monitor and evaluate data from different perspectives, providing an accurate basis for subsequent anomaly detection and handling.

[0161] S402. Compare the data item-level standard compliance, table-level standard compliance, and application-level standard compliance within the logical interval with the corresponding thresholds. If any of them is below the corresponding threshold, generate the corresponding anomaly warning.

[0162] After setting the thresholds for each logical interval, the next key step is to monitor and evaluate the data within the logical interval in real time. It is necessary to calculate the data item-level standard compliance, table-level standard compliance, and application-level standard compliance within each logical interval respectively, and carefully compare these compliance levels with the corresponding thresholds set previously.

[0163] Comparison of data item-level standard compliance: For each data item within the logical interval, calculate its compliance according to the pre-set data item standards. For example, for a data item in date format, check whether it conforms to the specified date format (such as "YYYY - MM - DD"). If it does, the compliance in terms of format is 100%. Then check whether its value is within a reasonable time range, such as within the business time span allowed by the system. Calculate the overall standard compliance of this data item by integrating these factors. Then, compare this compliance with the data item-level standard compliance threshold. If the compliance of this data item is below the threshold, the system will immediately generate a data item-level anomaly warning. This warning will detail information such as the name of the data item, the name of the table it belongs to, the specific non-compliance situation, and the possible impact on the business, so that relevant personnel can conduct investigations and handling in a timely manner.

[0164] Table-level standard compliance comparison: Based on the compliance assessment at the data item level, a comprehensive assessment of the entire table is further conducted to calculate the table-level standard compliance. This requires considering multiple factors, such as data integrity (whether there are a large number of missing values), data consistency (whether the logical relationships between different fields are correct), and data accuracy (whether the data conforms to the actual situation), etc. Through a series of complex algorithms and rules, the standard compliance of the table is obtained. Then, this compliance is compared with the table-level standard compliance threshold. If the table-level standard compliance is lower than the threshold, the system will generate a table-level anomaly warning. This warning will include information such as the name of the table, the main non-compliant aspects, and the possible impacts on relevant business modules. Relevant personnel can conduct a comprehensive review and repair of the data in the entire table based on this warning to ensure the quality and usability of the data.

[0165] Application-level standard compliance comparison: Finally, from the overall perspective of the application system, the application-level standard compliance is calculated. This requires considering the usage of data in various business processes and the degree of support of the data for business objectives. For example, in a financial statement generation system, it will be checked whether the financial data can generate statements in a timely and accurate manner, and whether the data in the statements conforms to financial regulations and business requirements. Through the evaluation of these aspects, the application-level standard compliance is obtained. It is compared with the application-level standard compliance threshold. If it is lower than the threshold, the system will generate an application-level anomaly warning. This warning will emphasize the potential impacts of the anomaly on the entire application system and business operations, prompting the relevant team to take measures quickly to ensure the normal operation of the business.

[0166] Through this meticulous comparison and warning mechanism, problems existing in data at different levels can be discovered in a timely manner, providing a strong guarantee for the improvement of data quality and the stable development of the business.

[0167] Examples of warning rules are as follows:

[0168] Definition of warning rules: Warn about the deviation degree of data standards in different application systems in the user acceptance test environment.

[0169] Warning level: Intermediate;

[0170] Warning frequency: Every Thursday;

[0171] Processing logic of warning rules: Obtain the deviation degree of data standards in different application systems in the user acceptance test environment. For deviation degrees in the range of -5% to 0%, no warning will be given. For deviation degrees in the range of -5% to -10%, a blue prompt will be given. For deviation degrees in the range of -15% to -10%, a yellow warning will be given. For deviation degrees greater than -15%, a red alarm will be given.

[0172] Supplementary note: The data standard deviation index describes the deviation between the compliance of the application-level data standard formed after each system's calculation and the application system benchmark value. The application system benchmark value is a fixed value set according to the business carried by the application system and the importance of the application system, and is adjusted quarterly. The specific calculation formula is: data standard deviation (ρ) = (application-level data standard compliance - application system benchmark value) / application system benchmark value * 100%.

[0173] S403. Analyze the boundary values of each logical interval using the One-Class SVM model, and determine the abnormal metadata based on the boundary values.

[0174] After generating an abnormal warning through threshold comparison, in order to more accurately identify and locate the abnormal metadata, the One-Class SVM (One-Class Support Vector Machine) model is introduced. One-Class SVM is a powerful machine learning algorithm, especially suitable for finding the boundary of normal data in a dataset, thereby identifying abnormal data that deviates from the normal range.

[0175] The core idea of the One-Class SVM model is to find an optimal hyperplane in a high-dimensional space to enclose most of the normal data within a region, and the boundary of this region is the so-called boundary of normal data. The model determines the distribution range of normal data by maximizing the distance between the normal data and this hyperplane. During the training process, the model only uses normal data for learning, thereby constructing a model that can describe the characteristics of normal data.

[0176] For each logical interval, take the metadata within it as input and analyze it using the One-Class SVM model. The model will find the boundary values of normal data in the feature space according to the characteristics of the input metadata. These boundary values represent the characteristic range of normal metadata within this logical interval. For example, in a logical interval containing customer transaction amounts, the model will analyze the distribution of transaction amounts and determine the upper and lower limits of normal transaction amounts, and these upper and lower limits are the boundary values.

[0177] After obtaining the boundary values of each logical interval, the abnormal metadata can be determined based on these boundary values. For each piece of metadata within the logical interval, compare its characteristics with the boundary values. If the characteristics of the metadata exceed the range defined by the boundary values, then this piece of metadata is determined to be abnormal metadata. For example, if a customer's transaction amount far exceeds the upper limit of the normal transaction amount determined by the One-Class SVM model, then the metadata corresponding to this transaction record will be marked as abnormal.

[0178] By analyzing the boundary values using the One-Class SVM model to determine abnormal metadata, it is possible to more accurately identify those abnormal data points hidden in a large amount of data, providing precise targets for further data processing and problem-solving. This method not only improves the accuracy of anomaly detection but also helps to deeply understand the internal characteristics and distribution laws of the data, providing strong support for optimizing data management strategies.

[0179] In an application scenario, please refer to Figure 2 , metadata management mainly includes three major parts: basic metadata control, enhanced metadata control, and integration of the scheduling platform.

[0180] Among them, basic metadata control includes modules such as metadata attribute definition, metadata collection, metadata catalog, and metadata alignment.

[0181] Metadata attribute definition is to provide a comprehensive and detailed description of metadata. The description of metadata is divided into various attributes, and the process of refining the attribute definitions is carried out. Currently, the various defined attributes of our bank's metadata are mainly divided into several categories of information such as management information, technical information, data standard alignment information, and data resource archiving information. Among them, management information mainly includes the source system to which this type of metadata belongs and the corresponding business management department, technical department, and corresponding responsible persons in that system; technical information includes database information, schema information, table-level information, field-level information, index information, etc.; data standard alignment information mainly includes the detailed result information formed after the consistency comparison and verification of metadata with data standards, including whether the names are consistent, whether the data types are consistent, whether the data lengths are consistent, whether the data precisions are consistent, whether the code items are consistent with the data standards, etc.; data resource archiving information mainly includes the classification, attribution, and other related information of this type of metadata archived to data resources.

[0182] Among them, management information mainly comes from ledger collection and maintenance, technical information mainly comes from the metadata catalog formed after the execution of the metadata collection function in the basic metadata control module, and data standard alignment information comes from the comparison result information after the collected metadata is aligned with the enterprise-level data standards, including whether a certain table and a certain field should be aligned and the alignment results, and the alignment results cover whether the data types match, whether the data lengths and data precisions match, etc.

[0183] Please refer to Figure 3, The improvement of metadata control mainly includes three major modules: metadata indicator management, metadata warning management, and comprehensive independent metadata analysis. Among them, the metadata indicator management function mainly includes functions such as metadata indicator application, metadata indicator maintenance, metadata indicator review, metadata indicator processing, and metadata indicator scheduling; the metadata warning management mainly includes functions such as metadata warning rule maintenance, metadata warning rule review, metadata warning rule processing, metadata warning rule configuration scheduling, and metadata warning signal handling; the comprehensive independent metadata analysis module mainly includes functions such as independent query (report) maintenance, independent query publishing, independent report publishing, large screen data assembly, large screen data publishing, and independent publishing review.

[0184] The metadata indicator application function is to initiate the definition and application process of metadata indicators. Through this function, information such as metadata indicator number, indicator name, indicator classification, indicator introduction, indicator type, indicator business rules, and indicator processing rules is maintained to complete the definition of metadata indicators. The indicator processing rules are carried out by means of independent selection and entry of logical operators. First, select the defined metadata attributes, and then form the processing logic of the metadata indicators through independent combination, arrangement, or logical processing of the formulas.

[0185] The metadata indicator application process is as follows: "Fill in the basic indicator information" -> "Basic attributes", fill in "Indicator name", "Indicator alias", "Indicator classification", "Indicator unit", and then click "Next". "Fill in the basic indicator information" -> "Business attributes", fill in "Business definition", "Business purpose", "Business rules", and then click "Next". "Fill in the basic indicator information" -> "Technical attributes", fill in "Data type", "Processing rules", and then click "Next". "Fill in the basic indicator information" -> "Management attributes", fill in "Contact person of the competent department", "Requirement submitter", and then click "Next". "Fill in the basic indicator information" -> "Data tagging", fill in "Whether it is a Class A indicator", "Whether it meets the data standard", "Whether it is a data standard", "Data standard number", "Data security classification", "Data security level", "Adjustment suggestions", "Data desensitization rules", and then click "Next". "Fill in the basic indicator information" -> "Quality requirements", fill in "Business quality requirements", "Technical quality requirements", "Verification script", and then click "Next".

[0186] The metadata indicator maintenance function is a process of modifying the definitions of metadata indicators that have not been audited or have been audited. After modification, it re-enters the audit process. Through the metadata indicator maintenance function, the metadata indicator name, indicator classification, indicator introduction, indicator type, indicator business rules, and indicator processing rules can be modified, but the metadata indicator number cannot be modified. Metadata indicators that have been incorporated into the metadata indicator processing task cannot be modified.

[0187] The metadata indicator audit function is a function for auditing the initiated metadata indicator applications or maintenance. After the auditor checks various information of the maintained metadata indicators, the audit action is completed, and the status of the audited metadata indicator changes to valid and can be incorporated into the metadata indicator processing task.

[0188] The metadata indicator processing function is a process of incorporating the metadata indicator processing rules into the metadata indicator processing task. Through the execution of the task, finally, the designed and maintained metadata indicators can be generated. A metadata indicator processing task can incorporate multiple metadata indicators.

[0189] An example of the metadata indicator processing task is as follows:

[0190] Task name: Data standard compliance series calculation task;

[0191] Metadata indicators included: Data item-level data standard compliance, table-level data standard compliance, application-level data standard compliance.

[0192] Explanation of the processing logic:

[0193] a. Data item-level data standard compliance:

[0194] Part of the variable definition explanation:

[0195] APPNAME = APPLIST.ALL / / This indicator processing selects the full application list;

[0196] SCHNAME = APPNAME.ALL / / This indicator processing selects all schemas under the application;

[0197] TABNAME = SCHNAME.ALL / / This indicator processing selects all tables under the schema;

[0198] COLNAME = TABNAME.ALL / / This indicator processing selects all columns under the table;

[0199] A1 = Basic indicator. Whether the data types match;

[0200] A2 = Basic indicator. Whether the data lengths match;

[0201] A3 = Basic Index. Whether the data precision is consistent;

[0202] A4 = Basic Index. The consistency of code items;

[0203] A5 = Basic Index. The compliance of data item names;

[0204] P = [50%, 20%, 10%, 15%, 5%] / / Define the weights used for different basic indexes here;

[0205] B1 = Composite Index. The compliance of data standards at the data item level:

[0206] Processing logic part:

[0207] <LOOP COLNAME>;

[0208] B1 = (IF(A1, 1, 0) * P[1] + IF(A2, 1, 0) * P[2] + IF(A3, 1, 0) * P[3] + IF(A4, A4, 0) * P[4] + IF(A5, A5, 0) * P[5]) / (IF(A1, P[1], 0) + IF(A2, P[2], 0) + IF(A3, P[3], 0) + IF(A4, P[4], 0) + IF(A5, P[5], 0));

[0209] <LOOP END>.

[0210] b. The compliance of data standards at the table level:

[0211] Variable definition and description part:

[0212] APPNAME = APPLIST.ALL / / The processing of this index selects the full application list;

[0213] SCHNAME = APPNAME.ALL / / The processing of this index selects all schemas under the application;

[0214] TABNAME = SCHNAME.ALL / / The processing of this index selects all tables under the schema;

[0215] COLP = Composite Index. The importance of data items;

[0216] B1 = Composite Index. The compliance of data standards at the data item level;

[0217] C1 = Composite Index. The compliance of data standards at the table level;

[0218] Processing logic part:

[0219] <LOOP TABNAME>;

[0220] C1C1 = IF(B1, B1, 0) * COLP++;

[0221] C1C2 = IF(B1, COLP, 0)++;

[0222] C1 = C1C1 / C1C2;

[0223] <LOOP END>。

[0224] c. Application - level data standard compliance:

[0225] Variable definition description part:

[0226] APPNAME = APPLIST.ALL / / This indicator processing selects the full application list;

[0227] TABP = Composite indicator. Table importance;

[0228] C1 = Composite indicator. Table - level data standard compliance;

[0229] D1 = Composite indicator. Application - level data standard compliance;

[0230] Processing logic part:

[0231] <LOOP APPNAME>;

[0232] D1D1 = IF(C1, C1, 0) * TABP++;

[0233] D1D2 = IF(C1, TABP, 0)++;

[0234] D1 = D1D1 / D1D2;

[0235] <LOOP END>。

[0236] The metadata indicator scheduling function incorporates information such as the maintenance execution cycle and execution time of the well - maintained metadata indicator processing tasks into the scheduling platform for unified scheduling and execution, and finally generates the metadata indicator processing calculation results. The task scheduling execution results are uniformly queried, displayed, and stored in the database.

[0237] The metadata warning rule maintenance function needs to enter information such as the warning rule number, warning rule name, warning level, warning cycle, and warning environment, and needs to select the already - maintained and used metadata data indicators. After the entry is completed, it is submitted to enter the warning rule review process. Additionally, the metadata warning rule maintenance function can be used to modify the rules that have not yet been incorporated into the metadata warning rule processing tasks and submit them for review again.

[0238] Review the maintained metadata warning rules. After passing the review, the metadata warning rules can be added to the rule processing task for subsequent warning scheduling.

[0239] Through the metadata warning rule processing function, add the reviewed metadata warning rules to the processing task for subsequent scheduling execution and monitoring. For rules with the same warning cycle and warning environment, they can be batch-added to the same warning rule processing task.

[0240] For the maintained metadata warning rule processing task, maintain information such as the execution cycle and execution time and incorporate them into the scheduling platform for unified scheduling execution, finally generating the metadata index processing calculation results. The task scheduling execution results are uniformly queried, displayed, and stored in the database.

[0241] Color-code the warning information generated by the metadata warning rule processing task according to different levels. For yellow-level warnings, send email notifications, and for red-level warnings, send both email and SMS notifications. After the relevant responsible person processes and maintains the signals, reschedule the execution of the metadata warning rule processing task the next day until it is confirmed that the warning signal has been cleared.

[0242] The metadata comprehensive analysis module mainly conducts independent processing and analysis based on relevant information such as the generated metadata definitions, metadata indicators, metadata index processing tasks, metadata warning rules, and metadata warning processing tasks to meet the requirements of actual work.

[0243] In some embodiments, the metadata management system may include multiple functional modules composed of computer program segments. The computer programs of each program segment in the metadata management system can be stored in the memory of the computer device and executed by at least one processor to perform the functions of metadata management (see Figure 1 description).

[0244] In this embodiment, according to the functions it performs, the metadata management system can be divided into multiple functional modules, as Figure 4 shown. The functional modules of system 400 may include: an attribute generation module 410, a data storage module 420, an index calculation module 430, and an exception warning module 440. The modules referred to in the present invention refer to a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0245] The attribute generation module is used to extract the characteristic information of the metadata and generate attribute information based on the characteristic information;

[0246] A data storage module, configured to allocate a key prefix for metadata based on the attribute information, and store the metadata in a pre-allocated logical range of the key prefix in the form of key-value pairs;

[0247] An index calculation module, configured to calculate index values of the attribute information of the metadata in each logical range respectively by using a preset index rule;

[0248] An anomaly warning module, configured to confirm that the index value of a logical range exceeds a set index threshold, and generate a metadata anomaly warning.

[0249] Optionally, as an embodiment of the present invention, the attribute generation module includes:

[0250] A feature extraction unit, configured to extract feature information from an account book and a metadata directory by using an entity naming model, where the feature information includes attribution source information, technical information, and resource information; the attribution source information includes an attribution application system, business management identity information, and technical management identity information; the technical information includes database information, schema information, table-level information, field-level information, and index information; the resource information includes data resource attribution identity information, data resource status, and attribution development platform information;

[0251] A first setting unit, configured to set the attribution source information as the management attribute of the metadata, set the technical information as the technical attribute of the metadata, and set the resource information as the resource filing attribute of the metadata;

[0252] A second setting unit, configured to generate standard comparison information according to whether the technical information conforms to corresponding technical standard parameters according to the technical information and preset technical standard parameters, and set the standard comparison information as the data comparison attribute of the metadata; the technical standard parameters include a mapped standard name, a standard data type, a standard data length, a standard data precision, and a standard code item corresponding to a combination of a table name and a field name.

[0253] Optionally, as an embodiment of the present invention, the data storage module includes:

[0254] A data classification unit, configured to determine the data type of the metadata according to the attribute information;

[0255] An identifier allocation unit, configured to allocate a key prefix for the metadata according to the correspondence between the data type and the key prefix and the determined data type;

[0256] A configuration storage unit, configured to pre-allocate a logical range for each key prefix, and store the data type, key prefix, and logical range address of the metadata in a configuration file;

[0257] The encoded storage unit is used to serialize the metadata and store the processed metadata in the corresponding logical range.

[0258] Figure 5 The metadata management method provided by the embodiments of this application can be applied to devices. Those skilled in the art can understand that the device structure involved in the embodiments of the present invention does not constitute a limitation on the device. The device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. In the embodiments of the present invention, the device includes, but is not limited to, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown in the figure, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the embodiments of this application described and / or claimed herein.

[0259] Among them, the device 500 may include: a processor 510, a memory 520, and a communication unit 530. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation on the present invention. It can be a bus structure, a star structure, or may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0260] Among them, the memory 520 can be used to store the execution instructions of the processor 510. The memory 520 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disks, or optical disks. When the execution instructions in the memory 520 are executed by the processor 510, the device 500 can execute some or all of the steps in the above method embodiments.

[0261] The processor 510 is the control center of the storage device, connecting various parts of the entire electronic device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 520, and invoking data stored in the memory, it performs various functions of the electronic device and / or processes data. The processor may be composed of an integrated circuit (IC), for example, it may be composed of a single packaged IC, or it may be composed of multiple packaged ICs with the same or different functions connected together. For example, the processor 510 may include only a central processing unit (CPU). In the embodiments of the present invention, the CPU may be a single arithmetic core or may include multiple arithmetic cores.

[0262] The communication unit 530 is used to establish a communication channel so that the storage device can communicate with other devices. It receives user data sent by other devices or sends user data to other devices.

[0263] The present invention also provides a computer storage medium. Among them, the computer storage medium can store a program, and when the program is executed, it may include some or all of the steps in the embodiments provided by the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), etc.

[0264] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., which can store program codes, and includes several instructions to enable a computer device (which may be a personal computer, a server, or a second device, a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0265] For the same or similar parts among the various embodiments in this specification, reference can be made to each other. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the descriptions in the method embodiments.

[0266] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the system or module can be in electrical, mechanical or other forms.

[0267] The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0268] In addition, in each embodiment of the present invention, the functional modules can be integrated into a processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.

[0269] Although the present invention has been described in detail by referring to the drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions should all be within the scope of the present invention. / Any person familiar with the technical field of the present invention can easily think of changes or substitutions within the technical scope disclosed by the present invention, and they should all be covered within the protection scope of the present invention.

Claims

1. A metadata management method, characterized in that: include: Extracting feature information of metadata, and generating attribute information based on the feature information; assigning a key prefix to the metadata based on the attribute information, and saving the metadata in a pre-assigned logical interval of the key prefix in the form of a key-value pair; Calculate the index value of the attribute information of the metadata in each logical interval respectively by using the preset index rule; A threshold combination is set for each logical interval respectively, and the threshold combination includes a data item-level standard compliance threshold, a table-level standard compliance threshold, and an application-level standard compliance threshold; the data item-level standard compliance, table-level standard compliance, and application-level standard compliance in the logical interval are compared with the corresponding thresholds, and if they are lower than the corresponding thresholds, a corresponding abnormal warning is generated; the boundary value of each logical interval is analyzed using a One-Class SVM model, and abnormal metadata is determined based on the boundary value; The index values ​​of the attribute information of the metadata in each logical interval are calculated using the preset index rules, including: Setting the weight of each parameter item of the data benchmarking attribute, and calculating the weighted sum to obtain the data item-level standard compliance; the attribute information includes the data benchmarking attribute; Perform weighted summation based on preset data item weights and data item-level standard compliance to obtain table-level standard compliance; A weighted sum is performed according to the preset table weight and the table-level standard compliance to obtain the application-level standard compliance.

2. The method according to claim 1, characterized in that Extracting feature information of metadata and generating attribute information based on the feature information includes: The entity naming model is used to extract characteristic information from the ledger and metadata directory, wherein the characteristic information includes attribution source information, technical information, and resource information; the attribution source information includes attribution application system, business management identity information, and technical management identity information; the technical information includes database information, mode information, table-level information, field-level information, and index information; the resource information includes data resource attribution identity information, data resource status, and attribution development platform information; Set the attribution source information as the management attribute of the metadata, set the technical information as the technical attribute of the metadata, and set the resource information as the resource archiving attribute of the metadata; According to the technical information and pre-set technical standard parameters, standard comparison information is generated according to whether the technical information conforms to the corresponding technical standard parameters, and the standard comparison information is set as a data matching attribute of the metadata; the technical standard parameters include a mapping standard name corresponding to a combination of a table name and a field name, a standard data type, a standard data length, a standard data precision and a standard code item.

3. The method according to claim 1, characterized in that: Allocating a key prefix to the metadata based on the attribute information, and saving the metadata in the form of a key-value pair to a pre-allocated logical interval of the key prefix, including: Determine the data type of the metadata according to the attribute information; assigning a key prefix to the metadata according to a correspondence between the data type and the key prefix and the determined data type; Assign a logical interval to each key prefix in advance, and store the data type, key prefix, and logical interval address of the metadata in a configuration file; The metadata is serialized and stored in the corresponding logical interval.

4. A metadata management system, characterized in that: include: An attribute generation module, used to extract feature information of metadata and generate attribute information based on the feature information; A data storage module, configured to assign a key prefix to the metadata based on the attribute information, and save the metadata in a pre-assigned logical interval of the key prefix in the form of a key-value pair; An index calculation module, used to calculate the index value of the attribute information of the metadata in each logical interval respectively using a preset index rule; The abnormal warning module is used to set a threshold combination for each logical interval, wherein the threshold combination includes a data item-level standard compliance threshold, a table-level standard compliance threshold, and an application-level standard compliance threshold; compare the data item-level standard compliance, table-level standard compliance, and application-level standard compliance in the logical interval with the corresponding thresholds, and generate a corresponding abnormal warning if they are lower than the corresponding thresholds; use the One-Class SVM model to analyze the boundary value of each logical interval, and determine the abnormal metadata based on the boundary value; The index values ​​of the attribute information of the metadata in each logical interval are calculated using the preset index rules, including: Setting the weight of each parameter item of the data benchmarking attribute, and calculating the weighted sum to obtain the data item-level standard compliance; the attribute information includes the data benchmarking attribute; Perform weighted summation based on preset data item weights and data item-level standard compliance to obtain table-level standard compliance; A weighted sum is performed according to the preset table weight and the table-level standard compliance to obtain the application-level standard compliance.

5. The system according to claim 4, characterized in that The attribute generation module includes: A feature extraction unit is used to extract feature information from the ledger and metadata directory using an entity naming model, wherein the feature information includes attribution source information, technical information, and resource information; the attribution source information includes attribution application system, business management identity information, and technical management identity information; the technical information includes database information, mode information, table-level information, field-level information, and index information; the resource information includes data resource attribution identity information, data resource status, and attribution development platform information; A first setting unit is used to set the attribution source information as the management attribute of the metadata, set the technical information as the technical attribute of the metadata, and set the resource information as the resource archiving attribute of the metadata; The second setting unit is used to generate standard comparison information according to the technical information and pre-set technical standard parameters, and whether the technical information conforms to the corresponding technical standard parameters, and set the standard comparison information as a data matching attribute of the metadata; the technical standard parameters include a mapping standard name corresponding to a combination of a table name and a field name, a standard data type, a standard data length, a standard data precision and a standard code item.

6. The system according to claim 4, characterized in that The data storage module comprises: A data classification unit, used to determine the data type of the metadata according to the attribute information; an identification allocation unit, configured to allocate a key prefix to the metadata according to a correspondence between the data type and the key prefix and the determined data type; A configuration storage unit is used to pre-allocate a logical interval for each key prefix, and store the data type, key prefix and logical interval address of the metadata in a configuration file; The encoding storage unit is used to serialize the metadata and store the processed metadata in a corresponding logical interval.

7. A metadata management device, characterized in that: include: A memory for storing a metadata management program; A processor, configured to implement the steps of the metadata management method as described in any one of claims 1 to 3 when executing the metadata management program.

8. A computer-readable storage medium storing a computer program, characterized in that: The readable storage medium stores a metadata management program, and when the metadata management program is executed by a processor, the steps of the metadata management method according to any one of claims 1 to 3 are implemented.

Citation Information

Patent Citations

  • Visual data processing method, device and equipment and readable storage medium

    CN117112860A