Metadata processing method and apparatus

CN116126875BActive Publication Date: 2026-04-28MASHANG CONSUMER FINANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
MASHANG CONSUMER FINANCE CO LTD
Filing Date
2022-12-28
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing technologies, the determination of metadata reuse values ​​relies too heavily on the work experience of operations and maintenance personnel, resulting in insufficient accuracy in data management.

Method used

By acquiring the metadata generated during the execution of the target business, we can determine the data correlation, business cost, and data application frequency. Using these parameters, we can calculate the target reuse value to decide whether to store or delete the metadata.

Benefits of technology

This improves the accuracy of determining reusable metadata values, thereby improving the accuracy of data management and reducing errors and wasted storage space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116126875B_ABST
    Figure CN116126875B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a metadata processing method and device, the method comprises the following steps: obtaining metadata generated in an execution process of a target service; determining at least one of data correlation degree, service cost value and data application frequency according to the metadata; determining a target reuse value according to the at least one of the data correlation degree, the service cost value and the data application frequency, wherein the target reuse value represents a possibility size of metadata reuse; and storing or deleting the metadata according to the target reuse value. The present application improves the accuracy of the target reuse value determination, and further improves the accuracy of data management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a metadata processing method and apparatus. Background Technology

[0002] With the development of network technology, the amount of data has exploded, and more and more people are realizing the importance of data.

[0003] As the amount of data increases, in order to better manage the data, we can first determine the reuse value of the data, and then manage the data according to the reuse value.

[0004] However, when determining data reuse values, the data reuse values ​​are usually set manually by operations and maintenance personnel. This over-reliance on the work experience of operations and maintenance personnel reduces the accuracy of data reuse value determination, which in turn affects the accuracy of data management. Summary of the Invention

[0005] This application provides a metadata processing method and apparatus to improve the accuracy of data management.

[0006] In a first aspect, embodiments of this application provide a metadata processing method, including:

[0007] Obtain metadata generated during the execution of the target business process;

[0008] Based on the metadata, determine at least one of the following: data correlation degree, business cost value, and data application frequency;

[0009] A target reuse value is determined based on at least one of the data correlation degree, the business cost value, and the data application frequency, wherein the target reuse value represents the likelihood that the metadata will be reused;

[0010] The metadata can be stored or deleted based on the target reuse value.

[0011] Optionally, determining the data correlation degree based on the metadata includes:

[0012] Determine the number of target groups corresponding to each sub-service involved in the metadata, wherein the target service includes at least one sub-service, each sub-service involves at least one target group, and each target group performs different operations on the corresponding sub-service;

[0013] The data association value of each sub-service is determined based on the number of target groups corresponding to each sub-service and a preset number threshold;

[0014] The data association value of each sub-service is summed to obtain the data association degree.

[0015] Optionally, determining the business cost value based on the metadata includes:

[0016] The storage cost and hardware consumption cost are determined based on the metadata.

[0017] The storage cost and the hardware consumption cost are summed to obtain the business cost value.

[0018] Optionally, determining the storage cost value based on the metadata includes:

[0019] The storage consumption of the target service is determined based on the metadata.

[0020] Determine the unit price of storage costs;

[0021] The storage cost value is determined based on the storage consumption of the target service and the unit price of the storage cost.

[0022] Optionally, determining the storage consumption of the target service based on the metadata includes:

[0023] The target data table contained in the metadata is determined by a preset storage path, wherein the target business contains at least one sub-business, and the preset storage path is the storage path for the data corresponding to each sub-business.

[0024] Determine the initial storage amount for each target data table, and sum the initial storage amounts for each target data table to obtain the storage amount consumed by the target service.

[0025] Optionally, determining the unit price of storage cost includes:

[0026] Obtain the basic cost value within a preset time period, wherein the basic cost value is at least one of the following within the preset time period: host depreciation cost, data center rental cost, network facility depreciation cost, and personnel operation and maintenance cost;

[0027] The first ratio of storage costs is determined based on the procurement costs of the processor, memory, and disk.

[0028] The unit price of storage cost is determined based on the basic cost value, the first ratio, and the total storage amount, wherein the storage amount consumed by the target service is the storage consumed by the target service every day, and the total storage amount is the total storage consumed by the target service within a preset time period.

[0029] Optionally, determining the hardware consumption cost based on the metadata includes:

[0030] The total number of processors and total memory used during the operation of the target service are determined based on the metadata.

[0031] Determine the unit cost of the processor and the unit cost of the memory;

[0032] The processor cost is determined based on the total amount of processors consumed by the target service and the unit price of the processors; the memory cost is determined based on the total amount of memory consumed by the target service and the unit price of the memory.

[0033] The hardware consumption cost is determined based on the processor cost and the memory cost.

[0034] Optionally, determining the data application frequency based on the metadata includes:

[0035] Determine the number of creators, the number of users, the number of downstream tables associated with the metadata, and the number of upstream tables contained in the metadata;

[0036] The frequency of data application is determined based on the number of creators, the number of users, the number of downstream tables associated with the metadata, and the number of upstream tables.

[0037] Optionally, deleting the metadata based on the target reuse value includes:

[0038] If the target reuse value is lower than the preset reuse value threshold, a metadata deletion prompt will be generated and displayed.

[0039] In response to a touch operation that prompts for metadata deletion, the metadata corresponding to the target service is deleted.

[0040] Secondly, embodiments of this application provide a metadata processing apparatus, including:

[0041] The acquisition module is used to acquire metadata generated during the execution of the target business.

[0042] The processing module is used to determine at least one of data correlation degree, business cost value and data application frequency based on the metadata;

[0043] The processing module is further configured to determine a target reuse value based on at least one of the data correlation degree, the business cost value, and the data application frequency, wherein the target reuse value represents the likelihood of the metadata being reused;

[0044] The processing module is also used to store or delete the metadata based on the target reuse value.

[0045] Thirdly, embodiments of this application provide an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0046] The memory stores computer-executed instructions;

[0047] The processor executes computer execution instructions stored in the memory to implement the metadata processing method as described in any of the first aspects.

[0048] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the metadata processing method as described in any of the first aspects.

[0049] This application provides a metadata processing method and apparatus. By adopting the above scheme, metadata generated during the execution of a target business can be obtained first. Then, at least one of data relevance, business cost value, and data application frequency can be determined based on the metadata. Next, a target reuse value is determined based on at least one of these parameters. Metadata is then stored or deleted based on the target reuse value. By first determining parameters such as data relevance, business cost value, and data application frequency based on the metadata generated during the execution of the target business, and then determining the target reuse value based on these parameters, the determined target reuse value can reflect the implementation status of the target business from multiple dimensions, rather than solely relying on the work experience of maintenance personnel. This improves the accuracy of target reuse value determination, thereby improving the accuracy of data management. Attached Figure Description

[0050] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 A schematic diagram of the application system architecture for the data processing method provided in the embodiments of this application;

[0052] Figure 2 A flowchart illustrating the data processing method provided in an embodiment of this application;

[0053] Figure 3 This is a schematic diagram illustrating the application of the matrix diagram corresponding to the first sub-database provided in the embodiments of this application.

[0054] Figure 4 This is an application diagram of the matrix diagram corresponding to the first sub-database provided in another embodiment of this application;

[0055] Figure 5 This is a schematic diagram of the structure of the data processing apparatus provided in the embodiments of this application;

[0056] Figure 6 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0057] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0058] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can also include other sequential examples besides those illustrated or described. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0059] In related technologies, as more and more businesses can be implemented online, the metadata (also known as intermediary data or relay data, which describes data, mainly information describing data attributes, used to support functions such as indicating storage location, historical data, resource lookup, and file records) generated during the business implementation process is also increasing. To better manage metadata, we can first determine the reusability value of the metadata, and then manage the metadata based on the reusability value. For example, we can determine whether metadata can be reused through the reusability value, and then further process the data based on the determination result. However, there is no systematic evaluation method for determining whether metadata can be reused; generally, operations and maintenance personnel determine this based on their work experience, which is highly subjective and reduces the accuracy of determining data reusability values, thus affecting the accuracy of data management.

[0060] Based on the aforementioned technical issues, this application first determines parameters such as data relevance, business cost, and data application frequency based on the metadata generated during the execution of the target business. Then, it determines the target reuse value based on these parameters. This approach allows the determined target reuse value to reflect the implementation status of the target business from multiple dimensions, rather than solely relying on the work experience of operations and maintenance personnel. This achieves the technical effect of improving the accuracy of target reuse value determination, thereby enhancing the accuracy of data management.

[0061] Figure 1 This is a schematic diagram of the application system architecture for the metadata processing method provided in the embodiments of this application, such as... Figure 1 As shown, the application system may include a database and an electronic device. The database stores metadata generated during the execution of different target services (for example, transfer services, login services, account registration services, etc.). The metadata may be data describing service attribute information. The electronic device can obtain the metadata corresponding to the target service from the database, and then determine at least one of data correlation, business cost value, and data application frequency based on the metadata. It then determines a target reuse value based on at least one of these factors, and further operates on the metadata based on the determined target reuse value.

[0062] Optionally, if the target reuse value exceeds the preset reuse value threshold, it indicates that the metadata has a high degree of reusability, meaning the metadata has high data value. Therefore, the metadata can be retained in the database or backed up to another database for future reuse. If the target reuse value is above or below the reuse value threshold, it indicates that the metadata has a low degree of reusability, meaning the metadata has low data value. Therefore, the metadata can be deleted directly.

[0063] Among them, electronic devices can be a single server or a server cluster.

[0064] The technical solutions of this application will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0065] Figure 2 This is a flowchart illustrating a metadata processing method provided in an embodiment of this application. The method in this embodiment can be executed by an electronic device. Figure 2 As shown, the method in this embodiment may include:

[0066] S201: Obtain metadata generated during the execution of the target business.

[0067] In this embodiment, when determining the implementation status of a target service, the metadata generated during the execution of the target service can be obtained first, and then the implementation status of the target service can be represented based on the metadata. The target service can be a transfer service, an account registration service, or an account login service, etc. Taking a transfer service as an example, it can include four sub-services: determining the transfer amount, determining the sender's account, determining the recipient's account, and performing the transfer based on the determined transfer amount, sender's account, and recipient's account. Different sub-services correspond to different metadata.

[0068] In addition, the metadata corresponding to the target business can be stored in a data warehouse, and when retrieving the metadata corresponding to the target business, it can be directly retrieved from the data warehouse.

[0069] Furthermore, there can be various types of data warehouses. For example, a data warehouse can be Hive, HBase, or Kafka, etc., and the data warehouse can contain metadata such as database table names and storage locations. It can also include metadata related to the target business implementation process, such as the processing procedures involved in the target business implementation process, the target groups involved, the input and output of data corresponding to the processing procedures, and the creators and users of the data.

[0070] Optionally, if the data warehouse is Hive, the Hive-metastore interface or the hive-metastore-listener can be used to collect metadata in Hive in batches or in real time. Alternatively, metadata can be obtained through hive-hook or manual batch import.

[0071] Optionally, if the data warehouse is HBase, metadata can be collected by analyzing the ZooKeeper data in HBase, or by using the Admin-related APIs in HBase, or it can be obtained manually.

[0072] Optionally, if the data warehouse is Kafka, metadata can be collected through the management-related APIs provided by Kafka, or it can be obtained manually.

[0073] In addition, metadata can be collected through the relevant interfaces provided by various business departments.

[0074] S202: Determine at least one of the following based on metadata: data relevance, business cost value, and data application frequency.

[0075] In this embodiment, after obtaining the metadata, it can be further processed to represent the implementation status of the target business from different dimensions. Furthermore, the implementation status of the target business can be represented by at least one of data correlation, business cost value, and data application frequency.

[0076] Optionally, data correlation can represent the relationship between each sub-business within the target business and the target group. Correspondingly, the more target groups a sub-business corresponds to, the higher the correlation between the data; conversely, the fewer target groups a sub-business corresponds to, the lower the correlation between the data.

[0077] Optionally, the business cost value can represent the cost incurred in the implementation of the target business, which can include storage costs and hardware consumption costs. Hardware consumption costs can include processor costs and memory costs.

[0078] Optionally, data application frequency can represent how often data is applied. A higher data application frequency indicates that the data is applied more frequently, and vice versa. Specifically, data application frequency can be related to the number of creators, users, downstream tables associated with the metadata, and upstream tables contained in the metadata.

[0079] S203: Determine the target reuse value based on at least one of data relevance, business cost value, and data application frequency, wherein the target reuse value represents the likelihood of metadata being reused.

[0080] In this embodiment, after determining at least one of data relevance, business cost value, and data application frequency, a target reuse value representing the likelihood of metadata being reused can be determined based on at least one of these factors. This target reuse value can be used to analyze the implementation of target business processes, analyze data flow routes, or serve as a reference for data archiving and destruction.

[0081] Furthermore, when determining the target reuse value based on at least one of data relevance, business cost, and data application frequency, the target reuse value can be obtained by summing these three factors. Alternatively, different weight values ​​can be assigned to data relevance, business cost, and data application frequency, and then a weighted sum can be performed based on these weight values ​​to obtain the target reuse value. Additionally, any two of these factors can be summed. Furthermore, if only one of data relevance, business cost, or data application frequency is determined, then that factor can be directly used as the target reuse value.

[0082] S203: Store or delete metadata based on the target reuse value.

[0083] In this embodiment, after obtaining the target reuse value, it can be determined whether metadata needs to be retained based on the target reuse value.

[0084] Furthermore, the deletion of the metadata based on the target reuse value may specifically include:

[0085] If the target reuse value is lower than the preset reuse value threshold, a metadata deletion prompt will be generated and displayed.

[0086] In response to a touch operation that prompts for metadata deletion, the metadata corresponding to the target service is deleted.

[0087] Specifically, if the target reuse value is lower than the preset reuse value threshold, it indicates that the metadata has low reuse value, and the metadata can be directly deleted to save storage space. Alternatively, when the target reuse value is determined to be lower than the preset reuse value threshold, a metadata deletion prompt can be generated and displayed first. After receiving a touch operation that triggers the metadata deletion prompt, the metadata corresponding to the target business is then deleted, improving metadata security and reducing the possibility of accidental operations.

[0088] Furthermore, if the target reuse value is higher than or equal to the reuse value threshold, the metadata can be retained. The reuse value threshold can be customized according to the actual application scenario, and will not be discussed in detail here.

[0089] By adopting the above solution, we can first obtain the metadata generated during the execution of the target business. Then, based on the metadata, we can determine at least one of the following: data relevance, business cost value, and data application frequency. Next, we can determine the target reuse value based on at least one of these parameters. Then, we can store or delete the metadata based on the target reuse value. By first determining parameters such as data relevance, business cost value, and data application frequency based on the metadata generated during the execution of the target business, and then determining the target reuse value based on these parameters, the determined target reuse value can reflect the implementation status of the target business from multiple dimensions, rather than relying solely on the work experience of operations and maintenance personnel. This improves the accuracy of determining the target reuse value, thereby improving the accuracy of data management.

[0090] based on Figure 2 In addition to the method described herein, this specification also provides some specific implementation schemes of the method, which will be described below.

[0091] In another embodiment, determining the data correlation degree based on the metadata may specifically include:

[0092] Determine the number of target groups corresponding to each sub-service involved in the metadata, wherein the target service includes at least one sub-service, each sub-service involves at least one target group, and each target group performs different operations on the corresponding sub-service.

[0093] The data association value of each sub-service is determined based on the number of target groups corresponding to each sub-service and a preset number threshold.

[0094] The data association value of each sub-service is summed to obtain the data association degree.

[0095] In this embodiment, after obtaining the metadata, due to the large amount of data within it, the metadata can be processed to obtain parameters such as data correlation, business cost value, and data application frequency. Different parameters can represent the implementation status of the target business from different dimensions. Furthermore, parameters such as data correlation, business cost value, and data application frequency can have multiple representations. In one possible implementation, these parameters can be in the form of a data table, thus saving storage space. In another possible implementation, these parameters can also be in the form of a matrix diagram. The matrix diagram can be displayed directly on a display device, or it can be displayed directly on the device via a data query request. This matrix diagram provides a clear understanding of the specific situation during the implementation of the target business, offering a reliable basis for subsequent data operations and improving the accuracy of data processing.

[0096] Furthermore, when determining data correlation, one can first identify the sub-businesses involved in the metadata, then determine the number of target groups corresponding to each sub-business (i.e., the number of processors involved in the sub-business processing). Then, one can compare the relationship between the number of target groups corresponding to each sub-business and a preset threshold. If the number of target groups corresponding to a sub-business exceeds the preset threshold, the data correlation value for that sub-business is determined to be 1; otherwise, it is 0. The data correlation values ​​for each sub-business can then be summed to obtain the data correlation degree. Alternatively, after determining the number of target groups corresponding to each sub-business, this number can be used as the data correlation value for each sub-business, and the summation of these values ​​can then be performed to obtain the data correlation degree.

[0097] In addition, data correlation can also be presented in the form of a matrix diagram, for example, a PO (Process and Organize) matrix diagram.

[0098] For example, Table 1 is a PO matrix diagram. In this matrix diagram, the target business can include four sub-businesses, each corresponding to a target group. Data correlation = Data correlation value of user-entered production information (i.e., 1) + Data correlation value of risk control approval verification (i.e., 1*1) + Data correlation value of data stored in the business system (i.e., 1) + Data correlation value of collected metadata to the data warehouse (i.e., 1) = 4. A higher data correlation indicates more processing steps or more departments involved in data flow, and a greater correlation between the data.

[0099] Table 1. PO Matrix Diagram

[0100]

[0101] In another embodiment, determining the business cost value based on the metadata may specifically include:

[0102] The storage cost and hardware consumption cost are determined based on the metadata.

[0103] The storage cost and the hardware consumption cost are summed to obtain the business cost value.

[0104] In this embodiment, after obtaining the metadata, the business cost value can be determined based on the metadata. The business cost value may include storage costs and hardware consumption costs.

[0105] Furthermore, determining the storage cost value based on the metadata may specifically include:

[0106] The storage consumption of the target service is determined based on the metadata.

[0107] Determine the unit price of storage costs.

[0108] The storage cost value is determined based on the storage consumption of the target service and the unit price of the storage cost.

[0109] Specifically, after the metadata is generated, it needs to be stored. During the storage process, in order to better evaluate the metadata, the storage cost value corresponding to the metadata can be determined, that is, the cost consumed in storing the metadata.

[0110] In addition, when determining the storage cost value, the amount of storage consumed by the metadata can be determined first, then the unit price of storage cost can be determined, and finally the storage cost value can be determined based on the amount of storage consumed and the unit price of storage cost.

[0111] Furthermore, determining the storage consumption of the target service based on the metadata may specifically include:

[0112] The target data table contained in the metadata is determined by a preset storage path, wherein the target business contains at least one sub-business, and the preset storage path is the storage path for the data corresponding to each sub-business.

[0113] Determine the initial storage amount for each target data table, and sum the initial storage amounts for each target data table to obtain the storage amount consumed by the target service.

[0114] Specifically, when determining the storage consumption of a target service, one can first identify the target data tables contained in the metadata, then determine the initial storage consumption of each target data table, and finally sum the initial storage consumption of each target data table to obtain the storage consumption of the target service.

[0115] Optionally, when determining the target data table contained in the metadata, the target data table can be determined through a pre-defined storage path. For example, the target data table can be an SQL table, where the grouping field in the SQL can be a path, and the table prefix portion of the path can be extracted as the defined storage path. For instance, for the file " / Hive unified prefix / database name / table name / file name" in the table, " / Hive unified prefix / database name (department) / table (personal) name" can be used as a group to calculate the initial storage volume (T*days) of the target data table. Furthermore, after the tablespace analysis is completed, the initial storage volume of each target data table can be written to the table storage space record table in MySQL. Alternatively, the initial storage volume of each target data table can be summed to obtain the storage volume consumed by the target business, and this storage volume can be stored in the table storage space record table. Additionally, the cost consumption of tables, databases, responsible persons, and departments can be calculated based on dimensions such as the database to which the table belongs, the responsible person, and the department. For departments, responsible persons, databases, and tables with high cost consumption, conscious optimization can be guided.

[0116] Furthermore, determining the unit price of storage costs may specifically include:

[0117] Obtain the basic cost value within a preset time period, wherein the basic cost value is at least one of the following within the preset time period: host depreciation cost, data center rental cost, network facility depreciation cost, and personnel operation and maintenance cost.

[0118] The first ratio of storage costs is determined based on the procurement costs of the processor, memory, and disk.

[0119] The unit price of storage cost is determined based on the basic cost value, the first ratio, and the total storage amount, wherein the storage amount consumed by the target service is the storage consumed by the target service every day, and the total storage amount is the total storage consumed by the target service within a preset time period.

[0120] Specifically, after determining the storage volume consumed by the target business, the unit storage cost can be determined. Then, the storage volume consumed by the target business and the unit storage cost can be multiplied to obtain the storage cost value. When determining the unit storage cost, a base cost value within a preset time period can be obtained first. Then, based on the purchase costs of the processor, memory, and disk, the first ratio of storage cost can be determined. Finally, the unit storage cost is determined based on the base cost value, the first ratio, and the total storage volume. Optionally, the unit storage cost can be determined by (base cost value * first ratio) / total storage volume, where the total storage volume is the total storage consumed by the target business within the preset time period. The preset time period can be customized according to the actual application scenario; for example, the preset time period can be one month. Alternatively, the first ratio of storage cost can be determined by the disk purchase cost / (processor purchase cost + memory purchase cost + disk purchase cost).

[0121] Furthermore, determining the hardware consumption cost based on the metadata may specifically include:

[0122] The total number of processors and memory used in the operation of the target service are determined based on the metadata.

[0123] Determine the unit cost of the processor and the unit cost of the memory.

[0124] The processor cost is determined based on the total amount of processors consumed by the target service and the unit cost of the processors, and the memory cost is determined based on the total amount of memory consumed by the target service and the unit cost of the memory.

[0125] The hardware consumption cost is determined based on the processor cost and the memory cost.

[0126] Specifically, hardware costs can be determined based on metadata, including processor and memory costs. Optionally, processor costs can be determined by multiplying the total number of processors consumed by the target service by the unit price of the processor. Memory costs can be determined by multiplying the total amount of memory consumed by the target service by the unit price of the memory.

[0127] Furthermore, the total number of processors can be the total number of processors (cores * hours) consumed by the target service within a preset time period (for example, one day), and the total number of memory can be the total number of memory (G * hours) consumed by the target service within a preset time period (for example, one day). For example, the total number of processors and the total number of memory can be the total number of processors and the total number of memory consumed by all applications (i.e., the target service within the applications) during execution in the system.

[0128] Furthermore, the unit cost of a processor can be determined by dividing the total processor consumption by (base cost * second ratio). Similarly, the unit cost of memory can be determined by dividing the total memory consumption by (base cost * third ratio). The total processor consumption can be the total number of processors consumed in a month, and the total memory consumption can be the total number of memory units consumed in a month. The second ratio can be determined by dividing the memory purchase cost by (processor purchase cost + memory purchase cost + disk purchase cost), and the third ratio can be determined by dividing the processor purchase cost by (processor purchase cost + memory purchase cost + disk purchase cost).

[0129] In addition, business cost values ​​can be in the form of a matrix diagram. For example, it can be an RD (Resource and Data) matrix diagram, which can evaluate the production cost or usage cost of data based on the resources involved in the implementation of the target business.

[0130] For example, Table 2 is an RD matrix diagram. In this matrix diagram, the target business can include five sub-businesses, and each sub-business can correspond to a data table, namely data table A, data table B, data table C, data table D, and data table E. Each data table can contain three types of costs: storage cost, processor cost, and memory cost.

[0131] Table 2 RD Matrix Diagram

[0132]

[0133] In another embodiment, determining the data application frequency based on the metadata may specifically include:

[0134] Determine the number of creators, users, downstream tables associated with the metadata, and upstream tables contained in the metadata.

[0135] The frequency of data application is determined based on the number of creators, the number of users, the number of downstream tables associated with the metadata, and the number of upstream tables.

[0136] In this embodiment, data application frequency can represent the frequency of data application. The higher the application frequency, the more times the service is applied, which indicates that the service is more important or that the service has a wider application scope.

[0137] In addition, it is also possible to determine the number of creators (for example, an administrator account through which metadata can be added, deleted, or modified), the number of users (for example, an administrator account through which metadata can be queried or applied), the number of associated downstream tables (i.e., the data tables corresponding to the business that can only be executed after the target business is completed), and the number of associated upstream tables (i.e., the data tables corresponding to the business that is executed before the target business is executed).

[0138] Optionally, the data application frequency can be presented as a matrix diagram. This matrix diagram can visually represent the number of users and creators included in the metadata, as well as the number of downstream and upstream tables associated with the metadata table. Through a query request, the matrix diagram corresponding to the data application frequency can be directly displayed on the interface, enabling quick and intuitive identification of metadata tables with frequent usage or numerous creators (different metadata tables can correspond to different sub-businesses). This allows for the identification of sub-businesses with high application frequency, which can then be adjusted based on their application frequency within the target business. Furthermore, if the application frequency of a sub-business falls below a preset threshold within a certain time period, it indicates that the sub-business may be unnecessary, infrequently used, or have problematic processing logic for the implementation of the target business. Therefore, sub-businesses with application frequencies below the threshold within the preset time period can be adjusted to simplify the implementation process of the target business and improve its efficiency.

[0139] The matrix diagram corresponding to data application frequency can also be called a CU (Create and Use) matrix diagram. It can be used to mark two important attributes associated with a target business: the creator and user of the table. For example, the creator of Hive table data, and the users of Hive table data. There can be multiple creators and multiple users; creators and users can be the same or different. Different display styles can be used to show creators and users, or to display the number of creators and users (for example, creators can be represented in red, and users in blue, with deeper red and deeper blue as the number of creators and users increases). This allows operations personnel to intuitively understand the specific details of the metadata. Furthermore, the data application frequency can be obtained by summing the number of creators, the number of users, the number of downstream tables associated with the metadata, and the number of upstream tables.

[0140] For example, Figure 3 This is an application diagram illustrating the matrix graph corresponding to the data application frequency provided in the embodiments of this application, such as... Figure 3 As shown in this embodiment, there are six sub-data tables corresponding to sub-services. Each sub-data table can be a rectangular block in a matrix diagram. Each rectangular block can be divided into two parts, representing the user and creator corresponding to each sub-service, respectively. The user and creator can be represented using different styles. Alternatively, the user and creator corresponding to each sub-service can be represented numerically. For example, in the first rectangular block, the sub-service has 1 user and 3 creators.

[0141] also, Figure 4 This is an application illustration of a matrix diagram corresponding to the frequency of data applications provided in another embodiment of this application, such as... Figure 4 As shown, in this embodiment, in Figure 3 Based on the aforementioned embodiments, upstream and downstream tables associated with each sub-service can also be displayed. Optionally, when a touch operation is applied to the rectangular area corresponding to a sub-service (e.g., when a mouse or user's finger touches the rectangular area corresponding to a sub-service), the specific details of the upstream and downstream tables associated with the sub-service can be displayed. For example, when the user's finger touches the area of ​​the last rectangular area, if the sub-service corresponding to that rectangular area has 3 upstream tables and 2 downstream tables, the specific information of the tables associated with that sub-service can be displayed. Furthermore, table names can also be displayed, such as the last table being named Table F. Further, the specific information of the user and creator can also be displayed.

[0142] In another embodiment, after storing the metadata according to the target reuse value, the process may specifically include:

[0143] The metadata is labeled to obtain a training sample set.

[0144] The network model is trained based on the training sample set to obtain the target business processing model.

[0145] In this embodiment, if the target reuse value is higher than a preset reuse value threshold, it indicates that the data reuse value is relatively high, and the metadata can be stored for subsequent reuse. Correspondingly, the metadata can be labeled to obtain a training sample set. Then, the network model can be trained using the training sample set to obtain a target business processing model, which can automatically implement the target business. For example, the target business could be image reprocessing, such as adding special effects. The labeled metadata can be used to automatically train the network model, obtaining an image processing model capable of adding special effects to images. Then, special effects can be added to images using this image processing model.

[0146] In summary, by reprocessing the stored metadata, the utilization value of the metadata is improved, and accurate and rich training samples are provided for model training, thereby improving the efficiency and accuracy of model training.

[0147] Based on the same idea, this specification also provides an apparatus corresponding to the above method. Figure 5 This is a schematic diagram of the structure of the metadata processing device provided in the embodiments of this application, such as... Figure 5 As shown, the apparatus provided in this embodiment may include:

[0148] The acquisition module 501 is used to acquire metadata generated during the execution of the target business.

[0149] Processing module 502 is used to determine at least one of data correlation degree, business cost value and data application frequency based on the metadata.

[0150] The processing module 502 is further configured to determine a target reuse value based on at least one of the data correlation degree, the business cost value, and the data application frequency, wherein the target reuse value represents the likelihood of the metadata being reused.

[0151] The processing module 502 is also used to store or delete the metadata according to the target reuse value.

[0152] In this embodiment, the processing module 502 is further configured to:

[0153] If the target reuse value is lower than the preset reuse value threshold, a metadata deletion prompt will be generated and displayed.

[0154] In response to a touch operation that prompts for metadata deletion, the metadata corresponding to the target service is deleted.

[0155] In another embodiment, the processing module 502 is further configured to:

[0156] Determine the number of target groups corresponding to each sub-service involved in the metadata, wherein the target service includes at least one sub-service, each sub-service involves at least one target group, and each target group performs different operations on the corresponding sub-service.

[0157] The data association value of each sub-service is determined based on the number of target groups corresponding to each sub-service and a preset number threshold.

[0158] The data association value of each sub-service is summed to obtain the data association degree.

[0159] In another embodiment, the processing module 502 is further configured to:

[0160] The storage cost and hardware consumption cost are determined based on the metadata.

[0161] The storage cost and the hardware consumption cost are summed to obtain the business cost value.

[0162] Furthermore, the processing module 502 is also used for:

[0163] The storage consumption of the target service is determined based on the metadata.

[0164] Determine the unit price of storage costs.

[0165] The storage cost value is determined based on the storage consumption of the target service and the unit price of the storage cost.

[0166] Furthermore, the processing module 502 is also used for:

[0167] The target data table contained in the metadata is determined by a preset storage path, wherein the target business contains at least one sub-business, and the preset storage path is the storage path for the data corresponding to each sub-business.

[0168] Determine the initial storage amount for each target data table, and sum the initial storage amounts for each target data table to obtain the storage amount consumed by the target service.

[0169] Furthermore, the processing module 502 is also used for:

[0170] Obtain the basic cost value within a preset time period, wherein the basic cost value is at least one of the following within the preset time period: host depreciation cost, data center rental cost, network facility depreciation cost, and personnel operation and maintenance cost.

[0171] The first ratio of storage costs is determined based on the procurement costs of the processor, memory, and disk.

[0172] The unit price of storage cost is determined based on the basic cost value, the first ratio, and the total storage amount, wherein the storage amount consumed by the target service is the storage consumed by the target service every day, and the total storage amount is the total storage consumed by the target service within a preset time period.

[0173] In addition, the processing module 502 is also used for:

[0174] The total number of processors and memory used in the operation of the target service are determined based on the metadata.

[0175] Determine the unit cost of the processor and the unit cost of the memory.

[0176] The processor cost is determined based on the total amount of processors consumed by the target service and the unit cost of the processors, and the memory cost is determined based on the total amount of memory consumed by the target service and the unit cost of the memory.

[0177] The hardware consumption cost is determined based on the processor cost and the memory cost.

[0178] In another embodiment, the processing module 502 is further configured to:

[0179] Determine the number of creators, users, downstream tables associated with the metadata, and upstream tables contained in the metadata.

[0180] The frequency of data application is determined based on the number of creators, the number of users, the number of downstream tables associated with the metadata, and the number of upstream tables.

[0181] In another embodiment, the processing module 502 is further configured to:

[0182] The metadata is labeled to obtain a training sample set.

[0183] The network model is trained based on the training sample set to obtain the target business processing model.

[0184] The apparatus provided in this application embodiment can achieve the above-mentioned... Figure 2 The methods in the embodiments shown are similar in principle and technical effect, and will not be described again here.

[0185] Figure 6 A schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application, such as... Figure 6 As shown, the device 600 provided in this embodiment includes a processor 601 and a memory communicatively connected to the processor. The processor 601 and the memory 602 are connected via a bus 603.

[0186] In a specific implementation, the processor 601 executes the computer execution instructions stored in the memory 602, causing the processor 601 to execute the method in the above method embodiment.

[0187] The specific implementation process of processor 601 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0188] In the above Figure 6In the illustrated embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.

[0189] The memory may include high-speed RAM, and may also include non-volatile storage (NVM), such as at least one disk storage.

[0190] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0191] This application also provides a computer-readable storage medium storing computer-executable instructions. When a processor executes the computer-executable instructions, it implements the metadata processing method of the above-described method embodiments.

[0192] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the metadata processing method described above.

[0193] The aforementioned computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0194] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.

[0195] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0196] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A metadata processing method, characterized in that, include: Obtain metadata generated during the execution of the target business process; Based on the metadata, determine at least one of the following: data correlation degree, business cost value, and data application frequency; The business cost value includes: storage cost and hardware consumption cost. The storage cost is determined in relation to the amount of storage consumed by the target business during execution, and the hardware consumption cost is determined in relation to the total amount of processors and memory consumed by the target business during execution. A target reuse value is determined based on at least one of the data correlation degree, the business cost value, and the data application frequency, wherein the target reuse value represents the likelihood that the metadata will be reused; The metadata can be stored or deleted based on the target reuse value.

2. The method according to claim 1, characterized in that, The step of determining the data correlation degree based on the metadata includes: Determine the number of target groups corresponding to each sub-service involved in the metadata, wherein the target service includes at least one sub-service, each sub-service involves at least one target group, and each target group performs different operations on the corresponding sub-service; The data association value of each sub-service is determined based on the number of target groups corresponding to each sub-service and a preset number threshold; The data association value of each sub-service is summed to obtain the data association degree.

3. The method according to claim 1, characterized in that, Determining the business cost value based on the metadata includes: The storage cost and hardware consumption cost are determined based on the metadata. The storage cost and the hardware consumption cost are summed to obtain the business cost value.

4. The method according to claim 3, characterized in that, Determining the storage cost value based on the metadata includes: The storage consumption of the target service is determined based on the metadata. Determine the unit price of storage costs; The storage cost value is determined based on the storage consumption of the target service and the unit price of the storage cost.

5. The method according to claim 4, characterized in that, Determining the storage consumption of the target service based on the metadata includes: The target data table contained in the metadata is determined by a preset storage path, wherein the target business contains at least one sub-business, and the preset storage path is the storage path for the data corresponding to each sub-business. Determine the initial storage amount for each target data table, and sum the initial storage amounts for each target data table to obtain the storage amount consumed by the target service.

6. The method according to claim 4, characterized in that, The determination of the unit price of storage cost includes: Obtain the basic cost value within a preset time period, wherein the basic cost value is at least one of the following within the preset time period: host depreciation cost, data center rental cost, network facility depreciation cost, and personnel operation and maintenance cost; The first ratio of storage costs is determined based on the procurement costs of the processor, memory, and disk. The unit price of storage cost is determined based on the basic cost value, the first ratio, and the total storage amount, wherein the storage amount consumed by the target service is the storage consumed by the target service every day, and the total storage amount is the total storage consumed by the target service within a preset time period.

7. The method according to claim 3, characterized in that, Determining hardware consumption costs based on the metadata includes: The total number of processors and total memory used during the operation of the target service are determined based on the metadata. Determine the unit cost of the processor and the unit cost of the memory; The processor cost is determined based on the total amount of processors consumed by the target service and the unit price of the processors; the memory cost is determined based on the total amount of memory consumed by the target service and the unit price of the memory. The hardware consumption cost is determined based on the processor cost and the memory cost.

8. The method according to any one of claims 1-7, characterized in that, Determining the data application frequency based on the metadata includes: Determine the number of creators, the number of users, the number of downstream tables associated with the metadata, and the number of upstream tables contained in the metadata; The frequency of data application is determined based on the number of creators, the number of users, the number of downstream tables associated with the metadata, and the number of upstream tables.

9. The method according to any one of claims 1-7, characterized in that, The step of deleting the metadata based on the target reuse value includes: If the target reuse value is lower than the preset reuse value threshold, a metadata deletion prompt will be generated and displayed. In response to a touch operation that prompts for metadata deletion, the metadata corresponding to the target service is deleted.

10. A metadata processing apparatus, characterized in that, include: The acquisition module is used to acquire metadata generated during the execution of the target business. The processing module is used to determine at least one of data correlation degree, business cost value and data application frequency based on the metadata; The business cost value includes: storage cost and hardware consumption cost. The storage cost is determined in relation to the amount of storage consumed by the target business during execution, and the hardware consumption cost is determined in relation to the total amount of processors and memory consumed by the target business during execution. The processing module is further configured to determine a target reuse value based on at least one of the data correlation degree, the business cost value, and the data application frequency, wherein the target reuse value represents the likelihood of the metadata being reused; The processing module is also used to store or delete the metadata based on the target reuse value.

11. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the metadata processing method as described in any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the metadata processing method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Data processing method and related equipment

    CN113901153A

  • Virtual service switch

    US20070260723A1