Metadata management method and device, electronic equipment, storage medium and program product

By binding tenants to data topics, dynamic updates of metadata are solved, and the high cost problems caused by manual configuration in the existing technology are improved, and the scalability and adaptability of the system are improved.

CN120492465APending Publication Date: 2025-08-15BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510560199.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

When adding new business scenarios or data sources, the existing metadata management system requires operation and maintenance personnel to manually modify a large number of configurations, resulting in high development and maintenance costs and it is difficult to adapt to rapidly changing business needs.

Method used

By binding tenants to data topics, the corresponding metadata is automatically searched and updated in the metadata management platform to achieve dynamic updates.

Benefits of technology

It reduces development and maintenance costs, facilitates access to new tenants, and improves the scalability and adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492465A_ABST
    Figure CN120492465A_ABST
Patent Text Reader

Abstract

The invention provides a metadata management method and device, electronic equipment, a storage medium and a program product, and relates to the technical field of data processing.The method comprises the steps that a theme table of a target database is obtained, and the theme table comprises data themes bound by tenants; the first data theme is traversed, first metadata of the first data theme is obtained, the first data theme is a data theme in the theme table, and the first metadata is local metadata corresponding to the first data theme; and importing the first metadata into a theme table of the target database to update the metadata in the theme table. According to the method and the device, the corresponding metadata can be automatically searched and updated based on the data theme, so that the dynamic updating of the metadata is realized, the development and maintenance cost is reduced, and the subsequent access of new tenants is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a metadata management method, device, electronic device, storage medium, and program product. Background Art

[0002] Metadata digitally describes an enterprise's data, processes, and applications, providing context for the content of its digital assets and making it easier to understand, find, manage, and use. By employing scientific and effective mechanisms to manage metadata and providing metadata services to developers and business users, we can meet their business needs and support the development and maintenance of enterprise business systems and data analytics.

[0003] However, the current metadata management and recall design is relatively customized. When adding new business scenarios or data sources, operations and maintenance personnel need to manually modify a large amount of configuration when updating metadata. This results in high development and maintenance costs and makes it difficult to adapt to rapidly changing business needs. Summary of the Invention

[0004] In view of this, the present disclosure provides a metadata management method, apparatus, electronic device, storage medium, and program product to improve the problem of high development and maintenance costs when updating metadata and difficulty in adapting to rapidly changing business needs.

[0005] In a first aspect, the present disclosure provides a metadata management method, the method comprising:

[0006] Obtain a subject table of the target database, where the subject table includes tenant-bound data subjects;

[0007] Traversing the first data subject and obtaining first metadata of the first data subject, wherein the first data subject is a data subject in the subject table, and the first metadata is local metadata corresponding to the first data subject;

[0008] The first metadata is imported into a subject table of a target database to update the metadata in the subject table.

[0009] In a second aspect, the present disclosure provides a metadata management device, the device comprising:

[0010] An acquisition module is used to obtain a subject table of a target database, wherein the subject table includes data subjects bound to the tenant;

[0011] A traversal module, configured to traverse a first data subject and obtain first metadata of the first data subject, wherein the first data subject is a data subject in a subject table, and the first metadata is local metadata corresponding to the first data subject;

[0012] The updating module is used to import the first metadata into the subject table of the target database to update the metadata in the subject table.

[0013] In a third aspect, the present disclosure provides an electronic device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the metadata management method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0014] In a fourth aspect, the present disclosure provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable an electronic device to execute the metadata management method of the first aspect or any corresponding embodiment thereof.

[0015] In a fifth aspect, the present disclosure provides a computer program product, including computer instructions, which are used to enable an electronic device to execute the metadata management method of the first aspect or any corresponding embodiment thereof.

[0016] The metadata management method provided by this disclosure, after obtaining the subject table of a target database, traverses a first data subject and obtains first metadata for the first data subject. The first metadata is then imported into the subject table of the target database to update the metadata in the subject table. This embodiment binds tenants to data subjects. When a new tenant is added, the metadata management platform automatically searches for corresponding metadata based on the data subject, and then updates the metadata in the subject table. This enables dynamic metadata updates, reduces development and maintenance costs, and facilitates the subsequent onboarding of new tenants. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the related technologies, the following briefly introduces the drawings required for use in the specific embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 is a schematic diagram of a metadata management platform according to an embodiment of the present disclosure;

[0019] Figure 2 is a flowchart of a metadata management method according to an embodiment of the present disclosure;

[0020] Figure 3 is a flowchart of another metadata management method according to an embodiment of the present disclosure;

[0021] Figure 4is a flowchart of metadata update according to an embodiment of the present disclosure;

[0022] Figure 5 is a flowchart of a vector library update according to an embodiment of the present disclosure;

[0023] Figure 6 is a flowchart of another metadata management method according to an embodiment of the present disclosure;

[0024] Figure 7 is a flowchart of metadata search according to an embodiment of the present disclosure;

[0025] Figure 8 is a structural block diagram of a metadata management device according to an embodiment of the present disclosure;

[0026] Figure 9 Schematic diagram of the hardware structure of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0027] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present disclosure.

[0028] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0029] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0030] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0031] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0032] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0033] As mentioned in the background technology, the current metadata management and recall design is relatively customized and lacks scalability. If multiple tenants (new business scenarios or data sources) are to be connected in the future, operation and maintenance personnel will need to manually modify a large amount of configurations to update the new metadata to the backend server. This results in high development and maintenance costs, making it difficult to adapt to rapidly changing business needs.

[0034] In view of this, the present disclosure provides a metadata management method, device, electronic device, storage medium and program product. By binding tenants to data topics, the corresponding metadata can be automatically found and updated based on the data topics, thereby realizing dynamic updating of metadata, reducing development and maintenance costs, and facilitating the subsequent access of new tenants.

[0035] The metadata management method provided in the present disclosure can be applied to electronic devices provided with a metadata management platform. The electronic devices may be computers, servers, etc. For ease of understanding, the present disclosure first describes the structure of the metadata management platform.

[0036] like Figure 1 As shown, the metadata management platform includes a tenant layer 110, a data subject layer 120, and a metadata module (not shown). The tenant layer 110 includes tenants, the data subject layer 120 includes data subjects, and the metadata module includes metadata of the data subjects.

[0037] Specifically, tenants can refer to different user groups or organizations that use metadata management platform resources, such as enterprise users. Tenants can be one or more. Figure 1 Take the example where the tenant level 110 includes two tenants (Tenant 1 and Tenant 2), but the present invention is not limited thereto.

[0038] A data subject may represent the category of data content used by a tenant. For example, in the advertising field, a data subject may be advertising basic data or advertising material data, etc. There may be at least one data subject, and a tenant may be bound to at least one data subject.

[0039] For example, Figure 1As shown, the data topic hierarchy includes 6 data topics, namely data topic A1, data topic A2, data topic B1, data topic B2, data topic C1 and data topic C2. Tenant 1 can bind (or associate) data topic A1 and data topic B1, and tenant 2 can bind data topic A2 and data topic C2. It should be understood that Figure 1 The number of data topics and the binding relationship between data topics and tenants are just examples and are not limited to these.

[0040] Specifically, tenants can be created by the tenant manager (enterprise manager) in the user interface (UI) of the metadata management platform and bound to specific data topics based on actual needs. For example, the tenant manager can also modify, add, or delete data topics bound to the tenant through the UI.

[0041] Metadata is data about data. In a multi-tenant scenario, metadata can include descriptive information such as basic tenant information, subject information used by the tenant, data structure definition, and data access permissions.

[0042] Furthermore, if Figure 1 As shown, the metadata module includes a field level 131, a dimension enumeration level 132 and a vector library 133. The field level 131 includes data indicators and / or data dimensions of the data subject, the dimension enumeration level 132 includes enumeration values of the data dimensions, and the vector library 133 includes embedding vectors (embedding) of the field level and embedding vectors of the dimension enumeration level. The metadata includes at least one of the data indicators, data dimensions and enumeration values of the data dimensions.

[0043] Specifically, data indicators can refer to the specific data objects to be analyzed. For example, when the data subject is advertising basic data or advertising creative data, the data indicators can be consumption or the length of time the user enters the live broadcast room, etc. Consumption can be understood as the resource consumption related to advertising delivery.

[0044] Data dimensions can refer to the angles from which data are observed and analyzed. For example, when the data subject is basic advertising data, the data dimensions can be the advertiser's identity (ID), plan ID, or marketing goal, etc. The plan can refer to the delivery plan formulated by the advertiser when delivering an advertisement; when the data subject is advertising creative data, the data dimensions can be the creative ID or creative type, etc.

[0045] The enumeration value of a data dimension can refer to the specific enumerable values under the data dimension. For example, when the data dimension is a marketing goal, the enumeration value can be promoting live broadcasts or promoting products, etc.; when the data dimension is a material type, the enumeration value can be title, video, or breakthrough, etc. Breakthrough can refer to a material category that breaks the routine and attracts user attention.

[0046] A data topic may include at least one data indicator and / or at least one data dimension, e.g. Figure 1 As shown, data subject C2 may include data indicator 1, data dimension 1 and data dimension 2. A data dimension may include at least one enumeration value, for example, Figure 1 As shown, data dimension 2 may include enumeration value 1 and enumeration value 2.

[0047] The vector library is used to store and manage vector data obtained by metadata conversion, for example Figure 1 As shown, the vector library can include a field name vector library and an enumeration name vector library. The field name vector library stores a list of data subject IDs, a list of field IDs, field names, and field name embeddings. Fields can refer to data indicators and / or data dimensions. The enumeration name vector library stores a list of data subject IDs, a list of field IDs, a list of enumeration IDs, enumeration names, and enumeration name embeddings. Embedding is the process of converting high-dimensional, discrete, and complex data into a low-dimensional, continuous, and easily computable vector representation. It also refers to the vector obtained after the conversion.

[0048] Optionally, the vector library may be an ES (Elasticsearch) metadata vector library. Elasticsearch is an open source distributed search and analysis engine that is often used to process large amounts of text data and has powerful full-text search, real-time analysis, and data aggregation capabilities.

[0049] The metadata management method provided by the embodiment of the present disclosure is explained below with reference to the accompanying drawings.

[0050] According to an embodiment of the present disclosure, an embodiment of a metadata management method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in an electronic device such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0051] In this embodiment, a metadata management method is provided, which can be used in the above-mentioned electronic device. Figure 2 is a flowchart of a metadata management method according to an embodiment of the present disclosure. Figure 2 As shown, the process includes the following steps:

[0052] Step S201: Obtain a subject table of a target database.

[0053] The topic table includes the data topics bound to the tenant.

[0054] Specifically, the target database may be a database in a background server connected to the electronic device. A subject table is stored in the target database. The subject table may include data subjects bound to each tenant and metadata under the data subjects.

[0055] For example, the topic table may include data topic A1 bound to tenant 1, metadata under data topic A1, data topic B1 bound to tenant 1, metadata under data topic B1, data topic A2 bound to tenant 2, metadata under data topic A2, data topic C2 bound to tenant 2, and metadata under data topic C2.

[0056] Exemplarily, the target database may be a relational database, such as a Mysql database. After connecting to the target database, the subject table in the target database may be read using Structured Query Language (SQL) and stored in the memory of the electronic device.

[0057] Optionally, the electronic device can be triggered to retrieve the target database's subject table on a regular basis to periodically update the metadata in the subject table. Alternatively, the electronic device can be triggered to retrieve the target database's subject table via an interface. For example, the electronic device can be triggered to retrieve the target database's subject table at midnight every Sunday. Alternatively, an interface can be configured to trigger the electronic device to retrieve the target database's subject table when a user triggers the interface.

[0058] Step S202: traverse the first data subject and obtain first metadata of the first data subject.

[0059] The first data subject is a data subject in the subject table, the first metadata is local metadata corresponding to the first data subject, and the local metadata may be metadata under the data subject in the data subject hierarchy.

[0060] For example, if the topic table includes data topic A1, data topic A2, data topic B1, and data topic B2, the first data topic may be one of data topic A1, data topic A2, data topic B1, and data topic B2. If the first data topic is data topic A1, the first metadata may be the metadata under data topic A1 in the data topic hierarchy 120; if the first data topic is data topic B1, the first metadata may be the metadata under data topic B1 in the data topic hierarchy 120.

[0061] Traversing the first data subject may refer to accessing and processing all data elements or related contents contained in the first data subject one by one in a certain order, and then obtaining the first metadata of the first data subject from the metadata management platform.

[0062] For example, after obtaining the first metadata of the first data subject, the first metadata of the first data subject can be stored in the form of a key-value pair. The storage structure can be Map <String,List <metadatadictionary>>, where Key is the first data subject, the storage form of the first data subject is a string (String), Value is metadata, and the storage form can be a metadata dictionary list type (List <metadatadictionary>), metadata may include data indicators, data dimensions and / or enumeration values of data dimensions.

[0063] Step S203: import the first metadata into the subject table of the target database to update the metadata in the subject table.

[0064] Specifically, after the first metadata is obtained, the metadata under the data subject in the subject table is replaced with the newly obtained first metadata, thereby updating the metadata in the subject table to obtain an updated subject table.

[0065] The metadata management method provided in this embodiment, after obtaining the subject table of the target database, traverses the first data subject and obtains the first metadata of the first data subject. The first metadata is then imported into the subject table of the target database to update the metadata in the subject table. This embodiment binds tenants to data subjects. When a new tenant is added, the metadata management platform automatically searches for the corresponding metadata based on the data subject, and then updates the metadata in the subject table. This enables dynamic metadata updates, reduces development and maintenance costs, and facilitates the subsequent onboarding of new tenants.

[0066] In this embodiment, another metadata management method is provided, which can be used for the above electronic device. Figure 3 is a flowchart of another metadata management method according to an embodiment of the present disclosure. Figure 3 As shown, the process includes the following steps:

[0067] Step S301: Obtain a subject table of a target database.

[0068] For details, please see Figure 2 Step S201 of the illustrated embodiment will not be described in detail here.

[0069] Step S302: traverse the first data subject and obtain first metadata from the metadata management platform through a metadata query interface.

[0070] Specifically, the subject table is configured with a metadata query interface. After obtaining the subject table, the metadata management platform can be accessed through the metadata query interface to obtain the first metadata under the first data subject from the metadata management platform.

[0071] Exemplarily, the metadata query interface includes a business module (Psm), a functional module (Method), and configuration information (Config). The business module is provided with the Internet Protocol (IP) address of the metadata management platform. The functional module is used to provide the functions required to be executed by the metadata query interface, such as providing the function of querying metadata according to conditions. The configuration information is used to adjust and customize the operating behavior, connection parameters, etc. of the metadata query interface.

[0072] Step S303: traverse the first metadata under the first data topic.

[0073] The first metadata may include at least one of a data indicator, a data dimension, and an enumeration value of a data dimension.

[0074] Specifically, after obtaining the first metadata, the data indicators, data dimensions and / or enumerated values of the data dimensions under the first data subject may be accessed and processed one by one in a certain order.

[0075] Step S304: Aggregate the first metadata to obtain a target relationship table.

[0076] The target relationship table includes at least one of a field table, a field-subject relationship table, an enumeration table, an enumeration-dimension relationship table, and an embedded vector table.

[0077] Specifically, the field table is used to store metadata-related field information, and the field table may include data indicators and / or data dimensions; the field-subject relationship table is used to describe the association relationship between fields and data subjects, and the field-subject relationship table may include the correspondence between data indicators and / or data dimensions and data subjects; the enumeration table is used to store enumeration values of data dimensions, and the enumeration-dimension relationship table is used to describe the association relationship between enumeration values and data dimensions, and the enumeration-dimension relationship table may include the correspondence between data dimensions and enumeration values; the embedded vector table is used to store metadata in vector form.

[0078] Exemplarily, there are multiple first data topics and multiple first metadata, and the first metadata includes data indicators, data dimensions, and enumeration values of data dimensions. When traversing the first metadata under the first data topic, the data indicators are aggregated by indicator name, the data dimensions are aggregated by dimension name, and the enumeration values are aggregated by enumeration name to obtain the target relationship table.

[0079] For example, multiple first metadata may include multiple data indicators 1 and multiple data indicators 2. Aggregating data indicators by indicator name means combining multiple data indicators 1 from different data themes to form a single data indicator 1, and combining multiple data indicators 2 from different data themes to form a single data indicator 2. The process of aggregating data dimensions by dimension name and aggregating enumeration values by enumeration name is similar to the process of aggregating data indicators by indicator name and will not be repeated here.

[0080] Exemplarily, after obtaining the first metadata, the electronic device is also used to check whether the data indicators, data dimensions and / or enumeration values of the data dimensions exist in the embedded vector table. For data indicators, data dimensions and / or enumeration values that do not exist in the embedded vector table, the embedded vector model interface is called to obtain them, and finally imported into the embedded vector table.

[0081] Step S305: import the target relational table into the subject table of the target database.

[0082] Specifically, after obtaining the target relationship table, the target relationship table is imported into the subject table to replace the relationship table in the subject table, thereby updating the metadata in the subject table.

[0083] Step S306: Determine the current version number of the first metadata according to the update time.

[0084] Specifically, after importing the target relational table into the subject table of the target database, the minute-level update time can be determined as the current version number of the first metadata. For example, if the first metadata was updated at 23:55 on April 27, 2025, the current version number of the first metadata can be 202504272355.

[0085] Step S307: Delete the first metadata that is at least one version number earlier than the current version number.

[0086] Specifically, after updating the metadata in the subject table, delete redundant versions of the metadata. For example, delete the first metadata that is two version numbers older than the current version. For example, if the current version number is 202504272355, the previous version number is 202504202355, and the previous version number is 202504132355, then you can delete the first metadata of all version numbers before 202504132355, retaining three versions of the first metadata. Alternatively, you can delete the first metadata of all version numbers before 202504202355, retaining two versions of the first metadata.

[0087] The metadata management method provided in this embodiment aggregates the first metadata before importing it into the subject table when obtaining the first metadata, which can save space in the target database. After updating the subject table, the first metadata that is at least one version number older than the current version number is deleted, and at least two versions of the first metadata are retained. This can cope with extreme situations and can restore the subject table in the event of an update error.

[0088] In one example, the metadata update process can be as follows: Figure 4 As shown, first, the interface triggers or timer triggers the update process, then reads the subject table in the MySQL database and the metadata query interface corresponding to the subject table, then traverses each data subject and obtains the metadata of each data subject through the metadata query interface. After obtaining the metadata of each data subject, the data indicators / data dimensions / enumeration values under each data subject are traversed and the data indicators / data dimensions / enumeration values are aggregated to obtain the aggregated results. The aggregated results are then imported into the field table, field-subject relationship table, enumeration table, and enumeration-dimensional relationship table of the MySQL database. Finally, redundant versions of the metadata are deleted, such as deleting all metadata that are two version numbers older than the current version number.

[0089] Exemplarily, after importing the target relational table into the subject table of the target database (ie, step S305), the metadata management method further includes a vector library update method, that is, the metadata management method further includes steps a1 to a3:

[0090] Step a1: Import the target relational table into the Hive table.

[0091] Specifically, Hive tables, similar to tables in traditional relational databases, consist of rows and columns and are used to organize and store large amounts of data for querying, analysis, and processing. Hive tables allow you to manipulate data using a SQL-like language without having to write complex programs, significantly simplifying data processing.

[0092] Step a2: Build a target vector library based on the honeycomb table.

[0093] Specifically, the metadata in the Hive table is converted into embedded vectors to obtain the target vector library.

[0094] For example, metadata in a Hive table may be converted into an embedding vector based on a word embedding model or a deep learning-based model.

[0095] Step a3: import the target vector library into the vector library to update the vector library.

[0096] Specifically, after determining the target vector library, the target vector library can be imported into the vector library via REST, thereby updating the vector library. REST is a software architecture style based on the Hypertext Transfer Protocol (HTTP) that uses standard HTTP methods to operate resources on the network.

[0097] For example, when the target vector library is imported into the vector library, the target vector library carries a document unique identifier (doc_id) and a version number (version), which facilitates data management and version control.

[0098] In one example, the update process of the vector library can be as follows: Figure 5 As shown, first, the vector library update process is triggered regularly (for example, once every other day). After the trigger, the field table, field subject relationship table, enumeration table, enumeration dimension relationship table, and embedded vector table are imported into the corresponding Hive table. Then, the metadata imported into the Hive table is used to build the target vector library, and the target vector library is imported into the ES metadata vector library through the REST method. After that, the metadata management platform is called to switch the metadata version number to the latest version number.

[0099] For example, the metadata management platform can be called to switch the metadata version number to the latest version by using the Curl command (a command-line tool for transferring data). Specifically, the Curl command is an open source file transfer tool that works in command-line mode using Uniform Resource Locator (URL) syntax and supports multiple protocols.

[0100] In this embodiment, another metadata management method is provided, which can be used for the above-mentioned electronic device. Figure 6 is a flowchart of another metadata management method according to an embodiment of the present disclosure. Figure 6 As shown, the process includes the following steps:

[0101] Step S601: Obtain a subject table of a target database.

[0102] For details, please see Figure 2 Step S201 of the illustrated embodiment will not be described in detail here.

[0103] Step S602: traverse the first data subject and obtain first metadata of the first data subject.

[0104] For details, please see Figure 2 Step S202 of the illustrated embodiment will not be described in detail here.

[0105] Step S603: Import the first metadata into the subject table of the target database to update the metadata in the subject table.

[0106] For details, please see Figure 2 Step S203 of the illustrated embodiment will not be described in detail here.

[0107] Step S604: Acquire the target sentence input by the user.

[0108] For example, a search box can be displayed on the data search page of the metadata management platform. When a user performs an input operation in the search box, the target sentence entered by the user in natural language is obtained. Alternatively, a smart object can be displayed on the data search page of the metadata management platform. By clicking the smart object, a smart object page is displayed on the data search page. When a user performs an input operation on the smart object page, the target sentence entered by the user in natural language is obtained.

[0109] Step S605: Process the target sentence to determine the second data subject and second metadata.

[0110] The second data subject is the data subject corresponding to the target sentence, and the second metadata is the metadata corresponding to the target sentence.

[0111] For example, after obtaining the target sentence, the target sentence can be input into a pre-trained machine learning model, and the target sentence can be processed by the machine learning model to obtain the second data topic and second metadata corresponding to the target sentence.

[0112] In this embodiment, the machine learning model can be an existing machine learning model in the relevant technology, or it can be a machine learning model that is improved from an existing machine learning model in the relevant technology. The present disclosed embodiment does not impose any restrictions on this.

[0113] In some embodiments, the machine learning model can be a deep learning model (such as a Transformer model), and the Transformer model can be trained by constructing training samples, so that the Transformer model can accurately extract data topics and metadata in the target sentence.

[0114] For example, a sample sentence annotated with a label and used to query data can be obtained, where the label is used to indicate the real data subject and real metadata of the sample sentence. The sample sentence is input into the Transforme model to obtain the predicted data subject and predicted metadata, and the loss function value of the real value (real data subject and real metadata) and the predicted value (predicted data subject and predicted metadata) is calculated. The parameter value of the Transforme model is adjusted based on the loss function value until the Transforme model converges.

[0115] Step S606: Search in the target database according to the second data subject and the second metadata to obtain the target data subject and target metadata.

[0116] The target data subject is a data subject in the target database corresponding to the second data subject, and the target metadata is metadata under the target data subject.

[0117] Exemplarily, the above step S606 includes:

[0118] Step S6061: Search for a third data topic from the target database according to the second data topic.

[0119] The third data topic is a data topic in the target database whose similarity to the second data topic is greater than a similarity threshold. The similarity threshold can be determined by the designer based on requirements.

[0120] Specifically, similarity is calculated between the second data topic and each data topic extracted from the target database, and a data topic in the target database whose similarity is greater than a similarity threshold is determined as a third data topic.

[0121] Exemplarily, the similarity between the second data topic and the first data topic may be determined by means of cosine similarity or edit distance (Levenshtein distance).

[0122] Step S6062: Determine a distance score between the first metadata and the second metadata under the third data theme.

[0123] Specifically, a distance score between the first metadata and the second metadata may be determined based on the embedding vector corresponding to the first metadata and the embedding vector corresponding to the second metadata, wherein a smaller distance score indicates a higher similarity between the first metadata and the second metadata.

[0124] Exemplarily, when the metadata type is text data, the distance score between the first metadata and the second metadata can be determined based on edit distance or cosine similarity; when the metadata type is numerical data, the distance score between the first metadata and the second metadata can be determined based on Euclidean distance and Manhattan distance.

[0125] Step S6063: Determine the target data subject and target metadata based on the distance score.

[0126] Specifically, the third data topic with the smallest distance score may be determined as the target data topic, and the metadata under the target data topic may be determined as the target metadata.

[0127] For example, the second data topic is data topic D, and the topic table includes data topics A1, A2, B1, and B2. First, the similarity between data topic D and data topics A1, A2, B1, and B2 needs to be determined. If the similarity between data topics B1 and B2 and data topic D is greater than the similarity threshold, the third data topics are data topics B1 and B2. Then, the distance scores between the metadata under data topic B1 and the metadata under data topic B2 and the second metadata are determined. If the distance score between the metadata under data topic B1 and the second metadata is small, data topic B1 is the target data topic, and the metadata under data topic B1 is the target metadata.

[0128] The metadata management method provided in this embodiment, after obtaining a target sentence input by the user, processes the target sentence to obtain a second data subject and second metadata. This is then used to search the target database for the second data subject and second metadata, obtaining the target data subject and target metadata. Compared to the brute force traversal method directly based on embedded vectors in related technologies, this embodiment searches (recalls) data in the target database based on a combination of data subjects and metadata, simplifying the data search process and improving search accuracy and convenience.

[0129] In one example, the metadata lookup process can be as follows Figure 7 As shown in the figure, after the user question (i.e., the target sentence) is written in the user layer, the keywords of the user question can be split based on the machine learning model, and then N data topics corresponding to the user question and metadata under the N data topics can be extracted based on the keywords of the user question. Where N is an integer greater than or equal to 1, Figure 7 Taking N as 3 as an example, it includes data topic 1, data topic 2 and data topic 3, but is not limited to this.

[0130] After obtaining N data topics and corresponding metadata, the metadata layer is entered to search the target database for data topics that match the N data topics (i.e., the third data topic mentioned above). For example, if the similarity between data topic 1 and data topic A in the target database is greater than the similarity threshold, then data topic A matches data topic 1. After determining the third data topic, the distance score between the first metadata and the second metadata under the third data topic is determined, for example, Figure 7 As shown, the distance scores of the metadata (data indicators / data dimensions / enumeration values) under data topic 1 and the metadata under data topic A (i.e., the first metadata) are determined, and the metadata with the smallest distance score under each third data topic (i.e., Top-N data indicators / Top-N data dimensions / Top-N enumeration values) are determined.

[0131] After determining the metadata with the smallest distance score under each third data theme, the third data theme is filtered based on the distance threshold, that is, the third data theme with a distance score less than the distance threshold is screened out. The distance threshold can be determined by the designer according to the requirements. After filtering, the metadata under the third data theme that meets the threshold condition (the distance score is less than the distance threshold) is aggregated to obtain an aggregated result. For example, if the metadata of data theme A and the metadata of data theme B both meet the threshold condition, the aggregated result is the metadata of data theme A and the metadata of data theme B.

[0132] After obtaining the summary result, if there are multiple data topics in the summary result, the data topic with the lowest distance score is determined as the target data topic, and the metadata under the target data topic is determined as the target metadata.

[0133] After obtaining the target data subject and target metadata, the target metadata is input into the large model, which converts the target metadata into a domain-specific language (DSL), which is natural language that users can understand, and displays the query results to the user. The large model can be a deep learning model, such as the Bidirectional Encoder Representations from Transformers (BERT) model.

[0134] This embodiment also provides a metadata management device for implementing the aforementioned embodiments and preferred implementations. Details already described will not be repeated. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. While the devices described in the following embodiments are preferably implemented using software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0135] This embodiment provides a metadata management device, such as Figure 8 As shown, including:

[0136] An acquisition module 801 is configured to acquire a subject table of a target database, wherein the subject table includes tenant-bound data subjects;

[0137] A traversal module 802 is configured to traverse a first data subject and obtain first metadata of the first data subject, wherein the first data subject is a data subject in a subject table and the first metadata is local metadata corresponding to the first data subject;

[0138] The updating module 803 is configured to import the first metadata into the subject table of the target database to update the metadata in the subject table.

[0139] In some optional embodiments, the device is applied to a metadata management platform, which includes a tenant level, a data subject level and a metadata module. The tenant level includes tenants, the data subject level includes data subjects, the metadata module includes metadata of the data subjects, and the local metadata is the metadata under the data subjects in the data subject level.

[0140] In some optional implementations, the subject table is configured with a metadata query interface, and the traversal module 802 includes:

[0141] The first acquiring unit is configured to acquire first metadata from the metadata management platform through a metadata query interface.

[0142] In some optional embodiments, the metadata module includes a field level, a dimension enumeration level, and a vector library, the field level includes data indicators and / or data dimensions of the data subject, the dimension enumeration level includes enumeration values of the data dimensions, the vector library includes embedded vectors of the field level and embedded vectors of the dimension enumeration level, and the metadata includes at least one of the data indicators, data dimensions, and enumeration values of the data dimensions.

[0143] In some optional embodiments, the device further comprises:

[0144] A processing module, configured to traverse first metadata under a first data topic;

[0145] an aggregation module, configured to aggregate the first metadata to obtain a target relationship table, wherein the target relationship table includes at least one of a field table, a field-subject relationship table, an enumeration table, an enumeration-dimension relationship table, and an embedded vector table;

[0146] The update module 803 includes:

[0147] The first updating unit is used to import the target relational table into the subject table of the target database.

[0148] In some optional embodiments, the device further comprises:

[0149] The first import module is used to import the target relational table into the honeycomb table;

[0150] A construction module is used to construct a target vector library based on the honeycomb table;

[0151] The second import module is used to import the target vector library into the vector library to update the vector library.

[0152] In some optional embodiments, the device further comprises:

[0153] A first determining module, configured to determine a current version number of the first metadata according to an update time;

[0154] The deletion module is used to delete the first metadata that is at least one version number earlier than the current version number.

[0155] In some optional embodiments, the device further comprises:

[0156] A receiving module is used to obtain the target sentence input by the user;

[0157] A second determination module is configured to process the target sentence and determine a second data subject and second metadata, wherein the second data subject is the data subject corresponding to the target sentence and the second metadata is the metadata corresponding to the target sentence;

[0158] The search module is used to search in the target database according to the second data subject and the second metadata to obtain the target data subject and the target metadata.

[0159] In some optional implementations, the search module includes:

[0160] a search unit, configured to search a third data topic from the target database according to the second data topic, wherein the third data topic is a data topic in the target database whose similarity to the second data topic is greater than a similarity threshold;

[0161] a first determining unit, configured to determine a distance score between the first metadata and the second metadata under a third data topic;

[0162] The second determining unit is configured to determine the target data subject and the target metadata according to the distance score.

[0163] The metadata management device provided in the embodiments of the present disclosure can execute the metadata management method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method. The further functional description of each of the above modules and units is the same as that of the corresponding embodiments above, and will not be repeated here.

[0164] Figure 9 A schematic structural diagram of an electronic device provided in an embodiment of the present disclosure.

[0165] The following specific reference Figure 9 , which shows a schematic diagram of the structure of the electronic device suitable for implementing the embodiments of the present disclosure. The electronic device may include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a memory 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the electronic device are also stored in the RAM 903. The processor 901, ROM 902, and RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0166] Typically, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a memory 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 9 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all of the devices shown, and more or fewer devices may be implemented or possessed instead.

[0167] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via communication device 909, or installed from memory 908, or installed from ROM 902. When the computer program is executed by processor 901, the above-described functions defined in the metadata management method of the embodiment of the present disclosure are performed.

[0168] Figure 9 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0169] The embodiments of the present disclosure also provide a computer-readable storage medium. The above-mentioned method according to the embodiments of the present disclosure can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded via a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, processor or hardware, the metadata management method shown in the above embodiment is implemented.

[0170] A portion of the present disclosure may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present disclosure through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes but is not limited to a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium that can be accessed by the computer.

[0171] Although the embodiments of the present disclosure have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.< / metadatadictionary> < / metadatadictionary>

Claims

1. A metadata management method, characterized in that: The method comprises: Obtaining a subject table of a target database, wherein the subject table includes tenant-bound data subjects; Traversing a first data topic and obtaining first metadata of the first data topic, wherein the first data topic is a data topic in the topic table, and the first metadata is local metadata corresponding to the first data topic; The first metadata is imported into the subject table of the target database to update the metadata in the subject table.

2. The method according to claim 1, characterized in that The method is applied to a metadata management platform, which includes a tenant level, a data subject level, and a metadata module. The tenant level includes the tenants, the data subject level includes the data subjects, the metadata module includes the metadata of the data subjects, and the local metadata is the metadata under the data subjects in the data subject level.

3. The method according to claim 2, characterized in that The subject table is configured with a metadata query interface, and obtaining the first metadata of the first data subject includes: The first metadata is obtained from the metadata management platform through the metadata query interface.

4. The method according to claim 2, characterized in that The metadata module includes a field level, a dimension enumeration level and a vector library, the field level includes data indicators and / or data dimensions of the data subject, the dimension enumeration level includes enumeration values of the data dimensions, the vector library includes embedded vectors of the field level and embedded vectors of the dimension enumeration level, and the metadata includes at least one of data indicators, data dimensions and enumeration values of data dimensions.

5. The method according to claim 4, characterized in that Before importing the first metadata into the subject table of the target database, the method further includes: Traversing the first metadata under the first data subject; Aggregating the first metadata to obtain a target relationship table, wherein the target relationship table includes at least one of a field table, a field-subject relationship table, an enumeration table, an enumeration-dimension relationship table, and an embedded vector table; Importing the first metadata into the subject table of the target database includes: Import the target relational table into the subject table of the target database.

6. The method according to claim 5, characterized in that After importing the target relational table into the subject table of the target database, the method further includes: Importing the target relational table into the honeycomb table; Constructing a target vector library according to the honeycomb table; The target vector library is imported into the vector library to update the vector library.

7. The method according to any one of claims 1 to 6, characterized in that After importing the first metadata into the subject table of the target database, the method further includes: Determining a current version number of the first metadata according to the update time; Delete the first metadata that is at least one version number earlier than the current version number.

8. The method according to any one of claims 1 to 6, characterized in that After importing the first tuple into the subject table of the target database, the method further includes: Get the target sentence input by the user; Processing the target sentence to determine a second data subject and second metadata, wherein the second data subject is the data subject corresponding to the target sentence, and the second metadata is the metadata corresponding to the target sentence; A search is performed in the target database according to the second data subject and the second metadata to obtain a target data subject and target metadata.

9. The method according to claim 8, characterized in that The searching in the target database according to the second data subject and the second metadata to obtain the target data subject and the target metadata includes: Searching for a third data topic from the target database according to the second data topic, wherein the third data topic is a data topic in the target database whose similarity to the second data topic is greater than a similarity threshold; determining a distance score between the first metadata and the second metadata under the third data subject; The target data subject and the target metadata are determined according to the distance score.

10. A metadata management device, characterized in that: The device comprises: An acquisition module, configured to acquire a subject table of a target database, wherein the subject table includes data subjects bound to the tenant; A traversal module, configured to traverse a first data subject and obtain first metadata of the first data subject, wherein the first data subject is a data subject in the subject table, and the first metadata is local metadata corresponding to the first data subject; An updating module is configured to import the first metadata into a subject table of the target database to update the metadata in the subject table.

11. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the metadata management method according to any one of claims 1 to 9 by executing the computer instructions.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable an electronic device to execute the metadata management method according to any one of claims 1 to 9.

13. A computer program product, characterized in that The method comprises computer instructions for causing an electronic device to execute the metadata management method according to any one of claims 1 to 9.