A method, device, medium and equipment for constructing metadata based on container technology

Through the metadata construction method based on container technology, the problem of low processing efficiency of high concurrent metadata construction requests in the existing technology is solved, efficient resource utilization and system availability are achieved, and it is suitable for processing metadata construction needs with high concurrency and low latency.

CN119226589BActive Publication Date: 2025-06-20ZHEJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411763553.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-06-20
Estimated Expiration
2044-12-03

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently handle highly concurrent metadata construction requests, resulting in resource competition and conflicts, affecting the availability and efficiency of the system.

Method used

Using a metadata construction method based on container technology, the metadata construction tasks are isolated through containers, and multi-threaded and highly concurrent execution of metadata construction tasks is achieved to optimize system performance and resource utilization.

Benefits of technology

The construction efficiency of metadata construction tasks is improved through container technology, the availability and resource utilization of the system are enhanced, resource competition and conflict are reduced, and burst traffic can be dealt with by increasing node level expansion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119226589B_ABST
    Figure CN119226589B_ABST
Patent Text Reader

Abstract

This specification provides a method, apparatus, medium, and device for constructing metadata based on container technology. It receives a metadata construction request for the original data and determines the parameter information of the original data. It calls the container for constructing metadata, inputs the parameter information into the container to wake up the client tool of the container, and based on the client tool, selects the corresponding metadata processor, enabling the container to obtain the original data and construct the corresponding metadata according to the original data. According to the metadata constructed by the container, it updates the metadata index table according to the preset storage rules. By isolating each metadata construction task through the container, good resource isolation is achieved among the metadata construction tasks. Through the management and scheduling of the container, the metadata construction tasks are executed in a multi-threaded and highly concurrent manner, improving the construction efficiency of the metadata construction tasks while enhancing the availability and resource utilization rate of the system. It is easy to horizontally expand the system by adding nodes to handle sudden traffic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of data processing, and particularly to a method, apparatus, medium, and device for constructing metadata based on container technology. Background Art

[0002] With the development of computer technology, metadata, as data that describes data, provides detailed information about the data, which helps in the management and use of the data and greatly improves the efficiency of data usage. Therefore, the efficiency of metadata construction is crucial.

[0003] In the prior art, metadata is usually constructed by means of bare-metal deployment. For example, running a metadata construction program on a bare metal and directly using the underlying hardware resources to construct metadata, etc. Usually, a single service on a single server is used to process metadata construction requests.

[0004] However, with the development of big data, for some datasets that require high dynamics, such as real-time transaction data, sensor data, etc., or data that needs to be monitored in real time, such as system performance monitoring data, etc., users need the metadata to be updated in real time to ensure the efficiency of data usage and achieve timely decision-making. As a result, the construction and update requirements of metadata are often high-concurrency and low-latency. The existing metadata construction methods are difficult to efficiently process a large number of concurrent metadata construction requests. Therefore, this specification provides a method, apparatus, medium, and device for constructing metadata based on container technology. Summary of the Invention

[0005] This specification provides a method, apparatus, medium, and device for constructing metadata based on container technology to partially solve the above problems existing in the prior art.

[0006] This specification adopts the following technical solutions:

[0007] A method for constructing metadata based on container technology includes:

[0008] Receiving a metadata construction request for original data, and determining parameter information of the original data, where the parameter information at least includes a storage address of the original data;

[0009] Invoking a container for constructing metadata, inputting the parameter information into the container to wake up a client tool of the container, and based on a metadata processor called by the client tool, enabling the container to obtain the original data and construct corresponding metadata according to the original data;

[0010] Updating a metadata index table according to the metadata constructed by the container according to a preset storage rule.

[0011] Optionally, the parameter information further includes a data source type, specifically including:

[0012] Input the parameter information into the container to wake up the client tool of the container, and based on the metadata processor called by the client tool, enable the container to obtain the original data and construct corresponding metadata according to the original data. Specifically, it includes:

[0013] Input the parameter information into the container, enable the container to determine the data source type of the original data according to the parameter information, and select the metadata processor corresponding to the data source type from the client tools of the container; start the metadata construction task corresponding to the container through the selected metadata processor, enable the container to execute the metadata construction task to obtain the original data and construct corresponding metadata according to the original data.

[0014] Optionally, the metadata construction request carries the update frequency of the metadata of the original data, and the method further includes:

[0015] Input the parameter information and the update frequency into the container, enable the container to regularly obtain the original data from the storage address according to the input parameter information at the update frequency and construct the metadata of the obtained original data;

[0016] Update the metadata index table according to the metadata constructed by the container based on a preset storage rule. Specifically, it includes:

[0017] Regularly update the metadata index table according to the constructed metadata at the update frequency based on a preset storage rule.

[0018] Optionally, updating the metadata index table according to the metadata constructed by the container based on a preset storage rule specifically includes:

[0019] Determine the unique identifier of the metadata constructed by the container according to the original data;

[0020] Determine a first identifier set according to the unique identifiers of the metadata stored in the metadata index table;

[0021] Update the metadata index table according to the unique identifier and the stored first identifier set based on a preset storage rule.

[0022] Optionally, updating the metadata index table according to the metadata constructed by the container based on a preset storage rule specifically includes:

[0023] Determine each container used for constructing metadata;

[0024] Determine a second identifier set according to the unique identifiers corresponding to the metadata constructed by each container;

[0025] Compare each unique identifier included in the second identifier set with each unique identifier in the first identifier set, and determine each unique identifier that is included in the first identifier set but not included in the second identifier set as the identifier to be deleted;

[0026] Delete the metadata corresponding to the identifier to be deleted from the metadata index table, update the first identifier set according to the second identifier set, and update the metadata corresponding to each unique identifier in the second identifier set to the metadata index table according to the preset storage rules.

[0027] Optionally, before updating the metadata index table according to the metadata constructed by the container according to the preset storage rules, the method further includes:

[0028] Determine the metadata constructed by the container, and determine whether the metadata conforms to the preset format;

[0029] If not, convert the metadata into the preset format.

[0030] Optionally, the metadata is stored in the metadata index table in the form of key-value pairs, and the method further includes:

[0031] In response to a query instruction, determine the retrieval information of the original data to be retrieved according to the query instruction;

[0032] Retrieve in the metadata index table according to the retrieval information, and output at least one key-value combination according to the retrieval result.

[0033] This specification provides a metadata construction device based on container technology, including:

[0034] A determination module, configured to receive a metadata construction request for original data, and determine parameter information of the original data, where the parameter information at least includes a storage address of the original data;

[0035] A construction module, configured to call a container for constructing metadata, input the parameter information into the container to wake up a client tool of the container, and based on a metadata processor called by the client tool, enable the container to obtain the original data and construct corresponding metadata according to the original data;

[0036] A storage module, configured to update a metadata index table according to the metadata constructed by the container according to preset storage rules.

[0037] This specification provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, the above-mentioned metadata construction method based on container technology is implemented.

[0038] The above at least one technical solution adopted in this specification can achieve the following beneficial effects:

[0039] In the method for constructing metadata based on container technology described in this specification, a metadata construction request for the original data is received, and the parameter information of the original data is determined. A container for constructing metadata is called, and the parameter information is input into the container to wake up the client tool of the container, and a corresponding metadata processor is selected based on the client tool, so that the container obtains the original data and constructs the corresponding metadata according to the original data. According to the metadata constructed by the container, the metadata index table is updated according to the preset storage rules.

[0040] It can be seen from the above method that by isolating each metadata construction task through the container, good resource isolation is achieved between each metadata construction task. Through the management and scheduling of the container, the metadata construction tasks are executed in a multi-threaded and highly concurrent manner, improving the construction efficiency of the metadata construction tasks, while improving the availability and resource utilization rate of the system. It is easy to horizontally expand the system by adding nodes to cope with sudden traffic. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The drawings described herein are used to provide a further understanding of this specification, and constitute a part of this specification. The schematic embodiments of this specification and their descriptions are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings:

[0042] Figure 1 It is a schematic diagram of the process of a method for constructing metadata based on container technology provided by this specification;

[0043] Figure 2 It is a schematic diagram of a server for processing parallel metadata construction requests provided by this specification;

[0044] Figure 3 It is a schematic diagram of a container provided by this specification;

[0045] Figure 4 It is a schematic diagram of metadata construction and storage provided by this specification;

[0046] Figure 5 It is a schematic diagram of a metadata construction device based on container technology provided by this specification;

[0047] Figure 6 Corresponding to Figure 1 the schematic diagram of the electronic device. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] To make the objectives, technical solutions and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part rather than all of the embodiments of this specification. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this invention.

[0049] With the development of computer technology, metadata, as data that describes data, provides detailed information about the data, facilitates the management and use of the data, and greatly improves the efficiency of data use. However, with the development of big data, for some highly dynamic data sets, such as real-time transaction data, sensor data, etc., or data that needs to be monitored in real time, such as system performance monitoring data, etc., users need the metadata to be updated in real time to ensure the efficiency of data use and achieve timely decision-making, resulting in the construction and update requirements of metadata often being highly concurrent and low-latency. However, the existing method of constructing metadata by generating a single service on bare metal or by deploying a single service using traditional virtual machines will cause the server to be difficult to efficiently process a large number of concurrent metadata construction requests. Especially when multiple metadata construction tasks are running simultaneously, it may lead to resource competition and resource conflicts between the metadata construction tasks, and there may also be data competition and even system crashes. Therefore, this specification provides a metadata construction method based on container technology to dynamically adjust the execution order and resource allocation of tasks through container technology, optimize system performance, reduce resource competition and conflicts, and improve the reliability and stability of metadata construction.

[0050] The following will detail the technical solutions provided by each embodiment of this specification in conjunction with the drawings.

[0051] Figure 1 It is a schematic diagram of the process of a metadata construction method based on container technology provided by this specification. Specifically:

[0052] S100: Receive a metadata construction request for the original data, and determine the parameter information of the original data. The parameter information at least includes the storage address of the original data.

[0053] In one or more embodiments of this specification, there is no limitation on the specific device that executes the metadata construction method based on container technology. For example, it can be a mobile terminal, a server, etc. However, in container technology, usually a server is responsible for managing and scheduling containers. Therefore, in the following of this specification, the execution of the metadata construction method based on container technology by a server is taken as an example for description. Herein, the server can be a single device or composed of multiple devices. For example, a distributed server. This specification does not limit this, and it can also refer to the server of cloud services.

[0054] Since metadata is data that describes data, in order to execute the metadata construction task, the server should first obtain the parameter information of the original data to be constructed with metadata, so that the container can construct the metadata of the original data based on the parameter information of the original data.

[0055] Specifically, the server can determine the parameter information of the original data according to the metadata construction request. The parameter information is the data and information required to execute the metadata construction task, and at least includes the storage address of the original data. In addition, the parameter information can also include the data source type of the original data, the Uniform Resource Locator (URL) for accessing the storage address of the original data, the interface information for accessing the storage address of the original data, etc. The data source type is the type of the data storage platform, and the server can connect to the data storage platform through the storage address of the data source determined by the metadata construction task.

[0056] In addition, in one or more embodiments of this specification, there is no limitation on the specific manner in which the server determines the parameter information of the original data according to the metadata construction request. It can be that the parameter information of the original data is carried in the metadata construction request, or the storage address of the parameter information of the original data is carried in the metadata construction request. Of course, it can also be through other means. Since there are many available means, they will not be elaborated here in this specification.

[0057] S102: Invoke the container for constructing metadata, input the parameter information into the container to wake up the client tool of the container, and based on the metadata processor invoked by the client tool, enable the container to obtain the original data and construct the corresponding metadata according to the original data.

[0058] As the development of big data, the construction of metadata often needs to be highly concurrent. That is, the server usually receives a large number of metadata construction requests. To prevent resource competition and conflicts among a large number of metadata construction tasks when responding to each metadata construction request under high concurrency, container technology can be adopted to package resources such as the program for constructing metadata and the running environment of the program into an image. Then, in response to each metadata construction request, the server calls the container carrying the image to execute each metadata construction task, achieving isolation of the tasks executed in each container. At the same time, through the management and scheduling of containers, the system performance is optimized, and resource competition and conflicts are reduced.

[0059] Specifically, as Figure 2 shown, Figure 2 is a schematic diagram of a server for processing parallel metadata construction requests provided in this specification. For each metadata construction request, the server calls at least one container for processing the metadata construction request, inputs parameter information into the container to wake up a preset client tool (client) in the container, and based on the client, the container obtains the original data and constructs the corresponding metadata according to the original data. That is, when the server receives n metadata construction requests, the server calls at least n containers to simultaneously process each metadata construction request respectively. Through the sever and the clients in the called containers, a server-client architecture is implemented, so that when constructing metadata, the traditional single-server processing of parallel metadata construction requests is transformed into multiple clients simultaneously processing each metadata construction request. That is, through container technology, the single-server processing of parallel metadata construction requests is transformed into multiple servers simultaneously processing parallel metadata construction requests. Resource competition and resource conflicts, as well as the situation of system crashes caused by data competition, are reduced. At the same time, through the integration of different types of metadata processors by the client, the impact on the sever's execution of unified metadata storage services, etc., is reduced.

[0060] It should be noted that the container image for executing the metadata construction task is pre-constructed. The container image contains all the programs required for executing the metadata construction task and the running environment of the program, and each container image contains a corresponding client. The client can execute the metadata construction task by calling the metadata processor, and the server can interact with the clients in the called containers to make the corresponding containers execute the corresponding metadata construction tasks.

[0061] Of course, in one or more embodiments of this specification, there is no limitation on where the preset container image is specifically stored. It can be in the same device as the server or in a different device from the server. There is also no limitation on how many containers are specifically called to execute a metadata construction task, but at least one.

[0062] In addition, in one or more embodiments of this specification, there is no limitation on the specific form of the metadata construction request. It can be a string of codes, a string of data, a string of signals, or even a piece of text, etc. Since there are many forms of instructions, this specification does not list them one by one.

[0063] In addition, when the server calls a container to execute a metadata construction task, it can select a target running area according to the computing resources, running resources, or storage resources of the device where the server is located, and run the container there. If the server is a distributed server, the server can also determine the node with the most idle resources according to the resources of each node to run the container, or determine the running area of the called container in the device according to other methods. The server can set it according to actual needs, and this specification does not limit it.

[0064] Then, in order to construct metadata through a container, the server can input the parameter information of the original data into the container, so that the container executes the metadata construction task based on the input parameter information.

[0065] Specifically, the server inputs the obtained parameter information into the container. Then, based on the parameter information, the container determines the storage address of the original data, obtains the original data by accessing the storage address, and then constructs the metadata corresponding to the original data based on the original data.

[0066] It should be noted that in one or more embodiments of this specification, there is no limitation on the specific protocol for transmitting the metadata between the client and the server, such as HTTPS, etc. Of course, this specification also does not limit the protocol for subsequent transmission between the server and the storage.

[0067] S104: Update the metadata index table according to the metadata constructed by the container according to a preset storage rule.

[0068] In order to improve the accuracy and efficiency of retrieving the original data based on the constructed metadata in the future, the server can update the metadata index table according to the metadata constructed by each container according to a preset storage rule.

[0069] Specifically, the server determines the metadata indicating that the container has been built, and then, based on the built metadata, determines whether the metadata is stored in the metadata index table. If so, the built metadata replaces the metadata stored in the metadata index table. If not, the metadata is directly stored.

[0070] It should be noted that after each container builds the metadata, the built metadata is directly updated in the metadata index table correspondingly. The metadata corresponding to the task number of the metadata construction task can be updated in the metadata index table. Of course, based on other numbers, the metadata built by each container can also be updated in the metadata index table. For example, according to the number of the container, the metadata built by the container is updated. The built metadata can be stored in the metadata index table in the form of a key-value combination to be compatible with metadata of different data source types and facilitate subsequent retrieval of the original data based on the metadata. Of course, it can also be stored according to other forms of preset storage rules, which are not limited in one or more embodiments of this specification and can be set according to actual needs.

[0071] In addition, in one or more embodiments of this specification, there is no limitation on the specific preset storage rule adopted by the server to update the metadata index table. The preset storage rule refers to the content of the metadata stored in the metadata index table, such as the metadata construction task number, metadata, identification set, etc. The built metadata can also be stored using a metamodel. A metamodel is a storage rule for metadata that defines the structure and semantics of the metadata and provides a basis for the management and use of the metadata. In the metadata index table constructed according to the metamodel, the metadata number, metadata name, data source type, and metadata can be stored. The metadata includes corresponding key-value combinations stored according to the data source types to which different metadata belong.

[0072] Based on Figure 1 a metadata construction method based on container technology, which receives a metadata construction request for the original data and determines the parameter information of the original data. Invokes and runs a container for building metadata, inputs the parameter information into the container, enables the container to obtain the storage address of the original data according to the input parameter information, accesses the storage address to obtain the original data, and builds the corresponding metadata according to the original data. Updates the metadata index table according to the preset storage rule based on the metadata built by the container.

[0073] As can be seen from the above method, isolating each metadata construction task through containers enables good resource isolation between the metadata construction tasks. Through the management and scheduling of containers, the metadata construction tasks are executed with multi-threaded high concurrency. While improving the construction efficiency of the metadata construction tasks, the availability and resource utilization rate of the system are also improved. It can be easily horizontally scaled by adding nodes to handle sudden traffic.

[0074] In addition, when constructing metadata, different types of data have different processing flows according to their types, and different processing steps are required for different processing flows. Therefore, when the parameter information contains the data source type, when constructing metadata, the server can input the parameter data into the called container, so that the client in the called container can select the corresponding metadata processor according to the data source type contained in the parameter information, thereby realizing the execution of the metadata construction task on the raw data of different data source types. Furthermore, when packaging the container image, the image contains the client and the environment on which the client depends, and the client internally contains metadata processors of different types, and each type of processor is used to process different types of raw data. As Figure 3 shown, Figure 3 is a schematic diagram of a container provided in this specification. Among them, inside each container, there is a client, and each client contains metadata processors for processing different data.

[0075] That is, if the parameter information does not contain the data source type, a metadata processor in a preset type of client is selected to process the metadata construction task. If the data source type corresponding to the raw data is included in the parameter information, the server can determine and call the metadata processor of the corresponding type of client according to the data source type to process the metadata construction task of this data source type.

[0076] Furthermore, in order for the server to meet the needs of some highly dynamic data sets, such as real-time transaction data, sensor data, etc., or data that needs to be monitored in real time, such as system performance monitoring data, etc., and to prevent the server from frequently calling and deleting containers for the metadata construction tasks of the raw data at the same storage address, the server can also determine the update frequency of the metadata carried in the metadata construction request, that is, through a container, incrementally construct the metadata of the raw data at the same storage address regularly.

[0077] Specifically, input the parameter information into the container, so that the container regularly obtains the raw data from the storage address based on the update frequency of the metadata in the parameter information, and constructs the metadata of the raw data obtained each time for the raw data obtained each time.

[0078] It should be noted that when determining the metadata update frequency, in addition to determining based on the metadata construction request, the server can also execute the metadata construction task based on the preset metadata update frequency. Of course, the server can also select the corresponding update frequency from the preset update frequencies as the update frequency of this metadata construction task based on the data source type of the original data.

[0079] When the container regularly obtains the original data from the storage address based on this update frequency and constructs the metadata of the original data, that is, when updating the metadata index table, it can also update the metadata index table regularly according to the constructed metadata based on the determined update frequency according to the preset storage rules.

[0080] Of course, if there is no restriction on the construction frequency of the container to construct metadata, the server can only obtain all the original data at one time and complete the construction and storage of metadata once, that is, the one-time full-volume construction of metadata.

[0081] Furthermore, when updating the metadata index table according to the preset storage rules according to the update frequency, in order to achieve incremental metadata construction and at the same time reduce the occupation of storage by the number of expired metadata stored in the storage and the excessive noise caused to subsequent metadata retrieval, when the server updates the metadata stored in the metadata index table, it at least includes updating, inserting, and deleting the constructed metadata, that is, modifying the existing metadata in the metadata index table, inserting the metadata not stored in the metadata index table into the metadata index table for storage, and deleting the metadata stored in the metadata index table but not included in the metadata constructed this time.

[0082] Specifically, when updating the metadata, the server can determine the unique identifier of the metadata corresponding to the original data according to each original data, and then determine the first identifier set according to the unique identifiers of the metadata stored in the metadata index table. Then, by comparing the unique identifier of the metadata constructed this time with the unique identifiers in the first identifier set, the original data index table is updated.

[0083] It should be noted that when determining the unique identifier of the metadata corresponding to the original data, the unique identifier of the original data can be determined first, and then the unique identifier of the original data is used as the unique identifier of the metadata of the original data. Of course, the unique identifier of the original data can also be encrypted, and the encrypted unique identifier is used as the unique identifier of the metadata of the original data. This specification does not limit this.

[0084] When deleting metadata, the server can determine the containers used to build the metadata, and determine a second identifier set based on the unique identifiers corresponding to the metadata built from each container. By comparing each unique identifier included in the second identifier set with each unique identifier in the first identifier set, determine each unique identifier that is included in the first identifier set but not in the second identifier set as the identifier to be deleted. Delete the metadata corresponding to the identifier to be deleted from the metadata index table, update the first identifier set according to the second identifier set, and update the metadata corresponding to each unique identifier in the second identifier set to the metadata index table according to the preset storage rules.

[0085] When there are unique identifiers that are not included in the first identifier set but are included in the second identifier set, the server can directly store the metadata corresponding to these identifiers in the metadata index table according to the preset storage rules, and add these identifiers to the first identifier set to update the first identifier set. Among them, the first identifier set and the metadata are stored separately to reduce the impact of network congestion on storage efficiency due to their different update frequencies.

[0086] Furthermore, after the metadata update is completed, the server can retrieve the stored metadata in the storage according to the input information and output the retrieved key-value combinations. It should be noted that in one or more embodiments of this specification, there is no limitation on the specific retrieval method adopted by the server. It can be to convert the input information and each key-value combination into vectors, and determine the target key-value combination by calculating the matching degree between the vectors. It can also perform keyword matching, or fuzzy matching, or full-text retrieval on the input information and each key-value combination in the storage to determine the target key-value combination. It can be specifically set according to actual needs. In addition, the above input information refers to the information of the metadata. Of course, if the retrieval information of the original data is not stored in the metadata index table, the server also outputs an empty or other preset identifier instead of outputting the metadata.

[0087] Furthermore, in order to save storage space and protect data security, the server can also perform encoding and compression processing on the first identifier set, and decode and decompress the first identifier set when updating the metadata in the storage.

[0088] In addition, in order to facilitate subsequent indexing according to the metadata index table, the server can also unify the formats of the metadata built from each container based on the preset storage rules. Specifically, after the server converts the metadata built from each container into the format of the preset storage rules, it updates the metadata in the converted format to the metadata index table according to the preset storage rules.

[0089] In addition, when the container constructs metadata, it is also possible that the metadata construction fails due to network problems, configuration errors, dependency problems, etc. To reduce the impact of metadata construction failure on metadata retrieval, after determining that the metadata construction fails, the server notifies the user and sends the error log to the user so that the user can analyze and handle the cause of the failure.

[0090] It should be noted that in one or more embodiments of this specification, there is no limitation on the specific container technology used to implement metadata construction. It can be implemented through docker, kubernetes, argo workflows, etc. Of course, it is also possible to use a combination of existing container technologies, or other technologies, which can be specifically set according to actual needs.

[0091] This specification also provides an embodiment of constructing metadata based on container technology to describe the metadata construction method based on container technology.

[0092] Specifically, as Figure 4 shown, Figure 4 is a schematic diagram of metadata construction and storage provided in this specification. The server responds to each metadata construction request, where each metadata construction request carries the data source type and storage address of the original data of the metadata to be constructed. Then, for each metadata construction request, the server sends the corresponding parameter information to the container according to the data source type of the original data in the metadata construction request, wakes up the client corresponding to the data source type in the called container, and loads the services required for the container metadata construction task through the awakened client in the container.

[0093] Then, the client of the container uses the parameter information passed into the container as the task data required for the container to execute the metadata construction task, starts running the metadata construction task, and then the container obtains the original data from the storage address based on the storage address of the original data to implement metadata construction according to the original data. If the parameter information also includes the data range of the original data, then when the container accesses the storage address to obtain the original data, it should only obtain the data within this data range as the original data.

[0094] When the server updates the metadata constructed by each container to the storage, the server can also compare the second identifier set with the first identifier set after decoding and decompressing in the storage to determine the updated metadata and the metadata to be deleted. And update, insert, and delete the metadata. And update the first identifier set, and at the same time, the updated first identifier set can be encoded and compressed.

[0095] If the parameter information includes an update frequency, the container periodically executes the metadata construction task based on the update frequency, and determines the unique identifier of the constructed metadata according to the unique identifier of the original data. Then, the constructed metadata and the corresponding unique identifier are updated to the storage.

[0096] The above is the metadata construction method based on container technology provided by the embodiments of this specification. Based on the same idea, this specification also provides a corresponding metadata construction device based on container technology, as Figure 5 shown.

[0097] A determination module 400 is configured to receive a metadata construction request for original data, and determine parameter information of the original data, where the parameter information at least includes a storage address of the original data;

[0098] A construction module 401 is configured to call a container for constructing metadata, input the parameter information into the container to wake up a client tool of the container, and based on a metadata processor called by the client tool, enable the container to obtain the original data and construct corresponding metadata according to the original data;

[0099] A storage module 402 is configured to update a metadata index table according to the metadata constructed by the container based on a preset storage rule.

[0100] Optionally, the parameter information further includes a data source type. The construction module 401 is configured to input the parameter information into the container, enable the container to determine the data source type of the original data according to the parameter information, and select a metadata processor corresponding to the data source type from the client tools of the container; start a metadata construction task corresponding to the container through the selected metadata processor, and enable the container to execute the metadata construction task to obtain the original data and construct corresponding metadata according to the original data.

[0101] Optionally, an update frequency of the metadata of the original data is carried in the metadata construction request. The construction module 401 is configured to input the parameter information and the update frequency into the container, enable the container to regularly obtain the original data from the storage address according to the input parameter information and construct metadata of the obtained original data according to the update frequency; the storage module 402 is configured to regularly update the metadata index table according to the constructed metadata based on a preset storage rule according to the update frequency.

[0102] Optionally, the storage module 402 is configured to determine a unique identifier of the metadata for container construction according to the original data; determine a first identifier set according to the unique identifiers of the metadata stored in the metadata index table; and update the metadata index table according to the preset storage rules based on the unique identifier and the stored first identifier set.

[0103] Optionally, the storage module 402 is configured to determine each container for constructing metadata; determine a second identifier set according to the unique identifier corresponding to the metadata constructed by each container; compare each unique identifier included in the second identifier set with each unique identifier in the first identifier set, and determine each unique identifier that is included in the first identifier set but not included in the second identifier set as the identifier to be deleted; delete the metadata corresponding to the identifier to be deleted from the metadata index table, update the first identifier set according to the second identifier set, and update the metadata corresponding to each unique identifier in the second identifier set to the metadata index table according to the preset storage rules.

[0104] Optionally, the storage module 402 is configured to determine the metadata for container construction and determine whether the metadata conforms to a preset format; if not, convert the metadata into the preset format.

[0105] Optionally, the storage module 402 is configured to, in response to a query instruction, determine retrieval information of the original data to be retrieved according to the query instruction; perform retrieval in the metadata index table according to the retrieval information, and output at least one key-value combination according to the retrieval result.

[0106] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the element.

[0107] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0108] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above Figure 1A method for constructing metadata based on container technology is provided.

[0109] This specification also provides Figure 6 A schematic structural diagram of the electronic device shown. As shown in FIG. 6, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 A method for constructing metadata based on container technology described. Of course, in addition to the software implementation method, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and can also be hardware or a logic device.

[0110] In the 1990s, it was obvious to distinguish whether an improvement in a technology was an improvement in hardware (e.g., improvement in circuit structures such as diodes, transistors, switches, etc.) or an improvement in software (improvement in method processes). However, with the development of technology, many improvements in method processes today can be regarded as direct improvements in hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structures by programming the improved method processes into the hardware circuits. Therefore, it cannot be said that an improvement in a method process cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. The designer can program by himself to "integrate" a digital system on a piece of PLD, without having to ask the chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a hardware description language (HDL), and there is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply making a little logical programming of the method process with the above-mentioned several hardware description languages and programming it into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method process.

[0111] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.

[0112] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0113] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0114] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0115] The present invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowchart illustrations and / or block diagrams, and combinations of flows and / or blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to the processors of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing device create means for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0116] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0117] These computer program instructions may also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0118] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0119] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.

[0120] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0121] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0122] It should be understood by those skilled in the art that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0123] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0124] The various embodiments in this specification are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and for related parts, reference can be made to the partial description of the method embodiments.

[0125] The above is only the embodiments of this specification and is not used to limit this specification. For those skilled in the art, various modifications and changes can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this application.

Claims

1. A metadata construction method based on container technology, characterized in that: include: Receive metadata construction requests for each piece of original data, and determine parameter information of the original data corresponding to each metadata construction request, wherein the parameter information at least includes a storage address and a data source type of the corresponding original data; For each metadata construction request, calling a metadata construction container for processing the metadata construction request, inputting the parameter information into the container, allowing the container to determine the data source type of the original data according to the parameter information, and selecting a metadata processor corresponding to the data source type from a client tool of the container; Starting the metadata construction task corresponding to the container through the selected metadata processor, so that the container executes the metadata construction task to obtain the original data and construct corresponding metadata according to the original data; Based on the metadata constructed by each container, the metadata index table is updated according to the preset storage rules.

2. The method according to claim 1, characterized in that: The metadata construction request carries the update frequency of the metadata of the original data, and the method further includes: Inputting the parameter information and the update frequency into the container, so that the container periodically obtains the original data from the storage address and constructs metadata of the obtained original data according to the input parameter information and the update frequency; Based on the metadata constructed by each container, the metadata index table is updated according to the preset storage rules, including: According to the update frequency, the metadata index table is updated regularly based on the constructed metadata and according to the preset storage rules.

3. The method according to claim 2, characterized in that: Based on the metadata constructed by each container, the metadata index table is updated according to the preset storage rules, including: Determine, based on the original data, a unique identifier of metadata constructed by the container; Determine a first identification set according to the unique identification of each metadata stored in the metadata index table; According to the unique identifier and the stored first identifier set, the metadata index table is updated according to a preset storage rule.

4. The method according to claim 3, characterized in that: Based on the metadata constructed by each container, the metadata index table is updated according to the preset storage rules, including: Determining containers for constructing metadata; Determine a second identifier set according to the unique identifier corresponding to the metadata constructed by each container; According to the comparison between each unique identifier included in the second identifier set and each unique identifier in the first identifier set, each unique identifier included in the first identifier set but not included in the second identifier set is determined as an identifier to be deleted; The metadata corresponding to the to-be-deleted identifier is deleted from the metadata index table, the first identifier set is updated according to the second identifier set, and the metadata corresponding to each unique identifier in the second identifier set is updated to the metadata index table according to a preset storage rule.

5. The method according to claim 1, characterized in that: According to the metadata constructed by each container, before updating the metadata index table according to the preset storage rule, the method further includes: Determining metadata constructed by the container, and judging whether the metadata conforms to a preset format; If not, the metadata is converted into the preset format.

6. The method according to claim 1, characterized in that The metadata is stored in the metadata index table in the form of key values, and the method further includes: In response to a query instruction, determining retrieval information of the original data to be retrieved according to the query instruction; A search is performed in the metadata index table according to the search information, and at least one key value combination is output according to the search result.

7. A metadata construction device based on container technology, characterized in that: include: A determination module, configured to receive metadata construction requests for each piece of original data, and determine parameter information of the original data corresponding to each metadata construction request, wherein the parameter information at least includes a storage address and a data source type of the original data; A construction module is used for calling a container for constructing metadata for processing each metadata construction request, inputting the parameter information into the container, so that the container determines the data source type of the original data according to the parameter information, and selecting a metadata processor corresponding to the data source type from a client tool of the container; Starting the metadata construction task corresponding to the container through the selected metadata processor, so that the container executes the metadata construction task to obtain the original data and construct corresponding metadata according to the original data; The storage module is used to update the metadata index table according to the metadata constructed by each container and the preset storage rules.

8. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Metadata storage method and device for unstructured data, medium and equipment

    CN117349401A

  • Data analysis method and device, computer equipment and storage medium

    CN117851463A