A data processing method and device, electronic equipment, and storage medium

By acquiring and classifying the file structure under the object storage path, a metadata set is generated, which solves the problem that traditional metadata management methods are difficult to meet management needs in data lakes, and realizes efficient metadata discovery and management.

CN116795785BActive Publication Date: 2026-01-23CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211370082.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-03
Publication Date
2026-01-23
Estimated Expiration
2042-11-03

AI Technical Summary

Technical Problem

Traditional metadata management methods are insufficient to meet the management needs of data assets in object-store-based data lakes. They lack metadata management methods for complex directory hierarchies and file formats, and cannot provide production-level metadata services.

Method used

By obtaining the file structure under the target object's storage path, traversing and classifying it, determining the file type, generating a metadata set, constructing metadata using the path list and file type information, and optimizing metadata management by combining machine learning techniques.

Benefits of technology

It improves metadata discovery efficiency, simplifies data management in data lakes, and meets the metadata discovery needs of object storage construction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116795785B_ABST
    Figure CN116795785B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a data processing method, comprising: obtaining a file structure under a target object storage path; traversing the file structure to obtain a first path list, wherein the first path list comprises at least one path from a root node to a target node in the file structure; determining a second path list based on the types of files in the first path list, wherein the files under each path in the second path list belong to the same type; determining a first metadata set based on the second path list to obtain metadata under the target object storage path. Embodiments of the present application also provide a data processing device, an electronic device and a storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of electronic devices, and relates to but is not limited to a data processing method and device, electronic device, and storage medium. BACKGROUND

[0002] A traditional metadata management method constructs a unified view of metadata through manual registration or automatic discovery, and in the traditional metadata management method, the source of metadata is mostly a database or a single file, and there is a lack of a method for managing metadata of an object storage complex directory hierarchy and file format, and it is difficult to meet the management needs of data assets in a data lake based on object storage and the needs of a data analysis engine at a higher level to provide production-level metadata services. SUMMARY

[0003] The present application provides a data processing method and device, electronic device, and storage medium.

[0004] The technical solution of the present application embodiment is as follows:

[0005] The present application provides a data processing method, which comprises: acquiring a file structure under a target object storage path; traversing the file structure to obtain a first path list, wherein the first path list comprises at least one path from a root node to a target node in the file structure; determining a second path list based on the types of files in the first path list, wherein the files under each path in the second path list belong to the same type; and determining a first metadata set based on the second path list to obtain metadata under the target object storage path.

[0006] The present application provides a data processing device, which comprises: an acquisition module configured to acquire a file structure under a target object storage path; a traversal module configured to traverse the file structure to obtain a first path list, wherein the first path list comprises at least one path from a root node to a target node in the file structure; a determination module configured to determine a second path list based on the types of files in the first path list, wherein the files under each path in the second path list belong to the same type; and determine a first metadata set based on the second path list to obtain metadata under the target object storage path.

[0007] The present application provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program capable of running on the processor, and the processor implements the steps in the above method when executing the program.

[0008] The embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the steps in the method.

[0009] The technical scheme provided by the embodiment of the present application has at least the following beneficial effects:

[0010] In the embodiment of the present application, the file structure under the target object storage path is obtained; the file structure is traversed to obtain a first path list, wherein the first path list includes at least one path from a root node to a target node in the file structure; based on the type of the file in the first path list, a second path list is determined, wherein the files under each path in the second path list belong to the same type; based on the second path list, a first metadata set is determined to obtain the metadata under the target object storage path. In this way, by using the file structure and the path list, the directory structure of the file storage can be determined, the metadata discovery demand in the data lake constructed based on the object storage is met, the metadata discovery efficiency is improved, and the data management in the data lake is simplified. BRIEF DESCRIPTION OF DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0012] Figure 1A A flowchart of a data processing method provided by the embodiment of the present application is shown in the figure;

[0013] Figure 1B A file structure provided by the present application is shown in the figure;

[0014] Figure 1C An application scenario diagram of a data processing method provided by the embodiment of the present application is shown in the figure;

[0015] Figure 1D An application scenario diagram of a data processing method provided by the embodiment of the present application is shown in the figure;

[0016] Figure 2 An optional architecture diagram of an execution system of a data processing method provided by the embodiment of the present application is shown in the figure;

[0017] Figure 3 A composition structure diagram of a data processing device provided by the embodiment of the present application is shown in the figure;

[0018] Figure 4A flowchart of a data processing method provided by an embodiment of the present application is shown in FIG. 1.

[0019] Figure 5 A component structure diagram of a data processing device provided by an embodiment of the present application is shown in FIG. 2.

[0020] Figure 6 A hardware entity diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 3. DETAILED DESCRIPTION

[0021] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application but not all embodiments of the present application. The following embodiments are used to describe the present application but not to limit the scope of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0022] In the following description, “some embodiments” are described, which describe a subset of all possible embodiments, but it can be understood that “some embodiments” can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0023] It should be noted that the terms “first\second\third” involved in the embodiments of the present application are only to distinguish similar objects and do not represent a specific order of the objects. It can be understood that “first\second\third” can be interchanged in a specific order or sequence as allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0024] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as generally understood by those skilled in the art to which the embodiments of the present application belong. It should also be understood that terms such as those defined in a general dictionary should be understood as having a meaning consistent with that in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as such herein.

[0025] With the development of big data, the data volume is expanding, and the data format is becoming more and more complex. How to ensure that the original information in the data is not lost, and to deeply mine the value of the data to meet the changing needs of the future, has become a common concern of enterprises. In this case, "data lake" has emerged. "Data lake" is a repository or system that stores data in its original format, allowing data to be stored according to the original information in the data without prior structured processing. A data lake can store structured data, unstructured data, and binary data, etc. Currently, the industry generally builds a unified data lake storage layer based on object storage to meet the needs of massive, diverse, and low-cost data storage. However, implementing unified data storage is far from enough. Unlike the data in the data warehouse, which is "pre-modeled", the data in the data lake is directly piled up without processing, so it is necessary to provide a unified metadata management service based on the data lake to efficiently support the rapid positioning and convenient application of massive data assets.

[0026] To help understand the technical solutions of the embodiments of the present application, the following introduces the terms involved in the embodiments of the present application:

[0027] Distributed storage: it is through the cooperative work of multiple servers, each server connects several physical media, and provides storage services for multiple systems together. In order to meet different access needs, a distributed storage system can provide file storage, block storage and object storage services at the same time. Among them, the data in block storage is accessed by bytes, and the data content and format of block storage are not defined, which is embodied as a volume or a hard disk. Block storage is responsible for data reading and writing, so the read-write performance is very high, and it is suitable for systems with high response time requirements, such as databases, etc. In file storage, data is stored and accessed in the form of files, and is organized according to the directory structure, which is embodied as a directory and a file, and is convenient for sharing, such as Network File System (NFS), File Transfer Protocol (FTP), etc. In object storage, data and metadata are packaged into a whole object and stored in a data pool, which is embodied as a Universally Unique Identifier (UUID). In the process of reading data, through a certain UUID, a whole object including data and metadata can be accessed. Object storage is suitable for video and image data with large data volume and fast increase speed.

[0028] File structure traversal methods include Depth-First Search (DFS) and Breadth-First Search (BFS). DFS traverses nodes based on their depth, prioritizing nodes generated later. BFS traverses nodes based on their depth, prioritizing nodes generated earlier. When there is more than one node at the bottom layer, the earliest generated node is selected for traversal.

[0029] Figure 1A A flowchart illustrating a data processing method provided in this application embodiment is shown in Figure 1. The method includes at least the following steps:

[0030] Step S101: Obtain the file structure under the target object storage path.

[0031] Here, the file structure may include a root directory, subdirectories under the root directory, and files under the root directory and subdirectories, wherein the root directory is the storage path of the target object.

[0032] Here, the target object storage path can be a path under a Windows system, such as "C:\Users\Downloads"; or a path under a Linux / Unix system, such as " / etc / systemd / system.conf".

[0033] Here, the root directory is represented as the root node in the file structure, and the subdirectories and files under the root directory and its subdirectories are represented as child nodes of the root node. For example, Figure 1B A schematic diagram of a file structure provided for this application; such as Figure 1B As shown, the root node 10 includes four child nodes Node11, Node12, Node13 and Node14; Node12 includes child nodes Node121 and Node122; Node13 includes child node Node131, and each child node includes at least one file.

[0034] Step S102: Traverse the file structure to obtain a first path list, wherein the first path list includes at least one path from the root node in the file structure to the target node.

[0035] Here, the traversal method for the file structure can be depth-first traversal or breadth-first traversal; no limitation is made here. For example, for... Figure 1BThe file structure shown is traversed to obtain a path list: <root.node11, root.node12.node121, root.node12.node122, root.node13, root.node13.node131, root.node14>.

[0036] In step S103, a second path list is determined based on the types of the files in the first path list; the files under each path in the second path list belong to the same type.

[0037] Here, the types of the files can include at least one of the following: JavaScript Object Notation (JSON), Comma-Separated Values (CSV), columnar storage format file Parquet, and eXtensible Markup Language (XML).

[0038] In step S104, a first metadata set is determined based on the second path list.

[0039] In one implementation manner, the step S104 includes: a step S1041 of performing file parsing on the files in the second path list to obtain first data information included in the files; a step S1042 of sampling the first data information to obtain second data information; and a step S1043 of determining the first metadata set based on the second data information.

[0040] Here, the sampling can be full-quantity collection or sampling collection, wherein the full-quantity collection scans full-quantity data of the first data information, takes a long time for a job, and has high result accuracy, and is generally used in a small data size scenario; the sampling collection scans part of the data of the first data information, takes less time for a job, and has low result accuracy, and is generally used in a large data size scenario.

[0041] In the above embodiment, a file structure under a target object storage path is acquired; the file structure is traversed to obtain a first path list, wherein the first path list includes at least one path from a root node to a target node in the file structure; based on the types of files in the first path list, a second path list is determined, wherein the files under each path in the second path list belong to the same type; based on the second path list, a first metadata set is determined to obtain metadata under the target object storage path. In this way, the directory structure of file storage can be determined through the file structure and the path list, the metadata discovery demand in the data lake constructed based on object storage is met, the metadata discovery efficiency is improved, and the data management in the data lake is simplified.

[0042] In an implementable manner, the step S1043 of determining the first metadata set based on the second data information includes: determining path information of the second data information; filtering the second data information with the same path information; and merging the filtered second data information based on the path information to obtain at least one table structure; wherein each field in each table structure is an element in the first metadata set.

[0043] For example, the path information with different file metadata under the same path information is further filtered, and the files with the same path information are merged, and finally each path is mapped to a table structure.

[0044] Here, the attribute information of the first metadata set includes file type, path depth, path information, and field information. Here, the file type can include four types of JSON, CSV, Parquet, and XML. Here, the first metadata set includes at least three types of metadata: field number, field name, and field type.

[0045] For example, as shown in Figure 1C A document element (meta) includes two table structures table1 and table2. Table1 is a table structure corresponding to the path information of root.node1.file1, which includes metadata of three field names a, b, and c, and the field type of a is integer (int), and the field types of b and c are string types. There is file3 under the root.node1 path, which includes metadata of three field names a, b, and c, but the field type of b is long. In this case, file3 includes different file metadata, and file3 needs to be filtered in the process of generating table1.

[0046] For example, as shown in Figure 1CAs shown, table2 is a table structure corresponding to the path information of root.node2.file2. When file4 of the csv type is included in root.node2, since it is not the same as the path information of file2, file4 needs to be filtered in the process of generating table2.

[0047] In another implementable manner, the step S104 of determining the first metadata set based on the second path list comprises: performing file parsing on the files in the second path list to obtain first data information included in the files; and deleting redundant metadata in the first data information to obtain the first metadata set.

[0048] Here, after the files in the second path list are parsed, data in all the files under the second path list, i.e., the first data information, can be obtained. Here, deleting the redundant metadata in the first data information can delete the metadata that is redundant in any one of the number of fields, the field name, and the field type in the first data information, reduce the number of metadata, and improve the metadata retrieval efficiency in the case of retrieving metadata under the second path list.

[0049] In one implementable manner, the method further comprises at least one of the following: a step S105 of performing path merging on the metadata in the first metadata set based on the path information of each metadata in the first metadata set to obtain at least one second metadata set;

[0050] a step S106 of performing semantic merging on the metadata in the first metadata set based on the semantic information of each metadata in the first metadata set to obtain at least one third metadata set.

[0051] In one implementable manner, the step S105 of performing path merging on the metadata in the first metadata set based on the path information of each metadata in the first metadata set to obtain at least one second metadata set comprises: a step S1051 of determining the path information of each metadata in the first metadata set; and a step S1052 of performing path merging on the metadata in the first metadata set based on the attribute information and the path information of each metadata to obtain at least one second metadata set; wherein the metadata in each second metadata set comprises the same path information.

[0052] Here, the path information can include parent node information and path depth. The attribute information of the metadata can include file type, field attribute. Illustratively, for metadata without path merging, traverse metadata with the same path depth and same direct parent node path as the metadata, and check whether the file type, field attribute, etc. of the metadata are consistent. If consistent, perform path merging, and record the merged partition information. For metadata that has already been merged, traverse other metadata with the same path depth and same direct parent node information as the metadata, and check whether the merged partition information is consistent. If consistent, continue to perform path merging, and record the merged partition information.

[0053] Illustratively, in the case of multiple files in the same path information, for example, multiple files under the A node and multiple files under the B node, the A node and the B node include the same path information, A and B are merged by path, and are allocated in the same partition, such as the partition_field shown in the figure. In this way, the files under the same path can be merged by path merging. Figure 1D

[0054] In an implementable manner, the step S106, based on the semantic information of each metadata in the first metadata set, performing semantic merging on the metadata in the first metadata set to obtain at least one third metadata set, includes: step S1061, obtaining scene information corresponding to the file structure under the target object storage path; step S1062, based on the scene information, determining semantic information of each metadata in the first metadata set in the scene information; step S1063, based on the semantic information, performing semantic merging on the metadata in the first metadata set to obtain at least one third metadata set.

[0055] Here, the scene information is business information, and the determination of the semantic information of each metadata in the first metadata set in the scene information based on the scene information includes: based on the business information, determining the business content of the field name of each metadata in the first metadata set.

[0056] Correspondingly, based on the semantic information, performing semantic merging on the metadata in the first metadata set to obtain at least one third metadata set includes: adding the same business label to the field name of the same business content; based on the business label, determining at least one third metadata set, wherein the metadata in each third metadata set includes the same business label.

[0057] ​For example, machine learning techniques such as clustering algorithms and association learning are used to detect the business content corresponding to field names (columns in the table structure) in object storage. Appropriate business tags are selected from a business tag library and applied to columns with the same business content. Columns with the same business content are merged to obtain a deduplicated third-party metadata set.

[0058] In the above embodiments, by semantically merging the metadata in the first metadata set based on the semantic information of each metadata in the first metadata set, at least one third metadata set is obtained. In this way, the metadata fields can be semantically merged, metadata redundancy can be reduced, and the metadata fields can have business attributes, so that the third metadata set is a metadata set that meets the business requirements.

[0059] Figure 2 A schematic diagram of an optional architecture for an execution system of a data processing method provided in an embodiment of this application; as shown. Figure 2 As shown, the architecture includes at least: a data development component 21, a metadata management component 22, a data storage component 23, and a data source component 24. The data source involved in this application is the object storage data source 241 in the data source component 24, and the component that implements the data processing method involved in this application is the auto-discovery component 221 in the metadata management component 22.

[0060] Figure 3 This is a schematic diagram of the composition structure of a data processing device provided in an embodiment of this application; as shown below. Figure 3 As shown, the device 300 includes: an object storage interface and a metadata discovery module, wherein the object storage interface is used to form an enterprise-level data catalog of discovered metadata for use by upper-layer data development; the metadata discovery module includes: a path exploration module 301, a file type identification and filtering module 302, a file sampling and parsing module 303, an initial metadata dataset generation module 304, an intelligent data detection module 305, and a metadata optimization and construction module 306, wherein:

[0061] The path exploration module 301 is used to traverse the object storage tree directory structure (file structure) with the input object storage path as the root node and output a list of available paths.

[0062] The file type identification and filtering module 302 is configured to identify file types and filter paths based on the identification results. The file type identification mainly takes a path list as input, identifies file types in each path, including JSON, CSV, Parquet and XML; the metadata discovery module defaults that all files in the same path belong to one table, so the files in the same path need to have the same file type, and the path filtering module is configured to filter paths with different file types in the same path based on the above identification results, and output a filtered path list (second path list).

[0063] The file sampling and parsing module 303 is configured to sample and parse files. The file sampling is implemented by a file sampler, which supports full collection or sampling collection of files according to the configuration. The full collection mode scans full data of files (first data information), which takes a long time and has high accuracy, and is generally used in small data scenarios. The sampling collection mode only scans part of the data of files, which takes less time and has lower accuracy, and is generally used in large data scenarios. The file parsing is implemented by a file parser, which parses the output of the file sampler. The parser provides JSON, CSV, Parquet and XML format parsing capabilities, and outputs a first metadata set, including field number, field name and field type. The metadata discovery module defaults that all files in the same path belong to one table, so the files in the same path need to have the same file type, field number, field name and field type.

[0064] The initial metadata generation module 304 is configured to integrate the output of the file type identification and filtering module and the file sampling and parsing module.

[0065] The intelligent data detection 305 is configured to detect the specific business content of the file columns (field names) in the object storage by machine learning techniques such as clustering algorithm and association learning, and select appropriate business tags from the business tag library and apply them to the columns with the same business content.

[0066] The metadata optimization and construction module 306 is configured to merge partitions and optimize based on semantic information. As shown in FIG. 6, the partition merging includes: Figure 4

[0067] In step S401, the partition state of the first metadata set is determined.

[0068] Here, any metadata of the initial metadata set (first metadata set) is selected, the path depth and path details are recorded, and then it is checked whether the metadata has been merged into a partition, to obtain the partition state.

[0069] ​In step S402, the merging partition is performed based on the partition state, and partition information is obtained.

[0070] Here, the partition information can include a partition level, metadata attributes, and the like.

[0071] For metadata that has not been subjected to path merging, metadata with the same path depth and direct parent node information is traversed to check whether the file type (meta[j]_file value), field attribute (meta[j]_value value), and the like are consistent. If consistent, path merging is performed, and the merged partition information is recorded. For metadata that has been subjected to path merging, metadata with the same path depth and direct parent node information is traversed to check whether the merged partition information is consistent. If consistent, path merging is continued, and the merged partition information is recorded.

[0072] In step S403, a second metadata set is determined according to the partition information. Each element in the second metadata set includes the same path information.

[0073] Based on the foregoing embodiments, an embodiment of the present application further provides a data processing apparatus. The control apparatus includes various modules, and can be implemented by a processor in an electronic device. Of course, the apparatus can also be implemented by a specific logic circuit. In the implementation process, the processor can be a central processing unit (CPU), a micro processing unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), and the like.

[0074] Figure 5 A component structure diagram of a data processing apparatus provided by an embodiment of the present application is shown in FIG. 5. As shown in FIG. 5, the apparatus 500 includes an acquisition module 501, a traversal module 502, and a determination module 503. In this embodiment, the acquisition module 501 is configured to acquire a file structure under a storage path of a target object. Figure 5 The acquisition module 501 is configured to acquire a file structure under a storage path of a target object.

[0075] The traversal module 502 is configured to traverse the file structure to obtain a first path list. The first path list includes at least one path from a root node to a target node in the file structure.

[0076]

[0077] ​The determining module 503 is configured to determine a second path list based on the types of the files in the first path list, wherein the files in each path in the second path list belong to the same type; and determine a first metadata set based on the second path list to obtain the metadata in the target object storage path.

[0078] In some possible embodiments, the determining module 503 is further configured to perform file analysis on the files in the second path list to obtain first data information included in the files; sample the first data information to obtain second data information; and determine the first metadata set based on the second data information.

[0079] In some possible embodiments, the determining module 503 is further configured to determine path information of the second data information; filter the second data information with the same path information; and merge the filtered second data information based on the path information to obtain at least one table structure; wherein each field in each table structure is an element in the first metadata set.

[0080] In some possible embodiments, the apparatus 500 further includes a merging module configured to at least one of: merge the metadata in the first metadata set based on path information of each metadata in the first metadata set to obtain at least one second metadata set; and merge the metadata in the first metadata set based on semantic information of each metadata in the first metadata set to obtain at least one third metadata set.

[0081] In some possible embodiments, the merging module is further configured to determine the path information of each metadata in the first metadata set; and merge the metadata in the first metadata set based on attribute information and the path information of each metadata to obtain at least one second metadata set; wherein the metadata in each second metadata set includes the same path information.

[0082] In some possible embodiments, the merging module is further configured to obtain scene information corresponding to a file structure in the target object storage path; determine semantic information of each metadata in the first metadata set in the scene information based on the scene information; and merge the metadata in the first metadata set based on the semantic information to obtain at least one third metadata set.

[0083] In some possible embodiments, the scene information is business information, and the merging module is further configured to: determine, based on the business information, business content of a field name of each piece of metadata in the first set of metadata; add a same business label to field names of a same business content; and determine at least one third set of metadata based on the business label, wherein each piece of metadata in each of the third sets of metadata comprises the same business label.

[0084] It should be noted that the above description of the device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects to the method embodiments. For technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0085] It should be noted that, in the embodiments of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, and includes a plurality of instructions for causing an electronic device to execute all or part of the methods described in the embodiments of the present application. The storage medium mentioned above includes: a U disk, a mobile hard disk, a read only memory (ROM), a magnetic disk or an optical disk, and various media that can store program codes. Thus, the embodiments of the present application are not limited to any specific hardware and software combination.

[0086] Correspondingly, the embodiments of the present application provide a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the steps of the method described in any of the above embodiments.

[0087] Correspondingly, the embodiments of the present application also provide a chip, which includes a programmable logic circuit and / or program instructions. When the chip is running, it is used to implement the steps of the method described in any of the above embodiments.

[0088] Correspondingly, the embodiments of the present application also provide a computer program product. When the computer program product is executed by a processor of an electronic device, it is used to implement the steps of the method described in any of the above embodiments.

[0089] Based on the same technical concept, the embodiments of the present application provide an electronic device for implementing the data processing method described in the above method embodiments. Figure 6 A hardware entity schematic diagram of an electronic device provided in the embodiments of the present application is shown in Figure 6As shown, the electronic device 600 includes a memory 610 and a processor 620, the memory 610 stores a computer program executable on the processor 620, and the processor 620 implements the steps in the method of any of the embodiments of the present application when executing the program.

[0090] The memory 610 is configured to store instructions executable by the processor 620 and applications, and can also cache data to be processed by the processor 620 and modules in the electronic device (for example, image data, audio data, voice communication data and video communication data) to be processed or having been processed, which can be implemented by FLASH or Random Access Memory (RAM).

[0091] The processor 620 implements the steps of any of the above data processing methods when executing the program. The processor 620 generally controls the overall operation of the electronic device 600.

[0092] The above processor can be at least one of an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a Central Processing Unit (CPU), a controller, a microcontroller, or a microprocessor. It can be understood that the electronic device implementing the above processor functions can also be other, and the embodiments of the present application are not limited specifically.

[0093] The computer storage medium / memory can be a Read Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electrically Erasable Programmable Read-Only Memory (EEPROM), a Ferromagnetic Random Access Memory (FRAM), a Flash Memory, a magnetic surface storage, an optical disc, a Compact Disc Read-Only Memory (CD-ROM), or the like memory; or can be various electronic devices including one or any combination of the above memories, such as a mobile phone, a computer, a tablet device, a personal digital assistant, and the like.

[0094] It should be noted that the above description of the storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0095] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that the size of the sequence number of the above processes in various embodiments of the present application does not mean the execution order, and the execution order of the processes should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above sequence number of the embodiments of the present application is only for description, not representing the advantages and disadvantages of the embodiments.

[0096] It should be noted that, in the present document, the terms "comprising", "containing", or any other similar term are intended to encompass non-exclusive inclusions, such that a process, method, article, or apparatus that comprises a list of elements does not necessarily include those elements only, but can include other elements not expressly listed, or can include elements inherent in such process, method, article, or apparatus. Without more limitations, an element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0097] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative, for example, the division of the units is only a logical functional division, and actual implementation can have another division manner, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed components can be through some interface, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0098] The units described above as separate components can or can not be physically separated, and the components shown as units can or can not be physical units; they can be located in one place or distributed on multiple network units; and part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments of the present application.

[0099] In addition, each functional unit in each embodiment of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be realized in the form of hardware or in the form of hardware plus software functional unit.

[0100] Alternatively, the integrated unit of the present application, if implemented in the form of a software function module and sold or used as an independent product, can also be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the embodiments of the present application can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for causing an apparatus to perform all or part of the methods described in the embodiments of the present application. The foregoing storage medium includes: mobile storage devices, ROM, magnetic or optical disks, and various other media that can store program codes.

[0101] The methods disclosed in the several method embodiments provided by the present application can be combined arbitrarily without conflict to obtain new method embodiments.

[0102] The features disclosed in the several method or device embodiments provided by the present application can be combined arbitrarily without conflict to obtain new method embodiments or device embodiments.

[0103] The above merely provides the implementation manners of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of the changes or replacements within the technical range disclosed by the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A data processing method, characterized in that, The method includes: Obtain the file structure under the target object's storage path; The file structure is traversed to obtain a first path list, wherein the first path list includes at least one path from the root node in the file structure to the target node; Based on the types of files contained in the first path list, a second path list is determined, wherein the files under each path in the second path list belong to the same type. Based on the second path list, a first metadata set is determined to obtain the metadata under the target object storage path; The step of determining the first metadata set based on the second path list includes: The files in the second path list are parsed to obtain the first data information contained in the files; The first data information is sampled to obtain the second data information; Determine the path information of the second data information; Filter the second data information with the same path information; Based on the path information, the filtered second data information is merged to obtain at least one table structure; wherein each field in each table structure is an element in the first metadata set.

2. The method as described in claim 1, characterized in that, The method further includes at least one of the following: Based on the path information of each metadata in the first metadata set, the metadata in the first metadata set is merged by path to obtain at least one second metadata set; Based on the semantic information of each metadata in the first metadata set, the metadata in the first metadata set is semantically merged to obtain at least one third metadata set.

3. The method as described in claim 2, characterized in that, The step of merging the paths of metadata in the first metadata set based on the path information of each metadata in the first metadata set to obtain at least one second metadata set includes: Determine the path information for each metadata in the first metadata set; Based on the attribute information and path information of each metadata, the metadata in the first metadata set is merged by path to obtain at least one second metadata set; wherein the metadata in each second metadata set includes the same path information.

4. The method as described in claim 2, characterized in that, The semantic merging of metadata in the first metadata set based on the semantic information of each metadata in the first metadata set to obtain at least one third metadata set includes: Obtain scene information corresponding to the file structure under the storage path of the target object; Based on the scene information, determine the semantic information of each metadata in the first metadata set within the scene information; Based on the semantic information, the metadata in the first metadata set is semantically merged to obtain at least one third metadata set.

5. The method as described in claim 4, characterized in that, The scenario information is business information. The step of determining the semantic information of each metadata in the first metadata set in the scenario information based on the scenario information includes: determining the business content of the field name of each metadata in the first metadata set based on the business information. Correspondingly, based on the semantic information, the metadata in the first metadata set is semantically merged to obtain at least one third metadata set, including: adding the same business tag to the field names of the same business content; and determining at least one third metadata set based on the business tag, wherein the metadata in each third metadata set includes the same business tag.

6. A data processing apparatus, characterized in that, The device includes: The acquisition module is used to obtain the file structure under the storage path of the target object; The traversal module is used to traverse the file structure to obtain a first path list, wherein the first path list includes at least one path from the root node in the file structure to the target node. The determining module is configured to determine a second path list based on the types of files in the first path list, wherein the files under each path in the second path list belong to the same type; and to determine a first metadata set based on the second path list to obtain the metadata under the target object storage path. The determining module is further configured to parse the files in the second path list to obtain first data information included in the files; sample the first data information to obtain second data information; determine the path information of the second data information; filter the second data information with the same path information; and merge the filtered second data information based on the path information to obtain at least one table structure; wherein each field in each table structure is an element in the first metadata set.

7. An electronic device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Metadata clustering management method and module applied to distributed file system

    CN103198153A

  • Big data environment oriented metadata organization method and system

    CN105550371A