Unstructured Data Lake Inflow Method, Apparatus, Electronic Device, and Storage Medium
By generating extended metadata and performing data retrieval and positioning operations, the problems of slow metadata retrieval and insufficient data validity in the unstructured data entry method are solved, and efficient data management and real-time guarantees are achieved.
Patent Information
- Application Number
- CN202211208748.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-09-30
AI Technical Summary
The existing unstructured data entry method has problems such as slow metadata retrieval, inability to modify, not supporting fuzzy search, failure to write data, and lack of data validity guarantee, resulting in inadequate management efficiency.
By generating extended metadata, including data key values, data versions and data validity, data retrieval checks and positioning operations, ensuring the effectiveness management of data before storage, and realizing data uniqueness and version evolution.
It improves the efficiency of unstructured data management, ensures the real-time and standardization of data, avoids data duplication and errors, and supports data writing, updating and deletion operations.
Smart Images

Figure CN115543198B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method, device, electronic device and computer-readable storage medium for unstructured data to enter a data lake. Background Art
[0002] Currently, there are mainly two ways to store unstructured data:
[0003] The first way is to write the original unstructured data file into the distributed storage system of the data lake. The disadvantage of this way is that it lacks the metadata information of the file (basic information of the file, such as file ownership, file usage scope, file description information, etc.), and the file cannot be found through the metadata, which is likely to form a "data swamp".
[0004] The second way is to separately enter the original unstructured file and its metadata into the data lake. The file is written into the distributed file system HDFS, and the metadata is written into the Hive data warehouse and the file storage path is saved. The metadata information is retrieved through HiveSQL to find the file storage address. This way also has the following disadvantages:
[0005] 1. The retrieval of metadata by Hive is slow;
[0006] 2. The metadata information cannot be modified;
[0007] 3. Fuzzy search of metadata is not supported;
[0008] 4. The unstructured file and the metadata are separately written into different storage systems of the data lake, and the prior art lacks data validity guarantee. For example, writing the metadata is successful while writing the file fails, resulting in data errors;
[0009] 5. There are problems such as the lack of a standard for structured data to enter the data lake and difficult data management.
[0010] Therefore, it is urgent to improve the method for unstructured data to enter the data lake and enhance the management efficiency of unstructured data. Summary of the Invention
[0011] The present invention provides a method, device, electronic device and computer-readable storage medium for unstructured data to enter a data lake, and its main purpose is to enhance the management efficiency of unstructured data.
[0012] To achieve the above object, a method for unstructured data to enter a data lake provided by the present invention includes:
[0013] Obtain a data processing request, and parse the data processing request to obtain the data to be processed, the metadata to be processed and the processing type;
[0014] Generate extended metadata of the data to be processed, wherein the extended metadata includes a data key value, a data version and data validity;
[0015] When the processing type is write, use the data key value and the data version to perform data duplicate check on the existing data in the preset storage system. When the data duplicate check passes, store the data to be processed in the data storage space of the preset storage system, store the corresponding metadata to be processed and the extended metadata in the metadata storage space of the preset storage system, and set the data validity of the data to be processed as valid.
[0016] When the processing type is change or delete, use the data key value and the data version to locate the existing target object in the preset storage system, set the data validity of the target object to invalid, store the data to be processed with the processing type of change in the data storage space of the preset storage system and store the corresponding metadata to be processed and the extended metadata in the metadata storage space of the preset storage system, set the data validity of the data to be processed as valid, and delete the target object with the processing type of delete.
[0017] Optionally, generating the extended metadata of the data to be processed, where the extended metadata includes a data key value, a data version, and data validity, includes:
[0018] Calculate the hash value of the data to be processed, and use the hash value as the data key value of the data to be processed;
[0019] Determine whether the data key value is unique in the preset storage system;
[0020] When the data key value is unique, set the data version corresponding to the data to be processed as the first version;
[0021] When the data key value is not unique, obtain the highest data version corresponding to the data key value, and obtain the data version corresponding to the data to be processed by adding 1 to the highest data version;
[0022] Initialize the data validity corresponding to the data to be processed as invalid;
[0023] Collect the data key value, the data version, and the data validity to obtain the extended metadata corresponding to the data to be processed.
[0024] Optionally, the using the data key value and the data version to perform data duplicate check on the existing data in the preset storage system includes:
[0025] Query in the preset storage system for a data key value that matches the data key value or a data version that matches the data version or a combination of both;
[0026] In the preset storage system, when no data key value matching the data key value is found, a message indicating that the data duplicate check passes is returned;
[0027] In the preset storage system, when a data key value matching the data key value is found but no data version matching the data version is found, a message indicating that the data duplicate check passes is returned;
[0028] In the preset storage system, when a data key value matching the data key value is found and a data version matching the data version is found, a message indicating that the data duplicate check fails is returned.
[0029] Optionally, the positioning of the existing target object in the preset storage system by using the data key value and the data version includes:
[0030] Using the data key value and the data version corresponding to the data to be processed as an index;
[0031] In the preset storage system, querying for a data key value and a data version matching the index;
[0032] Taking the data, metadata, and extended metadata corresponding to the matched data key value and data version as the target object.
[0033] Optionally, after storing the data to be processed in the data storage space of the preset storage system and storing the corresponding metadata to be processed and the extended metadata in the metadata storage space of the preset storage system, the method further includes:
[0034] Respectively obtaining the data storage address of the data to be processed in the data storage space and the metadata storage addresses of the corresponding metadata to be processed and the extended metadata in the metadata storage space;
[0035] Creating a mapping relationship between the data storage address and the metadata storage address according to the correspondence between the data to be processed and the metadata to be processed and the extended metadata.
[0036] Optionally, after returning a message indicating that the data duplicate check fails when a data key value matching the data key value is found and a data version matching the data version is found in the preset storage system, the method includes:
[0037] Rejecting the data processing request.
[0038] To solve the above problems, the present invention further provides an unstructured data lake device, and the device includes:
[0039] A data processing request parsing module, configured to obtain a data processing request, and parse the data processing request to obtain data to be processed, metadata to be processed, and a processing type;
[0040] An extended metadata generation module, configured to generate extended metadata for the data to be processed, where the extended metadata includes a data key value, a data version, and data validity;
[0041] A data writing module, configured to, when the processing type is writing, use the data key value and the data version to perform data duplicate check verification on existing data in a preset storage system. When the data duplicate check verification passes, store the data to be processed in a data storage space of the preset storage system, store the corresponding metadata to be processed and the extended metadata in a metadata storage space of the preset storage system, and set the data validity of the data to be processed to valid;
[0042] A data update and deletion module, configured to, when the processing type is change or deletion, use the data key value and the data version to locate an existing target object in the preset storage system, set the data validity of the target object to invalid, store the data to be processed with the processing type of change in the data storage space of the preset storage system and store the corresponding metadata to be processed and the extended metadata in the metadata storage space of the preset storage system, set the data validity of the data to be processed to valid, and delete the target object with the processing type of deletion.
[0043] Optionally, the data update and deletion module locates the target object through the following operations:
[0044] Locating an existing target object in the preset storage system by using the data key value and the data version includes:
[0045] Using the data key value and the data version corresponding to the data to be processed as an index;
[0046] In the preset storage system, query data key values and data versions that match the index;
[0047] Using the data, metadata, and extended metadata corresponding to the matched data key values and data versions as the target object.
[0048] To solve the above problems, the present invention further provides an electronic device, where the electronic device includes:
[0049] A memory, storing at least one computer program; and
[0050] A processor that executes the program stored in the memory to implement the unstructured data lake-in method described above.
[0051] To solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one computer program is stored, and the at least one computer program is executed by a processor in an electronic device to implement the unstructured data lake-in method described above.
[0052] In the embodiment of the present invention, the data key values and data versions of the extended metadata are used to perform duplicate checking and positioning operations on the existing data in the preset storage system, and to manage the evolution of the data versions in the preset storage system. Before processing the data to be processed, the data validity of the existing data corresponding to the data to be processed in the preset storage system is set to invalid. When the writing and updating of the data to be processed are completed, the data validity of the data to be processed is set to valid, which can ensure the real-time nature of the unstructured data in the preset storage system. Through the above means, the standardized management of unstructured data is realized, and the management efficiency of unstructured data is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a schematic flowchart of the unstructured data lake-in method provided by an embodiment of the present invention;
[0054] Figure 2 It is a detailed implementation flowchart of one of the steps in the unstructured data lake-in method provided by an embodiment of the present invention;
[0055] Figure 3 It is a detailed implementation flowchart of one of the steps in the unstructured data lake-in method provided by an embodiment of the present invention;
[0056] Figure 4 It is a functional module diagram of the unstructured data lake-in device provided by an embodiment of the present invention;
[0057] Figure 5 It is a schematic structural diagram of an electronic device for implementing the unstructured data lake-in method provided by an embodiment of the present invention.
[0058] The realization, functional characteristics and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0060] An embodiment of the present application provides a method for unstructured data to enter the lake. The execution subject of the method for unstructured data to enter the lake includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided in the embodiment of the present application. In other words, the method for unstructured data to enter the lake can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0061] Referring to Figure 1 As shown, it is a schematic flowchart of the method for unstructured data to enter the lake provided by an embodiment of the present invention. In this embodiment, the method for unstructured data to enter the lake includes:
[0062] S1. Obtain a data processing request, and parse the data processing request to obtain the data to be processed, the metadata to be processed, and the processing type;
[0063] In the embodiment of the present invention, the data processing request refers to an operation request for unstructured data initiated through a preset API interface, including a write request, a change request, a delete request, etc. for unstructured data. The preset API interface refers to a standardized data interface for users generated according to the operation specifications of unstructured data defined according to actual business needs.
[0064] The processing type refers to the data operation requirements to be realized by the data processing request. For example, writing the data to be processed or the metadata to be processed into a specified storage space of the system, or performing corresponding update, delete, etc. operations on the data already stored in the system.
[0065] In the embodiment of the present invention, the data to be processed is unstructured data. Unstructured data is relative to structured data. Generally, unstructured data can be text files such as images, videos, audios, documents, reports, etc., and audio files. Unstructured data with diverse data formats cannot be expressed and stored using the conventional two-dimensional table structure in structured data.
[0066] In the embodiment of the present invention, the metadata to be processed refers to the metadata corresponding to the data to be processed. For example, if the data to be processed is an image, the metadata to be processed is the information described around the image, such as the image name, the image generation time, the image pixel information, the image storage path, etc. These information constitute the metadata of the image.
[0067] In an embodiment of the present invention, the to-be-processed data, the to-be-processed metadata, and the processing type can be identified from the data processing request according to the data specifications of a preset API interface.
[0068] S2. Generate extended metadata for the to-be-processed data, where the extended metadata includes a data key value, a data version, and data validity.
[0069] In an embodiment of the present invention, the extended metadata is data that provides an overall description of the to-be-processed metadata and the to-be-processed data.
[0070] Among them, the data key value is data used to identify the uniqueness of the to-be-processed data and the to-be-processed metadata, and can be represented using a globally unique serial number.
[0071] The data version is a means for version management of the to-be-processed data and the to-be-processed metadata. By defining different version numbers for different data contents of the same data object, a version evolution relationship between the data contents of the same data object at different times is formed.
[0072] The data validity generally includes two values, namely valid or invalid. When the value corresponding to the data validity is valid, it indicates that the current data object is available; when the value corresponding to the data validity is invalid, it indicates that the current data object is unavailable.
[0073] In an embodiment of the present invention, the extended metadata may further include information such as the creation time, creator, last modification time, data source, association relationship with other data, data storage path, data type, or data format corresponding to the to-be-processed metadata and the to-be-processed data.
[0074] Specifically, referring to Figure 2 as shown, the generation of the extended metadata for the to-be-processed data, where the extended metadata includes a data key value, a data version, and data validity, includes:
[0075] S21. Calculate the hash value of the to-be-processed data, and use the hash value as the data key value of the to-be-processed data.
[0076] S22. Determine whether the data key value is unique in a preset storage system.
[0077] When the data key value is unique, then execute S23. Set the data version corresponding to the to-be-processed data to the first version.
[0078] When the data key value is not unique, execute S24, obtain the highest data version corresponding to the data key value, and obtain the data version corresponding to the data to be processed by adding 1 to the highest data version;
[0079] S25. Initialize the data validity corresponding to the data to be processed as invalid;
[0080] S26. Aggregate the data key value, the data version, and the data validity to obtain the extended metadata corresponding to the data to be processed.
[0081] In an embodiment of the present invention, the preset storage system refers to an unstructured data storage space pre-constructed or pre-specified according to actual business requirements, and may be a cloud storage server or a blockchain storage system.
[0082] In an embodiment of the present invention, only after writing the data to be processed into the preset storage system or completing the corresponding data update operation according to the data to be processed, can the data validity corresponding to the data to be processed be set to valid, and only by doing so can it conform to the result of the actual data operation.
[0083] In another optional embodiment of the present invention, the constituent fields, the definition of each field, and the acquisition method of each field of the extended metadata can be determined according to the pre-constructed metadata extension specification, and then the corresponding extended metadata can be generated for different data to be processed.
[0084] In an embodiment of the present invention, through the extended metadata, the uniqueness, version evolution, and data validity of the data to be processed and the metadata to be processed can be managed, and the management efficiency of unstructured data can be improved.
[0085] S3. When the processing type is writing, use the data key value and the data version to perform data duplicate check verification on the existing data in the preset storage system. When the data duplicate check verification passes, store the data to be processed in the data storage space of the preset storage system, store the corresponding metadata to be processed and the extended metadata in the metadata storage space of the preset storage system, and set the data validity of the data to be processed to valid;
[0086] In an embodiment of the present invention, different unstructured data processing operations are adopted for different data processing types.
[0087] In an embodiment of the present invention, the duplicate check verification of the existing data is realized by querying the data key value or the data version that matches the data key value or the mutually matching data version in the preset storage system.
[0088] Specifically, the data duplicate check and verification of the existing data in the preset storage system using the data key value and the data version includes:
[0089] Query in the preset storage system for a data key value that matches the data key value or the data version, or for mutually matching data versions;
[0090] In the preset storage system, when no data key value that matches the data key value is found, return a message indicating that the data duplicate check and verification has passed;
[0091] In the preset storage system, when a data key value that matches the data key value is found, but no data version that matches the data version is found, return a message indicating that the data duplicate check and verification has passed;
[0092] In the preset storage system, when a data key value that matches the data key value is found and a data version that matches the data version is found, return a message indicating that the data duplicate check and verification has failed.
[0093] In the embodiments of the present invention, the data query verification passing includes two cases:
[0094] When in the preset storage system, no data key value that matches the data key value is found, it means that the unstructured data corresponding to the data key value is not included in the preset storage system.
[0095] When in the preset storage system, a data key value that matches the data key value is found, but no data version that matches the data version is found, it means that the unstructured data stored in the preset storage system is the same as the data key value, but the data version corresponding to the stored unstructured data is different from the data version corresponding to the data to be processed.
[0096] In the above two cases, it indicates that the data to be processed does not conflict with the data stored in the preset storage system, and the data to be processed and the metadata to be processed can be stored in the preset storage system.
[0097] Correspondingly, the data query verification failure means that when both the data key value and the data version already exist. This situation indicates that the unstructured data stored in the preset storage system and the metadata corresponding to the unstructured data are the same as the data to be processed and the metadata to be processed, and the data to be processed and the metadata to be processed cannot be stored in the preset storage system, otherwise there will be a situation of duplicate data loading. Therefore, the data processing request needs to be rejected.
[0098] In an embodiment of the present invention, the preset storage system includes a preset data storage space and a preset metadata storage space, which are respectively used for storing unstructured data and the metadata corresponding to the unstructured data. Separating the storage of unstructured data and metadata can achieve the independence of the write, read, update, or delete operations of the two, and improve data management efficiency.
[0099] In another alternative embodiment of the present invention, after storing the data to be processed in the data storage space of the preset storage system and storing the corresponding metadata to be processed and the extended metadata in the metadata storage space of the preset storage system, the method further includes:
[0100] Obtain the data storage address of the data to be processed in the data storage space and the metadata storage addresses of the corresponding metadata to be processed and the extended metadata in the metadata storage space respectively;
[0101] Create a mapping relationship between the data storage address and the metadata storage address according to the correspondence between the data to be processed, the metadata to be processed, and the extended metadata.
[0102] In an embodiment of the present invention, according to the mapping relationship, corresponding data, metadata, and extended metadata can be quickly queried in the data storage space or the metadata storage space.
[0103] In an embodiment of the present invention, after the data to be processed, the metadata to be processed, and the extended metadata are written, the data validity in the extended metadata can be set to valid, indicating that the newly written data is available, which is in line with the actual situation of data operation.
[0104] S4. When the processing type is change or deletion, use the data key value and the data version to locate the existing target object in the preset storage system, set the data validity of the target object to invalid, store the data to be processed with the processing type of change in the data storage space of the preset storage system and store the corresponding metadata to be processed and the extended metadata in the metadata storage space of the preset storage system, set the data validity of the data to be processed to valid, and delete the target object with the processing type of deletion.
[0105] It can be understood that update or delete operations usually modify or delete existing data or existing metadata in the preset storage system.
[0106] In an embodiment of the present invention, the target object includes target data, target metadata corresponding to the target data, and target extended metadata.
[0107] Specifically, refer to Figure 3 shown, the positioning of the existing target object in the preset storage system by using the data key value and the data version includes:
[0108] S41. Use the data key value and the data version corresponding to the data to be processed as an index;
[0109] S42. In the preset storage system, query the data key value and data version that match the index;
[0110] S43. Use the data, metadata, and extended metadata corresponding to the matched data key value and data version as the target object.
[0111] In the embodiment of the present invention, the data key value and the data version corresponding to the data to be processed are used as an index, and the target extended metadata that matches the data key value and the data version is queried through this index. Then, based on the target extended metadata, the target data and target metadata corresponding to the target extended metadata are located to obtain the target object.
[0112] In the embodiment of the present invention, before updating or deleting the target object, it is necessary to set the data validity in the extended metadata corresponding to the target object to invalid, indicating that the target object that needs to be changed currently is unavailable.
[0113] When the processing type is update, after storing the data to be processed, the metadata to be processed, and the extended metadata in a preset data storage space or a preset metadata storage space, set the data validity field in the extended metadata to valid, indicating that the currently available data is the data newly written into the data storage space and the metadata corresponding to the data.
[0114] When the processing type is deletion, delete the target data, the target metadata, and the target extended metadata.
[0115] In the embodiment of the present invention, during the process of updating or deleting the target data and the target metadata, if an exception occurs, resulting in the failure of the update operation or the deletion operation, the data validity in the target extended metadata corresponding to the target object can be restored to valid, that is, the availability of the existing target object is restored.
[0116] In the embodiments of the present invention, the data key value and data version of the extended metadata are used to perform duplicate checking and positioning operations on the existing data in the preset storage system, and to manage the evolution of the data version in the preset storage system. Before processing the data to be processed, the data validity of the existing data corresponding to the data to be processed in the preset storage system is set to invalid. When the writing and updating of the data to be processed are completed, the data validity of the data to be processed is set to valid, which can ensure the real-time nature of the unstructured data in the preset storage system. Through the above means, the standardized management of unstructured data is realized, and the management efficiency of unstructured data is improved.
[0117] As Figure 4 shown, it is a functional module diagram of an unstructured data lake-in device provided by an embodiment of the present invention.
[0118] The unstructured data lake-in device 100 of the present invention can be installed in an electronic device. According to the functions implemented, the unstructured data lake-in device 100 may include a data processing request parsing module 101, an extended metadata generation module 102, a data writing module 103, and a data update and deletion module 104. The modules of the present invention may also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.
[0119] In this embodiment, the functions of each module / unit are as follows:
[0120] The data processing request parsing module 101 is used to obtain a data processing request, and parse the data processing request to obtain the data to be processed, the metadata to be processed, and the processing type;
[0121] The extended metadata generation module 102 is used to generate the extended metadata of the data to be processed, where the extended metadata includes a data key value, a data version, and data validity;
[0122] The data writing module 103 is used to, when the processing type is writing, perform data duplicate checking on the existing data in the preset storage system by using the data key value and the data version. When the data duplicate checking passes, store the data to be processed in the data storage space of the preset storage system, store the corresponding metadata to be processed and the extended metadata in the metadata storage space of the preset storage system, and set the data validity of the data to be processed to valid;
[0123] The data update and deletion module 104 is used to, when the processing type is change or deletion, locate the existing target object in the preset storage system by using the data key value and the data version, set the data validity of the target object to invalid, store the to-be-processed data with the processing type of change into the data storage space in the preset storage system and store the corresponding to-be-processed metadata and the extended metadata into the metadata storage space in the preset storage system, set the data validity of the to-be-processed data to valid, and delete the target object with the processing type of deletion.
[0124] Specifically, each module in the unstructured data lake-in device 100 in the embodiments of the present invention adopts the same technical means as those in the Figures 1 to 3 unstructured data lake-in method described above and can produce the same technical effects, which will not be elaborated here.
[0125] As Figure 5 shown, it is a schematic structural diagram of an electronic device for implementing the unstructured data lake-in method provided by an embodiment of the present invention.
[0126] The electronic device 1 may include a processor 10, a memory 11 and a bus, and may further include a computer program stored in the memory 11 and executable on the processor 10, such as the unstructured data lake-in program.
[0127] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. The memory 11 may be an internal storage unit of the electronic device 1 in some embodiments, such as the mobile hard disk of the electronic device 1. The memory 11 may also be an external storage device of the electronic device 1 in other embodiments, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 1. Further, the memory 11 may include both the internal storage unit and the external storage device of the electronic device 1. The memory 11 can be used not only to store application software installed in the electronic device 1 and various types of data, such as the code of the unstructured data lake-in program, but also to temporarily store data that has been output or will be output.
[0128] In some embodiments, the processor 10 may be composed of an integrated circuit. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips, etc. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and circuits, and executing various functions of the electronic device 1 and processing data by running or executing programs or modules (such as unstructured data lake-in programs, etc.) stored in the memory 11, and calling the data stored in the memory 11.
[0129] The bus may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is configured to enable connection communication between the memory 11 and at least one processor 10, etc.
[0130] Figure 5 Only the electronic device with components is shown. Those skilled in the art can understand that Figure 5 the shown structure does not constitute a limitation on the electronic device 1, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0131] For example, although not shown, the electronic device 1 may further include a power source (such as a battery) for powering each component. Preferably, the power source may be logically connected to the at least one processor 10 through a power management device, so as to implement functions such as charge management, discharge management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, a power status indicator, etc. The electronic device 1 may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0132] Furthermore, the electronic device 1 may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device 1 and other electronic devices.
[0133] Optionally, the electronic device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device 1 and to display a visual user interface.
[0134] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.
[0135] The unstructured data lake-in program stored in the memory 11 of the electronic device 1 is a combination of multiple instructions, which can be implemented when running in the processor 10:
[0136] Obtain a data processing request, and parse the data processing request to obtain the data to be processed, the metadata to be processed, and the processing type;
[0137] Generate extended metadata for the data to be processed, where the extended metadata includes a data key value, a data version, and data validity;
[0138] When the processing type is write, use the data key value and the data version to perform a data duplicate check on the existing data in the preset storage system. When the data duplicate check passes, store the data to be processed in the data storage space of the preset storage system, store the corresponding metadata to be processed and the extended metadata in the metadata storage space of the preset storage system, and set the data validity of the data to be processed to valid;
[0139] When the processing type is change or delete, use the data key value and the data version to locate the existing target object in the preset storage system, set the data validity of the target object to invalid, store the data to be processed with the processing type of change in the data storage space of the preset storage system and store the corresponding metadata to be processed and the extended metadata in the metadata storage space of the preset storage system, set the data validity of the data to be processed to valid, and delete the target object with the processing type of delete.
[0140] Furthermore, if the modules / units integrated in the electronic device 1 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory).
[0141] The present invention also provides a computer-readable storage medium. The readable storage medium stores a computer program, and when the computer program is executed by a processor of an electronic device, it can implement:
[0142] Obtain a data processing request, and parse the data processing request to obtain the data to be processed, the metadata to be processed, and the processing type;
[0143] Generate extended metadata for the data to be processed, where the extended metadata includes a data key value, a data version, and data validity;
[0144] When the processing type is write, use the data key value and the data version to perform data duplicate check on the existing data in the preset storage system. When the data duplicate check passes, store the data to be processed in the data storage space of the preset storage system, store the corresponding metadata to be processed and the extended metadata in the metadata storage space of the preset storage system, and set the data validity of the data to be processed to valid;
[0145] When the processing type is change or delete, use the data key value and the data version to locate the existing target object in the preset storage system, set the data validity of the target object to invalid, store the data to be processed with the processing type of change in the data storage space of the preset storage system and store the corresponding metadata to be processed and the extended metadata in the metadata storage space of the preset storage system, set the data validity of the data to be processed to valid, and delete the target object with the processing type of delete.
[0146] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of hardware plus software functional modules.
[0147] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.
[0148] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
[0149] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer, etc.
[0150] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0151] In addition, it is obvious that the word "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices stated in the system claims can also be implemented by one unit or device through software or hardware. Words such as "second" are used to denote names and do not denote any particular order.
[0152] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for unstructured data to enter the lake, characterized in that, The method includes: Obtain a data processing request, and parse the data processing request to obtain the data to be processed, the metadata to be processed, and the processing type; Generate extended metadata for the data to be processed, where the extended metadata includes a data key value, a data version, and data validity; When the processing type is write, use the data key value and the data version to perform data duplicate check on the existing data in the preset storage system. When the data duplicate check passes, store the data to be processed in the data storage space of the preset storage system, store the corresponding metadata to be processed and the extended metadata in the metadata storage space of the preset storage system, and set the data validity of the data to be processed to valid; When the processing type is change or delete, use the data key value and the data version to locate the existing target object in the preset storage system, set the data validity corresponding to the target object to invalid, store the data to be processed with the processing type of change in the data storage space of the preset storage system and store the corresponding metadata to be processed and the extended metadata in the metadata storage space of the preset storage system, set the data validity of the data to be processed to valid, and delete the target object with the processing type of delete.
2. The method for unstructured data to enter the lake according to claim 1, wherein, The generating the extended metadata for the data to be processed, where the extended metadata includes a data key value, a data version, and data validity, includes: Calculate the hash value of the data to be processed, and use the hash value as the data key value of the data to be processed; Determine whether the data key value is unique in the preset storage system; When the data key value is unique, set the data version corresponding to the data to be processed to the first version; When the data key value is not unique, obtain the highest data version corresponding to the data key value, and obtain the data version corresponding to the data to be processed by adding 1 to the highest data version; Initialize the data validity corresponding to the data to be processed to invalid; Collect the data key value, the data version, and the data validity to obtain the extended metadata corresponding to the data to be processed.
3. The method for unstructured data to enter the lake according to claim 1, characterized in that, The using the data key value and the data version to perform data duplicate check on the existing data in the preset storage system includes: Query in the preset storage system for a data key value that matches the data key value or a data version that matches the data version or a data version that matches each other; In the preset storage system, when no data key value that matches the data key value is found, return a message indicating that the data duplicate check passes; In the preset storage system, when a data key value that matches the data key value is found but no data version that matches the data version is found, return a message indicating that the data duplicate check passes; In the preset storage system, when a data key value that matches the data key value is found and a data version that matches the data version is found, return a message indicating that the data duplicate check fails.
4. The unstructured data lake-inlet method according to claim 1, wherein Locating an existing target object in the preset storage system by using the data key value and the data version includes: Using the data key value and the data version corresponding to the data to be processed as an index; In the preset storage system, querying for data key values and data versions that match the index; Taking the data, metadata, and extended metadata corresponding to the matched data key values and data versions as the target object.
5. The method for unstructured data to enter the lake according to claim 1, characterized in that, After storing the data to be processed in the data storage space of the preset storage system and storing the corresponding metadata to be processed and the extended metadata in the metadata storage space of the preset storage system, the method further includes: Obtaining the data storage address of the data to be processed in the data storage space and the metadata storage addresses of the corresponding metadata to be processed and the extended metadata in the metadata storage space respectively; Creating a mapping relationship between the data storage address and the metadata storage address according to the correspondence between the data to be processed and the metadata to be processed and the extended metadata.
6. The method for unstructured data to enter the lake as claimed in claim 3, wherein After, in the preset storage system, when a data key value that matches the data key value is found and a data version that matches the data version is found, returning a message indicating that the data duplicate check fails, the method includes: Rejecting the data processing request.
7. An unstructured data lake-inlet device, characterized in that, The apparatus includes: A data processing request parsing module, configured to obtain a data processing request and parse the data processing request to obtain data to be processed, metadata to be processed, and a processing type; An extended metadata generation module, configured to generate extended metadata for the data to be processed, where the extended metadata includes a data key value, a data version, and data validity; A data writing module, configured to, when the processing type is writing, perform a data duplicate check on existing data in a preset storage system by using the data key value and the data version, and when the data duplicate check passes, store the data to be processed in the data storage space of the preset storage system, store the corresponding metadata to be processed and the extended metadata in the metadata storage space of the preset storage system, and set the data validity of the data to be processed to valid; A data update and deletion module, configured to, when the processing type is change or deletion, locate an existing target object in the preset storage system by using the data key value and the data version, set the data validity of the target object to invalid, store the data to be processed with the processing type of change in the data storage space of the preset storage system and store the corresponding metadata to be processed and the extended metadata in the metadata storage space of the preset storage system, set the data validity of the data to be processed to valid, and delete the target object with the processing type of deletion.
8. The unstructured data lake-in device according to claim 7, characterized in that, The data update and deletion module locates the target object through the following operations: Locating an existing target object in the preset storage system by using the data key value and the data version includes: Use the data key value and the data version corresponding to the data to be processed as an index; In the preset storage system, query the data key value and data version that match the index; Use the data, metadata, and extended metadata corresponding to the matched data key value and data version as the target object.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the unstructured data lake-in method according to any one of claims 1 to 6.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the unstructured data lake-in method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Data storage method, device and equipment and storage medium
CN109710190A
Data processing method and device, storage medium and electronic device
CN111782635A