A data processing method and apparatus
By introducing a data lineage database into the data warehouse, the data in the distributed file system can be directly manipulated and lineage records can be recorded, which solves the problems of complex data warehouse structure and high cost, and realizes simplified data storage and flexible data processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
- Filing Date
- 2022-07-04
- Publication Date
- 2026-04-14
AI Technical Summary
Existing data warehouse technology architectures are complex, making it difficult to trace data lineage, maintain transaction consistency, and are costly and inflexible.
By using a data lineage database, data in the distributed file system can be manipulated directly, and operation records are saved in the data lineage database, simplifying data storage and calculation processes. Data can be traced through the lineage database, simplifying data table reconstruction and transaction consistency management.
It simplifies the data warehouse technical architecture, reduces construction costs, improves flexibility, and ensures transactional consistency in data processing and clear traceability of data sources.
Smart Images

Figure CN115168507B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a data processing method and apparatus. Background Technology
[0002] Existing data warehouse architectures process data in layers before providing it to users. Whether offline or real-time, data storage and computation involve aggregating and summarizing data from multiple sources before inputting it into the data warehouse. The data warehouse then uses data computation and storage tools to perform multi-layered processing, including null checks, anomaly detection, data splitting, and data extraction, before providing the data to users for offline analysis, data profiling, business queries, real-time analysis, real-time recommendations, real-time queries, and real-time risk control.
[0003] In the process of realizing this invention, the inventors discovered at least the following problems in the prior art:
[0004] Because existing data warehouse technology architectures are complex and require layered data processing, it is difficult to trace data lineage, maintain transaction consistency, and the construction cost of data warehouses is high and their flexibility is poor. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a data processing method and apparatus that can directly operate on data in a distributed file system and save operation records to a data lineage database. This allows data to meet all data storage and computation needs solely through the distributed file system and the data lineage database. The data warehouse architecture is simple, requiring no layered data processing, making data lineage tracing easy and maintaining operational consistency. Simultaneously, it significantly reduces the construction cost of the data warehouse system and offers high flexibility. Using the lineage database of the present invention for data tracing, and obtaining data lineage based on operation records, data tables can be quickly reconstructed, ensuring transactional consistency in data processing. It also supports rapid querying of historical and intermediate data during data reading, ensuring that the source of all data is clearly traceable.
[0006] To achieve the above objectives, according to one aspect of the present invention, a data processing method is provided, comprising:
[0007] In response to user operations on data, the operation statement of the current operation is parsed to obtain the parsing result, which includes the operation object;
[0008] Get the filename of the operation record of the previous operation of the operation object;
[0009] The parsing result and the filename of the operation record of the previous operation of the operation object are combined to generate the operation record of the current operation;
[0010] The operation record of the current operation is saved to the blood relationship database.
[0011] Optionally, the operation records are stored in the form of a structure table.
[0012] Optionally, the operation record is named with a timestamp.
[0013] Optionally, it also includes: merging operation records stored in the blood relationship database according to a set number of merge files, and naming the merged operation records according to the minimum and maximum timestamps in the file names of the operation records before merging.
[0014] Optionally, it also includes: compressing and storing the operation records stored in the blood relation database.
[0015] Optionally, it further includes: responding to a data source lookup request, searching for the operation record of the most recent operation on the data based on the data identifier; parsing the operation record, and determining the original data corresponding to the data based on the parsing result.
[0016] Optionally, the parsing result further includes an operation operator and an operation range; determining the original data corresponding to the data based on the parsing result includes: if the parsing result does not include the operation record file name of the previous operation of the operation object, then determining the original data corresponding to the data from the operation range based on the operation object and the operation operator; if the parsing result includes the operation record file name of the previous operation of the operation object, then parsing the operation record of the previous operation of the operation object, and repeating the above steps to determine the original data corresponding to the data based on the parsing result.
[0017] According to another aspect of the present invention, a data processing apparatus is provided, comprising:
[0018] The data operation parsing module is used to respond to user operations on data, parse the operation statement to obtain the parsing result, and the parsing result includes the operation object;
[0019] The operation record acquisition module is used to acquire the file name of the operation record of the previous operation of the operation object;
[0020] The operation record generation module is used to assemble the parsing result and the operation record file name of the previous operation of the operation object to generate the operation record of the current operation;
[0021] The operation record saving module is used to save the operation record of the current operation to the blood relationship database.
[0022] According to another aspect of the present invention, a data processing electronic device is provided.
[0023] A data processing electronic device includes: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the data processing method provided in the embodiments of the present invention.
[0024] According to another aspect of the present invention, a computer-readable medium is provided.
[0025] A computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the data processing method provided in the embodiments of the present invention.
[0026] One embodiment of the above invention has the following advantages or beneficial effects: By responding to user operations on data, the operation statement of the current operation is parsed to obtain the parsing result, which includes the operation object; the filename of the operation record of the previous operation of the operation object is obtained; the parsing result and the filename of the operation record of the previous operation of the operation object are combined to generate the operation record of the current operation; the technical solution of saving the operation record of the current operation to a lineage database allows direct operation on data in the distributed file system and saving the operation record to the data lineage database. This enables data to meet all data storage and computation needs solely through the distributed file system and the data lineage database. The data warehouse architecture is simple, requiring no layered data processing, and data lineage tracing is simple while maintaining operational transaction consistency. Simultaneously, it significantly reduces the construction cost of the data warehouse system and offers high flexibility. Using the lineage database of the present invention for data tracing, and obtaining data lineage based on operation records, data tables can be quickly reconstructed, ensuring transaction consistency in data processing. It supports rapid querying of historical data and intermediate data during data reading, ensuring that the source of all data is clearly traceable.
[0027] The further effects of the aforementioned unconventional alternative methods will be explained below in conjunction with specific implementation methods. Attached Figure Description
[0028] The accompanying drawings are provided to better understand the invention and are not intended to unduly limit the scope of the invention. Wherein:
[0029] Figure 1 This is a schematic diagram of the main steps of the data processing method according to an embodiment of the present invention;
[0030] Figure 2 This is a schematic diagram of the main modules of a data processing apparatus according to an embodiment of the present invention;
[0031] Figure 3 This is an exemplary system architecture diagram in which embodiments of the present invention can be applied;
[0032] Figure 4 This is a schematic diagram of the structure of a computer system suitable for implementing terminal devices or servers of the present invention. Detailed Implementation
[0033] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of the present invention, including various details to aid understanding. These details should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0034] With the rise of the Internet of Things (IoT) and smart devices, more and more data will be integrated into the system, potentially reaching hundreds of billions or even trillions of records. The collection and processing of IoT data presents significant challenges: the IoT industry generally generates large volumes of data, making data storage a huge cost burden for enterprises; data warehouses rely on batch processing, which cannot meet the demands of real-time data analysis; rapid business development necessitates more flexible and agile data acquisition methods; and the algorithmic models in the IoT industry require flexible data pipelines. However, existing data warehouse systems are complex in structure and require layered processing, making it difficult to trace data lineage, maintain transaction consistency, and resulting in high construction costs and poor flexibility.
[0035] To address the aforementioned technical problems in the existing technology, this invention provides a data processing method and apparatus. Based on the traditional data architecture, an adjustment is made by establishing a data lineage for the original dataset. Each data computation is considered a "commit" of an operation, and each commit records the data source and computation logic (as well as the computation status, including whether the computation was successful). This information is stored in a structured table (e.g., XML, JSON). If a user needs to extract data from a specific stage or point in time, they can first read the structured table of the operation records to obtain the rollback logic, and then perform reverse computation to obtain the result.
[0036] In existing technologies, building a data warehouse in a project faces dual pressures of cost and data traceability. The cost pressure stems from the need for multi-layered processing of data from various data sources during data warehouse construction. This primarily includes: a data aggregation layer that aggregates and merges data from data sources (e.g., stored in MySQL, HBase) and saves it to the distributed file system HDFS, where HDFS periodically cleans the data to remove empty or noisy data; subsequently, the data in HDFS undergoes data splitting and extraction before being stored in a data mart; finally, the data in the data mart is processed by the distributed computing framework SPARK and then saved again to different storage spaces within the data mart. This process involves multiple different data warehouses or databases for data storage, as well as multiple data processing or computing frameworks for data processing, thus leading to significant cost pressures in data warehouse construction. In particular, data marts often exhibit data stratification, with each layer storing the same data in multiple locations, and data points being persisted to multiple disks. The pressure of data traceability arises because distributed file systems require cleaning abnormal data to ensure data quality, and data marts require data aggregation across multiple dimensions to ensure data closely aligns with business needs. As a result, the most original data is discarded, making traceability difficult.
[0037] To address these two pain points, this invention incorporates a new data lineage database. After aggregating the source data, it stores it in a distributed file system. Then, it directly operates on the data in the distributed file system and saves the operation records to the data lineage database. This allows the data to meet all data storage and computation needs using only the distributed file system and the data lineage database.
[0038] Figure 1 This is a schematic diagram illustrating the main steps of a data processing method according to an embodiment of the present invention. Figure 1 As shown, the data processing method of this embodiment mainly includes the following steps S101 to S104.
[0039] Step S101: In response to user operations on data, the operation statement of the current operation is parsed to obtain the parsing result, which includes the operation object. Whenever a user submits an operation task (job) on data in the distributed file system, the data lineage database parses the operation statement (SQL statement) of the operation into three elements: operation operator, operation scope (time range and data domain), and operation object.
[0040] For example, the user's statement for manipulating data is as follows:
[0041] SELECT prod_id,quantity,item_price,quantity*item_price AS exp_priceFROM order items WHERE order_time=1654486405
[0042] After parsing the operation statement, we can extract the following: "Operator: *; Operation objects: quantity, item_price, prod_id, order items; Operation range: order_time = 1654486405".
[0043] Step S102: Obtain the filename of the operation record of the previous operation for the operation object. After each analysis of the user's operation statement, the analyzed operation record is saved, and the operation record has a filename. Simultaneously, a data table can be created to store the mapping relationship between the execution result of the operation statement (i.e., the operation result) and the operation record filename. In one embodiment of the present invention, the operation record is named using a timestamp. When obtaining the filename of the operation record of the previous operation for the operation object, the operation object can be used as the operation result of the previous operation, and the corresponding operation record filename can be found based on the mapping relationship between the operation result and the operation record filename.
[0044] Step S103: Combine the parsing result with the filename of the operation record of the previous operation for the object to generate the operation record for the current operation. The data lineage database combines the three elements included in the parsing result with the filename of the operation record of the previous operation for the object to obtain the operation record for the current operation.
[0045] Step S104: Save the operation record of the current operation to the lineage relationship database. After obtaining the operation record of the current operation, save the operation record as a file, and record it in the lineage relationship database with the timestamp (accurate to milliseconds) as the filename. In an embodiment of the present invention, the operation record is stored in the form of a structure table. This structure table is, for example, in XML format or JSON format, etc.
[0046] According to embodiments of the present invention, operation records stored in the kinship database are merged according to a set number of files to be merged, and the merged operation records are named based on the minimum and maximum timestamps in the filenames of the operation records before merging. For example, 10 operation record files can be merged into one file, that is, 10 XML files can be merged into one XML file, and the filename is named by concatenating the maximum and minimum timestamps of these 10 filenames. By merging the operation record files, the number of files can be reduced, and disk read efficiency can be improved. In addition, in another embodiment, the operation records stored in the kinship database can also be compressed for storage to further reduce data storage space and improve data query speed and efficiency.
[0047] According to another embodiment of the present invention, based on the aforementioned kinship database, data source retrieval can also be performed. In response to a data source retrieval request, the operation record of the most recent operation on the data is retrieved according to the data identifier; the operation record is parsed, and the original data corresponding to the data is determined based on the parsing result.
[0048] According to another embodiment of the present invention, the parsing result further includes an operation operator and an operation range; determining the original data corresponding to the data based on the parsing result includes: if the parsing result does not include the operation record file name of the previous operation of the operation object, then the original data corresponding to the data is determined from the operation range based on the operation object and the operation operator; if the parsing result includes the operation record file name of the previous operation of the operation object, then the operation record of the previous operation of the operation object is parsed, and the above steps are repeated to determine the original data corresponding to the data based on the parsing result. Specifically, if the parsing result includes the operation object, operation operator, and operation range, but does not include the operation record file name of the previous operation of the operation object, it indicates that the operation record is an operation on the most original data in the distributed data system, so the original data corresponding to the data is directly determined from the operation range based on the operation object and the operation operator. If the parsing result includes the operation record file name of the previous operation of the operation object in addition to the operation object, operation operator, and operation range, it indicates that secondary tracing or more levels of tracing are required. In this case, the operation record of the previous operation of the operation object needs to be parsed, and the previous steps are repeated to determine the original data corresponding to the data based on the parsing result.
[0049] When a user needs to trace the source of certain data, they only need to find the last operation record (XML file) of that data. By parsing the operation object, operation operator, and operation range within the XML file, the original data can be found. If second- or third-level tracing is required, the user can trace back to the previous level of XML through the last operation record in the XML file, thereby finding the original data. Using the lineage database of this invention for data tracing, the data lineage can be obtained from the XML file, allowing for rapid reconstruction of data tables, ensuring transactional consistency in data processing, supporting rapid querying of historical and intermediate data during data reading, and ensuring that the source of all data is clearly traceable.
[0050] Figure 2 This is a schematic diagram of the main modules of a data processing apparatus according to an embodiment of the present invention. Figure 2 As shown, the data processing device 200 of this embodiment mainly includes a data operation parsing module 201, an operation record acquisition module 202, an operation record generation module 203, and an operation record storage module 204.
[0051] The data operation parsing module 201 is used to respond to user operations on data by parsing the operation statement of the current operation to obtain a parsing result, the parsing result including the operation object;
[0052] The operation record acquisition module 202 is used to acquire the file name of the operation record of the previous operation of the operation object;
[0053] The operation record generation module 203 is used to assemble the parsing result and the operation record file name of the previous operation of the operation object to generate the operation record of the current operation;
[0054] The operation record saving module 204 is used to save the operation record of the current operation to the blood relationship database.
[0055] According to one embodiment of the present invention, the operation record is stored in the form of a structure table.
[0056] According to another embodiment of the invention, the operation record is named with a timestamp.
[0057] According to another embodiment of the present invention, the data processing device 200 further includes an operation record reprocessing module (not shown in the figure), which is used to: merge the operation records stored in the blood relationship database according to a set number of merge files, and name the merged operation records according to the minimum and maximum timestamps in the file names of the operation records before merging.
[0058] According to another embodiment of the present invention, the operation record reprocessing module (not shown in the figure) can also be used to compress and store the operation records stored in the blood relationship database.
[0059] According to another embodiment of the present invention, the data processing device 200 further includes a data source lookup module (not shown in the figure), configured to: in response to a data source lookup request, look up the operation record of the most recent operation of the data according to the data identifier; parse the operation record, and determine the original data corresponding to the data based on the parsing result.
[0060] According to another embodiment of the present invention, the parsing result further includes an operation operator and an operation range; the data source lookup module (not shown in the figure) can also be used to: if the parsing result does not include the operation record file name of the previous operation of the operation object, then determine the original data corresponding to the data from the operation range according to the operation object and the operation operator; if the parsing result includes the operation record file name of the previous operation of the operation object, then parse the operation record of the previous operation of the operation object, and repeat the above steps to determine the original data corresponding to the data according to the parsing result.
[0061] According to the technical solution of this invention, in response to user operations on data, the operation statement of the current operation is parsed to obtain the parsing result, which includes the operation object; the operation record file name of the previous operation of the operation object is obtained; the parsing result and the operation record file name of the previous operation of the operation object are assembled to generate the operation record of the current operation; and the operation record of the current operation is saved to the lineage database. This technical solution allows direct operation on data in the distributed file system and saves the operation record to the data lineage database. Data can be stored and computed using only the distributed file system and the data lineage database, resulting in a simple data warehouse architecture that eliminates the need for layered data processing. Data lineage tracing is simple and maintains operational transaction consistency. It also significantly reduces the construction cost of the data warehouse system and offers high flexibility. Using the lineage database of this invention for data tracing and obtaining data lineage based on operation records allows for rapid reconstruction of data tables, ensuring transaction consistency in data processing. It also supports rapid querying of historical and intermediate data during data reading, ensuring that the source of all data is clearly traceable.
[0062] Figure 3 An exemplary system architecture 300 is shown that can be applied to the data processing method or data processing apparatus of the present invention.
[0063] like Figure 3As shown, system architecture 300 may include terminal devices 301, 302, and 303, a network 304, and a server 305. Network 304 serves as the medium for providing communication links between terminal devices 301, 302, and 303 and server 305. Network 304 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0064] Users can use terminal devices 301, 302, and 303 to interact with server 305 via network 304 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 301, 302, and 303, such as database applications, data analysis tool applications, data processing applications, data warehouses, etc. (for example only).
[0065] Terminal devices 301, 302, and 303 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0066] Server 305 can be a server providing various services, such as a backend management server that supports data processing requests sent by users using terminal devices 301, 302, and 303 (this is just an example). The backend management server can respond to the received data processing requests and other data, allowing the user to perform operations on the data. It can parse the operation statements to obtain parsing results, including the operation object; obtain the filename of the operation record of the previous operation of the operation object; assemble the parsing results and the filename of the operation record of the previous operation of the operation object to generate the operation record of the current operation; save the operation record of the current operation to a lineage database, and then feed back the processing results (e.g., operation record, operation record saving result – this is just an example) to the terminal device.
[0067] It should be noted that the data processing method provided in the embodiments of the present invention is generally executed by server 305, and correspondingly, the data processing device is generally located in server 305.
[0068] It should be understood that Figure 3 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0069] The following is for reference. Figure 4 It shows a schematic diagram of the structure of a computer system 400 suitable for implementing terminal devices or servers of the present invention. Figure 4 The terminal device or server shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of the present invention.
[0070] like Figure 4 As shown, the computer system 400 includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 402 or programs loaded from storage section 408 into random access memory (RAM) 403. The RAM 403 also stores various programs and data required for the operation of the system 400. The CPU 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0071] The following components are connected to I / O interface 405: an input section 406 including a keyboard, mouse, etc.; an output section 407 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN card, modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to I / O interface 405 as needed. A removable medium 411, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 410 as needed so that computer programs read from it can be installed into storage section 408 as needed.
[0072] In particular, according to the embodiments disclosed in this invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 409, and / or installed from removable medium 411. When the computer program is executed by central processing unit (CPU) 401, it performs the functions defined above in the system of this invention.
[0073] It should be noted that the computer-readable medium shown in this invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0074] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0075] The units or modules described in the embodiments of the present invention can be implemented in software or hardware. The described units or modules can also be housed in a processor; for example, a processor can be described as including a data operation parsing module, an operation record acquisition module, an operation record generation module, and an operation record saving module. The names of these units or modules do not necessarily limit the specific unit or module itself; for example, the operation record saving module can also be described as "a module for saving the operation record of the current operation to a kinship database."
[0076] In another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiments; or it may exist independently and not assembled into the device. The computer-readable medium carries one or more programs, which, when executed by the device, cause the device to: parse an operation statement in response to a user's operation on data to obtain a parsing result, the parsing result including an operation object; obtain the filename of the operation record of the previous operation of the operation object; assemble the parsing result and the filename of the operation record of the previous operation of the operation object to generate the operation record of the current operation; and save the operation record of the current operation to a lineage database.
[0077] According to the technical solution of this invention, in response to user operations on data, the operation statement of the current operation is parsed to obtain the parsing result, which includes the operation object; the operation record file name of the previous operation of the operation object is obtained; the parsing result and the operation record file name of the previous operation of the operation object are assembled to generate the operation record of the current operation; and the operation record of the current operation is saved to the lineage database. This technical solution allows direct operation on data in the distributed file system and saves the operation record to the data lineage database. Data can be stored and computed using only the distributed file system and the data lineage database, resulting in a simple data warehouse architecture that eliminates the need for layered data processing. Data lineage tracing is simple and maintains operational transaction consistency. It also significantly reduces the construction cost of the data warehouse system and offers high flexibility. Using the lineage database of this invention for data tracing and obtaining data lineage based on operation records allows for rapid reconstruction of data tables, ensuring transaction consistency in data processing. It also supports rapid querying of historical and intermediate data during data reading, ensuring that the source of all data is clearly traceable.
[0078] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A data processing method, characterized in that, include: In response to user operations on data, the operation statement of the current operation is parsed to obtain the parsing result, which includes the operation object, operation operator, and operation range; The operation object is used as the operation result of the previous operation, and the operation record file name of the previous operation of the operation object is obtained according to the mapping relationship between the operation result and the operation record file name; The parsing result and the filename of the operation record of the previous operation of the operation object are combined to generate the operation record of the current operation; Save the operation record of the current operation to the blood relationship database; In response to a data source lookup request, the operation record of the most recent operation on the data is retrieved based on the data identifier; Parsing the operation record and determining the original data corresponding to the data based on the parsing result includes: if the parsing result does not include the operation record file name of the previous operation of the operation object, then determining the original data corresponding to the data from the operation range based on the operation object and the operation operator; if the parsing result includes the operation record file name of the previous operation of the operation object, then parsing the operation record of the previous operation of the operation object to obtain the parsing result, and repeating the step of determining the original data corresponding to the data based on the parsing result until the original data corresponding to the data is determined.
2. The method according to claim 1, characterized in that, The operation records are stored in the form of a structure table.
3. The method according to claim 1, characterized in that, The operation records are named using timestamps.
4. The method according to claim 3, characterized in that, Also includes: The operation records stored in the blood relationship database are merged according to the set number of merge files, and the merged operation records are named according to the minimum and maximum timestamps in the file names of the operation records before merging.
5. The method according to claim 1, characterized in that, Also includes: The operation records stored in the blood relation database are compressed and stored.
6. A data processing apparatus, characterized in that, include: The data operation parsing module is used to respond to user operations on data, parse the operation statement of the current operation to obtain the parsing result, which includes the operation object, operation operator and operation range; The operation record acquisition module is used to treat the operation object as the operation result of the previous operation and obtain the operation record file name of the previous operation of the operation object according to the mapping relationship between the operation result and the operation record file name; The operation record generation module is used to assemble the parsing result and the operation record file name of the previous operation of the operation object to generate the operation record of the current operation; An operation record saving module is used to save the operation record of the current operation to a blood relationship database; The data source lookup module is used to respond to a data source lookup request and look up the operation record of the most recent operation on the data based on the data identifier; Parsing the operation record and determining the original data corresponding to the data based on the parsing result includes: if the parsing result does not include the operation record file name of the previous operation of the operation object, then determining the original data corresponding to the data from the operation range based on the operation object and the operation operator; if the parsing result includes the operation record file name of the previous operation of the operation object, then parsing the operation record of the previous operation of the operation object to obtain the parsing result, and repeating the step of determining the original data corresponding to the data based on the parsing result until the original data corresponding to the data is determined.
7. An electronic device for data processing, characterized in that, include: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-5.
8. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Log processing method and device, electronic equipment and computer readable storage medium
CN113010480A
Big data storage and traceability system
CN114386098A