Data Processing Method, Apparatus, Electronic Device and Computer-Readable Storage Medium
By using query language in the Flink data processing engine to encrypt, edit and decrypt data sets, generate task processing sets and submit them to the data processing engine, the high entry threshold and insecurity of the Flink data processing engine is solved, and a lower threshold and higher security data processing process is achieved.
Patent Information
- Application Number
- CN202011645339.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-30
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-12-30
AI Technical Summary
The high entry barrier and unfriendly operation interface of the existing Flink data processing engine make it difficult for ordinary data processing personnel to operate and have poor security.
Through query language, the data set in the database in the interactive real-time computing platform is encrypted, edited and decrypted, and a task processing set is generated, and submitted to the data processing engine for parsing and execution, to generate a processing result set.
It lowers the threshold for data processing, improves the security of data processing, and improves the accuracy and interactivity rate of subsequent parallelized data processing.
Smart Images

Figure CN112632626B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, electronic device and computer-readable storage medium for data processing. Background Art
[0002] In the context of the booming development of big data, people have more and more occasions and opportunities to perform massive data processing. In the process of data processing, data processing engines are widely used, and the data processing engines can complete powerful real-time analysis. For example, the Flink data processing engine executes any stream data program in a data parallel and pipeline manner. The pipeline runtime system of the Flink data processing engine can execute batch processing and stream processing programs. In addition, the runtime of the Flink data processing engine itself also supports the execution of iterative algorithms and can perform multi-threaded reading and parallel execution of data.
[0003] Now, the Flink data processing engine is widely applied to stream data processing. However, for those who are not familiar with this product, there is a very high entry threshold. They need to understand the operating mechanism of the Flink data processing engine itself and also need to understand development code. For general data personnel, it is basically impossible to process stream data like processing HIVE data. Regarding the user interface, the Flink data processing engine product itself only provides an sql interactive interface based on linux shell. The single purpose leads to poor security, and at the same time, the operation is not friendly. It is necessary to log in to the server to operate, which is basically impossible for data developers without development skills to operate. Summary of the Invention
[0004] The present invention provides a data processing method, device and computer-readable storage medium, and its main purpose is to reduce the data processing threshold and improve the security of data processing.
[0005] To achieve the above object, a data processing method provided by the present invention includes:
[0006] Setting and encrypting and editing a data set in a database in an interactive real-time computing platform through a query language to obtain an encrypted data set;
[0007] Obtaining a task data set, decrypting the encrypted data set to obtain a decrypted data set, and adjusting parameters of the task data set and the decrypted data set to obtain a task processing set, where the task processing set includes an execution command and a task command;
[0008] Submitting the task processing set to a data processing engine in the interactive real-time computing platform, and parsing the execution command according to the task command to generate a parsed task set;
[0009] Submit the parsing task set to the data processing engine and run it to generate a processing result set.
[0010] Optionally, setting and encrypting and editing the data set in the database through a query language to obtain an encrypted data set, including:
[0011] Editing and setting the data set in the database through a query language. After successful editing and setting, create a source table, a dimension table, and a target table;
[0012] Performing encrypted editing on the source table to obtain an encrypted source table, and associating the encrypted source table with the dimension table and the target table to obtain the encrypted data set.
[0013] Optionally, performing encrypted editing on the source table to obtain an encrypted source table, and associating the encrypted source table with the dimension table and the target table to obtain the encrypted data set, including:
[0014] Obtain the ID of the source table, and create a first field according to the ID of the source table;
[0015] Use a pre-built encryption function to encrypt the ID of the source table to obtain a second field;
[0016] Use the first field and the second field to create an encrypted source table;
[0017] Associate the encrypted source table with the dimension table and the target table according to the ID of the source table to obtain the encrypted data set.
[0018] Optionally, decrypting the encrypted data set to obtain a decrypted data set, including:
[0019] Call the decryption function corresponding to the encryption function;
[0020] Decrypt the encrypted source table through the decryption function to obtain a decrypted source table and the ID of the decrypted source table;
[0021] Summarize the decrypted source table and the ID of the decrypted source table to obtain the decrypted data set.
[0022] Optionally, performing parameter adjustment on the task data set and the decrypted data set to obtain a task processing set, including:
[0023] Slice the task data set according to a pre-built data slicing function to obtain a slice set;
[0024] Traverse the slice set, compare it with the data in the decryption dataset, associate the data of the same type in the slice set with the data in the decryption dataset, and add command statements and command fields. Add task fields and task command statements to the data of different data types in the slice set and the decryption dataset;
[0025] Combine the command statements, command fields, task fields, and task command statements to obtain a task processing set.
[0026] Optionally, submitting the task processing set to the data processing engine, parsing the execution command according to the task command, and generating a parsed task set, including:
[0027] Obtain the task processing set, and parse the execution command through a pre-built parsing function to obtain a syntax tree;
[0028] Search and iterate through the syntax tree to obtain a set of dimension tables corresponding to the syntax tree;
[0029] Classify the set of dimension tables into a first set of dimension tables and a second set of dimension tables according to the dimension table structure and non-dimension table structure;
[0030] Select dimension tables from the first set of dimension tables and the second set of dimension tables through preset filtering conditions to obtain a parsed task set.
[0031] Optionally, submitting the parsed task set to the data processing engine and running it to generate a processing result set, including:
[0032] Obtain the parsed task set, and start the work management function and task management function in the data processing engine;
[0033] Use the data processing engine to transform the parsed task set into the form of a parallel data stream;
[0034] Run the work management function and the task management function to perform parallel processing on the parsed task set of the parallel data stream to generate a processing result set.
[0035] To solve the above problems, the present invention also provides a data processing device, and the device includes:
[0036] A data encryption module, configured to perform setting and encryption editing on a dataset in a database in an interactive real-time computing platform through a query language to obtain an encrypted dataset;
[0037] A data decryption module, configured to obtain a task dataset, decrypt the encrypted dataset to obtain a decrypted dataset, adjust parameters of the task dataset and the decrypted dataset to obtain a task processing set, and the task processing set includes an execution command and a task command;
[0038] A data parsing module, configured to submit the task processing set to a data processing engine in the interactive real-time computing platform, parse the execution command according to the task command, and generate a parsed task set;
[0039] A data processing module, configured to submit the parsed task set to the data processing engine and run to generate a processing result set.
[0040] Optionally, the data encryption module obtains an encrypted data set through the following operations:
[0041] Edit and set the data set in the database through a query language. After the edit and setting is successful, create a source table, a dimension table, and a target table;
[0042] Perform encrypted editing on the source table to obtain an encrypted source table, and associate the encrypted source table with the dimension table and the target table to obtain the encrypted data set.
[0043] Preferably, when generating the encrypted data set, the data encryption module is further configured to:
[0044] Obtain the ID of the source table, and create a first field according to the ID of the source table;
[0045] Use a pre-built encryption function to encrypt the ID of the source table to obtain a second field;
[0046] Create an encrypted source table by using the first field and the second field;
[0047] Associate the encrypted source table with the dimension table and the target table according to the ID of the source table to obtain the encrypted data set.
[0048] Preferably, the data decryption module obtains the decrypted data set by performing the following operations:
[0049] Call a decryption function corresponding to the encryption function;
[0050] Decrypt the encrypted source table through the decryption function to obtain a decrypted source table and the ID of the decrypted source table;
[0051] Summarize the decrypted source table and the ID of the decrypted source table to obtain the decrypted data set.
[0052] Preferably, the data encryption module obtains the task processing set through the following method:
[0053] Slice the task data set according to a pre-built data slicing function to obtain a slice set;
[0054] Traverse the slice set, compare it with the data in the decryption dataset, associate the data of the same type in the slice set with the data in the decryption dataset, and add command statements and command fields. Add task fields and task command statements to the data of different data types in the slice set and the decryption dataset;
[0055] Combine the command statements, command fields, task fields, and task command statements to obtain a task processing set.
[0056] Preferably, the data parsing module obtains the parsing task set through the following method:
[0057] Obtain the task processing set, and parse the execution command through a pre-built parsing function to obtain a syntax tree;
[0058] Search and iterate through the syntax tree to obtain a set of dimension tables corresponding to the syntax tree;
[0059] Classify the set of dimension tables into a first set of dimension tables and a second set of dimension tables according to the dimension table structure and non-dimension table structure;
[0060] Select dimension tables from the first set of dimension tables and the second set of dimension tables through preset filtering conditions to obtain a parsing task set.
[0061] Preferably, the data processing module obtains the processing result set through the following method:
[0062] Obtain the parsing task set, and start the work management function and task management function in the data processing engine;
[0063] Use the data processing engine to transform the parsing task set into the form of a parallel data stream;
[0064] Run the work management function and the task management function to perform parallel processing on the parsing task set of the parallel data stream, and generate a processing result set.
[0065] To solve the above problems, the present invention also provides an electronic device, which includes:
[0066] A memory that stores at least one instruction; and
[0067] A processor that executes the instructions stored in the memory to implement the data processing method described above.
[0068] To solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is executed by a processor in an electronic device to implement the data processing method described above.
[0069] In an embodiment of the present invention, a dataset in a database in an interactive real-time computing platform is set and encrypted and edited through a query language to obtain an encrypted dataset, ensuring the security of data under multi-threaded data processing. Further, a task dataset is obtained, the encrypted dataset is decrypted to obtain a decrypted dataset, and at the same time, parameters of the decrypted dataset are adjusted to obtain a task processing set, which can improve the accuracy of subsequent parallel data processing. In addition, since the acquisition of the task dataset is relatively simple, the threshold for data processing developers is greatly reduced, and the interaction rate is increased. Therefore, the data processing method, device, and computer-readable storage medium provided by the present invention can reduce the data processing threshold and improve the security of data processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 is a schematic flowchart of a data processing method provided by an embodiment of the present invention;
[0071] Figure 2 is a functional module diagram of a data processing device provided by an embodiment of the present invention;
[0072] Figure 3 is a schematic structural diagram of an electronic device for implementing a data processing method provided by an embodiment of the present invention;
[0073] The implementation, functional features, and advantages of the objectives of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0074] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0075] The execution subject of the data processing method provided by an embodiment of the present application includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by an embodiment of the present application. In other words, the data processing method can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc.
[0076] The present invention provides a data processing method. Refer to Figure 1 As shown, it is a schematic flowchart of a data processing method provided by an embodiment of the present invention. In this embodiment, the data processing method includes:
[0077] S1. Set and encrypt and edit a dataset in a database in an interactive real-time computing platform through a query language to obtain an encrypted dataset.
[0078] In one embodiment, the interactive real-time computing platform is used to establish data communication between the database and the data processing engine. In the interactive real-time computing platform, the database is a set of data stored together in a certain way, shareable by multiple users, with as little redundancy as possible, and independent of application programs, which can be regarded as an electronic filing cabinet - a place to store electronic files. Users can perform operations such as adding, querying, updating, and deleting data in the files. The data processing engine can use the currently publicly available Flink data processing engine. Obtain the Flink compression package, decompress and install the compression package, and configure it to generate the Flink data processing engine. The Flink data processing engine executes any stream data program in a data parallel and pipeline manner. The pipeline runtime system of the Flink data processing engine can execute batch processing and stream processing programs. In addition, the runtime of the Flink data processing engine itself also supports the execution of iterative algorithms.
[0079] Specifically, setting and encrypting and editing the data set in the database through a query language to obtain an encrypted data set includes:
[0080] Editing and setting the data set in the database through a query language. After the editing and setting is successful, create a source table, a dimension table, and a target table;
[0081] Performing encrypted editing on the source table to obtain an encrypted source table, and associating the encrypted source table with the dimension table and the target table to obtain the encrypted data set.
[0082] Preferably, the query language can use the currently publicly available Structured Query Language (SQL). SQL is the most widely used language in data processing, allowing users to concisely declare the required business logic. SQL belongs to a declarative language, and only needs to express the requirements clearly without the need to understand the specific implementation; SQL can be optimized, with multiple built-in query optimizers, and multiple query optimizers can translate the best execution plan for SQL. Editing and setting the data set in the database through SQL statement code to obtain a source table, a dimension table, a target table, etc.
[0083] In detail, the database is a set of data stored together in a certain way, shareable by multiple users, with as little redundancy as possible, and independent of application programs, which can be regarded as an electronic filing cabinet - a place to store electronic files. Users can perform operations such as adding, querying, updating, and deleting data in the files.
[0084] Specifically, in the embodiment of the present invention, since the subsequent collection of the task data set can be multi-threaded, in order to ensure data security, the data set is encrypted and edited through the query statement to obtain the encrypted data set.
[0085] Further, encrypting and editing the source table to obtain an encrypted source table, and associating the encrypted source table with the dimension table and the target table to obtain the encrypted data set, includes:
[0086] Obtain the ID of the source table, and create a first field according to the ID of the source table;
[0087] Use a pre-built encryption function to encrypt the ID of the source table to obtain a second field;
[0088] Create an encrypted source table using the first field and the second field;
[0089] Associate the encrypted source table with the dimension table and the target table according to the ID of the source table to obtain the encrypted data set.
[0090] Specifically, when creating a table (all the tables mentioned above), set the table type to a dynamic table. The fields in the dynamic table are of variable length. The dynamic table refers to a table type in which data is dynamically inserted as time and data change. The fields in a traditional static table are of fixed length and cannot automatically add fields as time and data change. When the changing data exceeds the fixed length of the static table, data overflow will occur.
[0091] For example, in the above embodiment, obtain the ID of the source table and denote it as the first field, use a pre-built encryption function (such as the base64 encryption function) to encrypt the ID to obtain a second field, and save the first field and the second field into the source table, so as to obtain an encrypted source table according to the ID of the source table; then associate the encrypted source table with the dimension table and the target table through the ID of the source table, which can initialize the data connection and create an available database connection pool; in addition, obtaining the encrypted data set based on the ID of the source table can effectively divide the data set, so as to ensure that parallel execution paths in subsequent data processing read and write by executing their respective data subsets.
[0092] S2. Obtain a task data set, decrypt the encrypted data set to obtain a decrypted data set, and perform parameter adjustment on the task data set and the decrypted data set to obtain a task processing set, where the task processing set contains execution commands and task commands.
[0093] Specifically, the task dataset refers to the dataset to be processed set by the user. The user can directly obtain the task dataset by entering the query statement on the preset query interface, which greatly reduces the development cost and enables non-R & D personnel to perform operations simply. If other task datasets need to be obtained, only the query statement code needs to be modified. At the same time, multiple task datasets are obtained through multiple query statements, and the task datasets are read in multiple threads, greatly improving the interaction rate.
[0094] Preferably, decrypting the encrypted dataset to obtain a decrypted dataset includes:
[0095] Invoking a decryption function corresponding to the encryption function;
[0096] Decrypting the encrypted source table through the decryption function to obtain a decrypted source table and the ID of the decrypted source table;
[0097] Summarizing the decrypted source table and the ID of the decrypted source table to obtain the decrypted dataset.
[0098] Specifically, in the embodiment of the present invention, by invoking a decryption function corresponding to the encryption function (such as a base64 decryption function), the second field in the encrypted source table is decrypted to obtain the ID of the encrypted source table, and whether the ID of the encrypted source table is the same as the first field in the encrypted source table is compared. If they are different, it is re-obtained. If they are the same, the decrypted source table is obtained, and the first field is the ID of the decrypted source table.
[0099] Further, adjusting the parameters of the task dataset and the decrypted dataset to obtain a task processing set specifically includes:
[0100] Slicing the task dataset according to a pre-constructed data slicing function to obtain a slice set;
[0101] Traversing the slice set, comparing it with the data in the decrypted dataset, associating the data of the same type in the slice set with the data in the decrypted dataset, and adding command statements and command fields, and adding task fields and task command statements to the data of different data types in the slice set and the decrypted dataset;
[0102] Combining the command statements, command fields, task fields, and task command statements to obtain a task processing set.
[0103] Specifically, through the parameter adjustment, the task dataset is sliced to obtain a slice set, and the slice set is traversed and compared and associated with the data in the decrypted dataset to obtain a task processing set. Since the data in the task processing set is split into multiple data subsets due to data splitting, the accuracy of subsequent task command parsing can be improved.
[0104] For example: the data in the task dataset is logically sliced according to the average slicing method through a preset slicing function (createInputSplites function) to obtain a slice set. The slice set stores the data information contained in each slice, including set information, cursor position, etc. Further, traverse the slice set, compare it with the data in the decryption dataset, associate the data of the same type (such as tables, text, etc.) in the slice set with the data in the decryption dataset, and add a command field and a command statement. Add a task field and a task command statement to the data in the slice set that does not belong to the same type as the data in the decryption dataset. Combine the command field, command statement, task field, and task command statement to obtain a task processing set. The task processing set contains the content data of each slice.
[0105] S3. Submit the task processing set to the data processing engine in the interactive real-time computing platform, and parse the execution command according to the task command to generate a parsed task set.
[0106] Further, the S3 includes:
[0107] Obtain the task processing set, and parse the execution command through a pre-built parsing function to obtain a syntax tree;
[0108] Search and iterate the syntax tree to obtain a dimension table set corresponding to the syntax tree;
[0109] Classify the dimension table set according to the dimension table structure and non-dimension table structure to obtain a first dimension table set and a second dimension table set;
[0110] Select dimension tables from the first dimension table set and the second dimension table set through preset filtering conditions to obtain a parsed task set.
[0111] Preferably, the data processing engine can use the currently publicly available Flink data processing engine, obtain the Flink compression package, decompress and install the compression package, and configure it to generate a Flink data processing engine. The Flink data processing engine executes any stream data program in a data parallel and pipeline manner. The pipeline runtime system of the Flink data processing engine can execute batch processing and stream processing programs. In addition, the runtime of the Flink data processing engine itself also supports the execution of iterative algorithms.
[0112] Specifically, the task processing engine parses the task processing set, and further filters the data subset in the task processing set through the syntax tree and the preset filtering conditions, so as to ensure that the subsequent data processing engine can more quickly transform the parsed task set into the form of parallel data streams.
[0113] Specifically, the syntax tree is a graphical representation of the execution command, representing the derivation result of the execution command, which is conducive to understanding the hierarchical structure of the execution command syntax.
[0114] S4. Submit the parsing task set to the data processing engine and run it to generate a processing result set.
[0115] Specifically, S4 includes:
[0116] Obtain the parsing task set and start the job management function and task management function in the data processing engine;
[0117] Use the data processing engine to transform the parsing task set into the form of parallel data streams;
[0118] Run the job management function and the task management function to perform parallel processing on the parsing task set of the parallel data stream and generate a processing result set.
[0119] For example, when querying unstructured data such as a financial statement set, based on the Flink data processing engine and the SQL, the SQL query function similar to structured data can be realized. The SQL processes the financial statement set to obtain a task processing set. The Flink data processing engine parses the task processing set to generate a parsing task set. The Flink data processing engine preprocesses the parsing task set and transforms it into the form of parallel data streams, and uses the Flink data processing engine to perform query processing on the parallel data streams to generate a result set. Through the Flink data processing engine, the query of the financial statement set is a parallel execution process and a multi-threaded operation, which greatly improves the interaction rate.
[0120] The query of the financial statement set is a parallel execution process. At the same time, according to user operations, multiple financial statement sets can be read in multiple threads, which greatly improves the interaction rate.
[0121] In the embodiments of the present invention, a dataset in a database in an interactive real-time computing platform is set and encrypted and edited through a query language to obtain an encrypted dataset, which ensures the security of data under multi-threaded data processing. Further, a task dataset is obtained, the encrypted dataset is decrypted to obtain a decrypted dataset, and at the same time, the parameters of the decrypted dataset are adjusted to obtain a task processing set, which can improve the accuracy of subsequent parallel data processing. In addition, since the acquisition of the task dataset is relatively simple, the threshold for data processing developers is greatly reduced, and the interaction rate is increased. Therefore, the data processing method, device and computer-readable storage medium proposed by the present invention can reduce the data processing threshold and improve the security of data processing.
[0122] As Figure 2 shown, it is a functional module diagram of a data processing device provided by an embodiment of the present invention.
[0123] The data processing device 100 of the present invention can be installed in an electronic device. According to the functions achieved, the data processing device may include a data encryption module 101, a data decryption module 102, a data parsing module 103 and a data processing module 104. The modules of the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.
[0124] In this embodiment, the functions of each module / unit are as follows:
[0125] The data encryption module 101 is used to set and encrypt and edit a dataset in a database in an interactive real-time computing platform through a query language to obtain an encrypted dataset;
[0126] The data decryption module 102 is used to obtain a task dataset, decrypt the encrypted dataset to obtain a decrypted dataset, and adjust the parameters of the task dataset and the decrypted dataset to obtain a task processing set, and the task processing set contains execution commands and task commands;
[0127] The data parsing module 103 is used to submit the task processing set to a data processing engine in the interactive real-time computing platform, and parse the execution command according to the task command to generate a parsed task set;
[0128] The data processing module 104 is used to submit the parsed task set to the data processing engine and run to generate a processing result set.
[0129] In this embodiment, the functions of each module / unit are as follows:
[0130] The data encryption module 101 sets and encrypts and edits a data set in a database in the interactive real-time computing platform through a query language to obtain an encrypted data set.
[0131] In an embodiment, the interactive real-time computing platform is used to establish data communication between the database and the data processing engine. In the interactive real-time computing platform, the database is a data set stored together in a certain way, can be shared by multiple users, has the smallest possible redundancy, and is independent of application programs, and can be regarded as an electronic filing cabinet - a place for storing electronic files. Users can perform operations such as adding, querying, updating, and deleting data in the files. The data processing engine can use the currently publicly available Flink data processing engine, obtain the Flink compressed package, decompress and install the compressed package, and configure it to generate the Flink data processing engine. The Flink data processing engine executes any stream data program in a data parallel and pipeline manner. The pipeline runtime system of the Flink data processing engine can execute batch processing and stream processing programs. In addition, the runtime of the Flink data processing engine itself also supports the execution of iterative algorithms.
[0132] Specifically, the setting and encrypting and editing of the data set in the database through the query language to obtain the encrypted data set includes:
[0133] Edit and set the data set in the database through the query language. After the edit and setting is successful, create a source table, a dimension table, and a target table;
[0134] Perform encrypted editing on the source table to obtain an encrypted source table, and associate the encrypted source table with the dimension table and the target table to obtain the encrypted data set.
[0135] Preferably, the query language can use the currently publicly available Structured Query Language (SQL). SQL is the most widely used language in data processing, allowing users to concisely declare the required business logic. SQL belongs to a declarative language. Just express the requirements clearly without having to understand the specific implementation; SQL can be optimized, with multiple built-in query optimizers, and multiple query optimizers can translate the best execution plan for SQL. Edit and set the data set in the database through SQL statement code to obtain a source table, a dimension table, a target table, etc.
[0136] In detail, the database is a data set stored together in a certain way, can be shared by multiple users, has the smallest possible redundancy, and is independent of application programs, and can be regarded as an electronic filing cabinet - a place for storing electronic files. Users can perform operations such as adding, querying, updating, and deleting data in the files.
[0137] Specifically, in the embodiments of the present invention, since the subsequent collection of the task data set can be multi-threaded, in order to ensure data security, the data set is encrypted and edited through the query statement to obtain the encrypted data set.
[0138] Further, encrypting and editing the source table to obtain an encrypted source table, and associating the encrypted source table with the dimension table and the target table to obtain the encrypted data set, including:
[0139] Obtain the ID of the source table, and create a first field according to the ID of the source table;
[0140] Use a pre-built encryption function to encrypt the ID of the source table to obtain a second field;
[0141] Use the first field and the second field to create an encrypted source table;
[0142] Associate the encrypted source table with the dimension table and the target table according to the ID of the source table to obtain the encrypted data set.
[0143] Specifically, when creating a table (all the tables mentioned above), set the table type to a dynamic table. The fields in the dynamic table are of variable length. The dynamic table refers to a table type in which data is dynamically inserted over time and as the data changes. The fields in a traditional static table are of fixed length and cannot add fields automatically over time and as the data changes. When the changing data exceeds the fixed length of the static table, data overflow will occur.
[0144] For example, in the above embodiment, obtain the ID of the source table and denote it as the first field, use a pre-built encryption function (such as the base64 encryption function) to encrypt the ID to obtain a second field, and save the first field and the second field into the source table, so as to obtain an encrypted source table according to the ID of the source table; then associate the encrypted source table with the dimension table and the target table through the ID of the source table, which can initialize the data connection and create an available database connection pool; in addition, obtaining the encrypted data set based on the ID of the source table can effectively partition the data set, so as to ensure that in subsequent data processing, parallel execution paths read and write by executing their respective data subsets.
[0145] The data decryption module 102 obtains the task data set, decrypts the encrypted data set to obtain a decrypted data set, and adjusts the parameters of the task data set and the decrypted data set to obtain a task processing set, which contains execution commands and task commands.
[0146] Specifically, the task dataset refers to the dataset to be processed set by the user. The user can directly obtain the task dataset by entering the query statement on the preset query interface, which greatly reduces the development cost and enables non-R&D personnel to perform operations very simply. If other task datasets need to be obtained, only the query statement code needs to be modified. At the same time, multiple task datasets are obtained through multiple query statements, and the task datasets are read in multiple threads, greatly improving the interaction rate.
[0147] Preferably, decrypting the encrypted dataset to obtain a decrypted dataset includes:
[0148] Invoking a decryption function corresponding to the encryption function;
[0149] Decrypting the encrypted source table through the decryption function to obtain a decrypted source table and the ID of the decrypted source table;
[0150] Summarizing the decrypted source table and the ID of the decrypted source table to obtain the decrypted dataset.
[0151] Specifically, in the embodiment of the present invention, by invoking a decryption function corresponding to the encryption function (such as a base64 decryption function), the second field in the encrypted source table is decrypted to obtain the ID of the encrypted source table, and the ID of the encrypted source table is compared with the first field in the encrypted source table. If they are different, re-obtain; if they are the same, the decrypted source table is obtained, and the first field is the ID of the decrypted source table. Further, adjusting the parameters of the task dataset and the decrypted dataset to obtain a task processing set specifically includes:
[0152] Slicing the task dataset according to a pre-constructed data slicing function to obtain a slice set;
[0153] Traversing the slice set, comparing it with the data in the decrypted dataset, associating the data of the same type in the slice set with the data in the decrypted dataset, and adding command statements and command fields, and adding task fields and task command statements to the data of different data types in the slice set and the decrypted dataset;
[0154] Combining the command statements, command fields, task fields, and task command statements to obtain a task processing set.
[0155] Specifically, through the parameter adjustment, the task dataset is sliced to obtain a slice set, and the slice set is traversed and compared with the data in the decrypted dataset to obtain a task processing set. Since the data in the task processing set is split into multiple data subsets due to data splitting, the accuracy of subsequent task command parsing can be improved.
[0156] For example: The data in the task dataset is logically sliced according to the average slicing method through a preset slicing function (createInputSplites function) to obtain a slice set. The slice set stores the data information contained in each slice, including set information, cursor position, etc. Further, traverse the slice set, compare it with the data in the decryption dataset, associate the data of the same type (such as tables, text, etc.) in the slice set with the data in the decryption dataset, and add a command field and a command statement. Add a task field and a task command statement to the data in the slice set that does not belong to the same type as the data in the decryption dataset. Combine the command field, command statement, task field, and task command statement to obtain a task processing set, and the task processing set contains the content data of each slice.
[0157] The data parsing module 103 submits the task processing set to the data processing engine in the interactive real-time computing platform, and parses the execution command according to the task command to generate a parsed task set.
[0158] Specifically, the data parsing module 103 parses the execution command according to the task command to generate a parsed task set, which specifically includes:
[0159] Obtain the task processing set, and parse the execution command through a pre-built parsing function to obtain a syntax tree;
[0160] Search and iterate the syntax tree to obtain a set of dimension tables corresponding to the syntax tree;
[0161] Classify the set of dimension tables according to the dimension table structure and non-dimension table structure to obtain a first set of dimension tables and a second set of dimension tables;
[0162] Select dimension tables from the first set of dimension tables and the second set of dimension tables through a preset filtering condition to obtain a parsed task set.
[0163] Preferably, the data processing engine can use the currently publicly available Flink data processing engine, obtain the Flink compression package, decompress and install the compression package, and configure it to generate the Flink data processing engine. The Flink data processing engine executes any stream data program in a data parallel and pipeline manner. The pipeline runtime system of the Flink data processing engine can execute batch processing and stream processing programs. In addition, the runtime of the Flink data processing engine itself also supports the execution of iterative algorithms.
[0164] Specifically, the task processing engine parses the task processing set, and further filters the data subset in the task processing set through the syntax tree and the preset filtering conditions, so as to ensure that the subsequent data processing engine can more quickly transform the parsed task set into the form of a parallel data stream.
[0165] Specifically, the syntax tree is a graphical representation of the execution command, representing the derivation result of the execution command, which is conducive to understanding the hierarchical structure of the execution command syntax.
[0166] The data processing module 104 submits the parsed task set to the data processing engine and runs to generate a processing result set.
[0167] Specifically, the data processing module 104 submits the parsed task set to the data processing engine and runs to generate a processing result set, which specifically includes:
[0168] Obtain the parsed task set, and start the work management function and task management function in the data processing engine;
[0169] Use the data processing engine to transform the parsed task set into the form of a parallel data stream;
[0170] Run the work management function and the task management function to perform parallel processing on the parsed task set of the parallel data stream to generate a processing result set.
[0171] For example, for querying unstructured data such as a financial statement set, based on the Flink data processing engine and the SQL, the SQL query function similar to structured data can be implemented. The SQL processes the financial statement set to obtain a task processing set. The Flink data processing engine parses the task processing set to generate a parsed task set. The Flink data processing engine preprocesses the parsed task set and transforms it into the form of a parallel data stream. The Flink data processing engine performs query processing on the parallel data stream to generate a result set. Through the Flink data processing engine, the query of the financial statement set is a parallel execution process and a multi-threaded operation, which greatly improves the interaction rate.
[0172] The query of the financial statement set is a parallel execution process. At the same time, according to user operations, multiple financial statement sets can be read in multiple threads, which greatly improves the interaction rate.
[0173] Such as Figure 3 shown, is a schematic structural diagram of an electronic device for implementing the data processing method provided by an embodiment of the present invention.
[0174] The electronic device 1 may include a processor 10, a memory 11, and a bus. It may also include a computer program stored in the memory 11 and executable on the processor 10, such as a data processing program 12.
[0175] Among them, the memory 11 includes at least one type of readable storage medium, which includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device 1, such as the mobile hard disk of the electronic device 1. In some other embodiments, the memory 11 may also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 1. Further, the memory 11 may also include both an internal storage unit and an external storage device of the electronic device 1. The memory 11 can be used not only to store application software installed on the electronic device 1 and various types of data, such as the code of the data processing program 12, etc., but also to temporarily store data that has been output or will be output.
[0176] In some embodiments, the processor 10 may be composed of integrated circuits. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including a combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines, and by running or executing programs or modules stored in the memory 11 (such as executing the data processing program, etc.), and calling data stored in the memory 11, to perform various functions of the electronic device 1 and process data.
[0177] The bus may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is set to enable connection and communication between the memory 11 and at least one processor 10, etc.
[0178] Figure 3Only an electronic device with components is shown. Those skilled in the art can understand that Figure 3 the shown structure does not constitute a limitation on the electronic device 1, and it may include fewer or more components than shown, or combine certain components, or have a different component arrangement.
[0179] For example, although not shown, the electronic device 1 may further include a power source (such as a battery) for supplying power to each component. Preferably, the power source can be logically connected to the at least one processor 10 through a power management device, so as to implement functions such as charge management, discharge management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, a power status indicator, etc. The electronic device 1 may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0180] Furthermore, the electronic device 1 may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device 1 and other electronic devices.
[0181] Optionally, the electronic device 1 may further include a user interface. The user interface may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device 1 and to display a visual user interface.
[0182] It should be understood that the embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.
[0183] The data processing program 12 stored in the memory 11 in the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can implement:
[0184] Setting and encrypting and editing a data set in a database in an interactive real-time computing platform through a query language to obtain an encrypted data set;
[0185] Obtain a task dataset, decrypt the encrypted dataset to obtain a decrypted dataset, adjust parameters of the task dataset and the decrypted dataset to obtain a task processing set, where the task processing set contains execution commands and task commands;
[0186] Submit the task processing set to a data processing engine in the interactive real-time computing platform, and parse the execution command according to the task command to generate a parsed task set;
[0187] Submit the parsed task set to the data processing engine and run it to generate a processing result set.
[0188] Specifically, for the specific implementation method of the above instructions by the processor 10, reference can be made to Figure 1 the description of the relevant steps in the corresponding embodiments, which will not be elaborated here.
[0189] Furthermore, if the modules / units integrated in the electronic device 1 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory).
[0190] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0191] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0192] In addition, in each embodiment of the present invention, the various functional modules can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of hardware plus software functional modules.
[0193] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.
[0194] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Accordingly, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced by the present invention. Any reference signs in the claims should not be construed as limiting the claims concerned.
[0195] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.
[0196] In addition, it is obvious that the word "including" does not exclude other units or steps, and the singular does not exclude the plural. A plurality of units or devices stated in the system claims can also be implemented by one unit or device through software or hardware. Words such as "second" are used to denote names and do not denote any particular order.
[0197] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A data processing method, characterized in that, the method includes: After editing and setting a data set in the database of an interactive real-time computing platform through a query language, creating a source table, a dimension table, and a target table, and creating a first field according to the ID of the source table; Using a pre-built encryption function to encrypt the ID of the source table to obtain a second field; Using the first field and the second field to create an encrypted source table; Associating the encrypted source table with the dimension table and the target table according to the ID of the source table to obtain an encrypted data set; Obtaining a task data set, decrypting the encrypted data set to obtain a decrypted data set, and adjusting parameters of the task data set and the decrypted data set to obtain a task processing set, where the task processing set contains execution commands and task commands; Submitting the task processing set to a data processing engine in the interactive real-time computing platform, and parsing the execution command according to the task command to generate a parsed task set; Submitting the parsed task set to the data processing engine and running it to generate a processing result set; The adjusting parameters of the task data set and the decrypted data set to obtain a task processing set includes: Slicing the task data set according to a pre-built data slicing function to obtain a slice set; Traversing the slice set, comparing it with the data in the decrypted data set, associating the data of the same type in the slice set with the data in the decrypted data set, and adding command statements and command fields, and adding task fields and task command statements to the data of different data types in the slice set and the decrypted data set; Combining the command statements, command fields, task fields, and task command statements to obtain a task processing set.
2. The data processing method according to claim 1, characterized in that, the decrypting the encrypted data set to obtain a decrypted data set includes: Invoking a decryption function corresponding to the encryption function; Decrypting the encrypted source table through the decryption function to obtain a decrypted source table and the ID of the decrypted source table; Summarizing the decrypted source table and the ID of the decrypted source table to obtain the decrypted data set.
3. The data processing method according to claim 1, characterized in that, the submitting the task processing set to the data processing engine and parsing the execution command according to the task command to generate a parsed task set includes: Obtaining the task processing set, and parsing the execution command through a pre-built parsing function to obtain a syntax tree; Searching and iterating the syntax tree to obtain a set of dimension tables corresponding to the syntax tree; Classifying the set of dimension tables according to the dimension table structure and non-dimension table structure to obtain a first set of dimension tables and a second set of dimension tables; Selecting dimension tables from the first set of dimension tables and the second set of dimension tables through a preset filtering condition to obtain a parsed task set.
4. The data processing method according to claim 1, characterized in that, the submitting the parsed task set to the data processing engine and running it to generate a processing result set includes: Obtaining the parsed task set, and starting a work management function and a task management function in the data processing engine; Use the data processing engine to transform the parsing task set into the form of a parallel data stream; Run the work management function and the task management function to perform parallel processing on the parsing task set of the parallel data stream and generate a processing result set.
5. A data processing device for implementing the data processing method according to any one of claims 1-4, characterized in that, the device includes: A data encryption module for setting and encrypting and editing a data set in a database of an interactive real-time computing platform through a query language to obtain an encrypted data set; A data decryption module for obtaining a task data set, decrypting the encrypted data set to obtain a decrypted data set, and adjusting parameters of the task data set and the decrypted data set to obtain a task processing set, where the task processing set includes an execution command and a task command; A data parsing module for submitting the task processing set to a data processing engine in the interactive real-time computing platform and parsing the execution command according to the task command to generate a parsing task set; A data processing module for submitting the parsing task set to the data processing engine and running to generate a processing result set.
6. An electronic device, characterized in that, the electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the data processing method according to any one of claims 1 to 4.
7. A computer-readable storage medium storing a computer program, characterized in that, the computer program, when executed by a processor, implements the data processing method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Data processing method, device and system
CN110502915A
Data processing method and device and storage medium
CN111382131A