Data processing system and method, electronic equipment and storage medium
By deploying independent cluster computing engines and data warehouses in the bank impairment system, the problems of insufficient computing resources and high storage pressure are solved, efficient data batch processing and real-time data updates are achieved, and the system's operation and maintenance efficiency and data accuracy are improved.
Patent Information
- Application Number
- CN202510837336.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-23
AI Technical Summary
In the bank impairment system, the insufficient computing resources, high operation and maintenance costs, high storage pressure caused by monthly detailed data processing, and high data quality and accuracy requirements, which cannot be effectively solved by the existing technology.
Deploy the computing engine in an independent cluster, build a data processing system with the data warehouse, execute batch tasks through independent computing resources, and transmit the batch results to the data warehouse in a structured storage format, decouple the computing and storage resources of the application database, update the configuration data tables in real time, and use target identifiers for intelligent routing and dynamic allocation of storage nodes.
It significantly improves data batch processing efficiency, shortens processing time, reduces operation and maintenance costs and storage pressure, ensures real-time and accuracy of data, and improves the dynamic adaptability and business stability of the system.
Smart Images

Figure CN120353801A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of data processing and cloud computing, and particularly to a data processing system, method, electronic device, and storage medium. Background Art
[0002] Currently, in the impairment system of a bank, it is necessary to provide impairment for the entire bank's assets on the application database side every month (the computing engine device is set on one side of the cluster where the application database is located, sharing computing resources and storage resources). The provision process is a detailed-by-detail one-by-one provision, so the data volume is large and the batch running time is long. The first-phase detailed data is about 20 million - 30 million records, and it takes about 2.5 - 3 hours to complete the entire monthly batch run. Due to the high requirements for data quality and data accuracy at the end of the month, multiple reruns may be involved. Three reruns will take about 7 - 9 hours. The computing resources of the application database are insufficient, the batch running efficiency is low, resulting in high operation and maintenance costs, long batch running time, and the data after batch running needs to be stored in the application database, resulting in a large storage pressure on the application database. Summary of the Invention
[0003] This application provides a data processing system, method, electronic device, and storage medium.
[0004] In one aspect, an embodiment of this application provides a data processing system, which includes: a data warehouse and a computing engine; the computing engine is deployed in an independent cluster; In response to receiving a target processing task, the computing engine obtains data to be processed from the data warehouse, and performs batch processing on the data to be processed based on configuration data corresponding to the data to be processed to obtain a batch processing result; After the target processing task is completed, the computing engine merges at least two sub-files to obtain a target file, and transmits the target file to the data warehouse. The sub-files are obtained by packing the preset number of batch processing results when the computing engine determines that the number of batch processing results is greater than the preset number, and the target file adopts a structured storage format; After receiving the target file, the data warehouse stores the batch processing results in the target file.
[0005] Wherein, the system further includes an application database: In response to receiving a target processing task, the computing engine reads a configuration data table from the application database and stores the configuration data table in memory.
[0006] Among them, the computing engine includes a storage device. When the computing engine determines that at least one piece of target information is received and the computing engine is not executing a processing task, it reads the update information corresponding to the target information from the application database, and updates the configuration data table stored in the storage device based on the update information. The target information indicates that the configuration data table in the application database has changed, and the configuration data table includes a plurality of configuration data.
[0007] Among them, the computing engine obtains at least one target identifier from the data to be processed, and obtains the configuration data corresponding to the data to be processed from the configuration data table in the memory based on the at least one target identifier. The target identifier at least includes a service identifier, and the service identifier represents the service type to which the data to be processed belongs.
[0008] Among them, the data warehouse further includes at least one storage node; The computing engine transfers the target file to the target storage node, and the target storage node is determined from the at least one storage node; After receiving the target file, the target storage node imports the batch processing result in the target file into the database, updates the data in the target data table of the database, and performs data refreshing so that the updated data in the target data table of the database becomes effective.
[0009] Among them, after the target processing task is completed, the computing engine generates a report based on the batch processing result and sends it to the application database; After receiving the report, the application database stores the report and displays it to the user.
[0010] Another aspect of the embodiments of the present application provides a data processing method. The method is applied to a data processing system, and the system includes a data warehouse and a computing engine. The computing engine is deployed in an independent cluster. The method includes: In response to receiving a target processing task, obtain data to be processed from the data warehouse; Perform batch processing on the data to be processed based on the configuration data corresponding to the data to be processed to obtain a batch processing result; After the target processing task is completed, merge at least two sub-files to obtain a target file, and transfer the target file to the data warehouse so that after the data warehouse receives the target file, it stores the batch processing result in the target file. The sub-files are obtained by packing the preset number of batch processing results when it is determined that the number of batch processing results is greater than the preset number, and the merged target file adopts a structured storage format.
[0011] Among them, the system further includes an application database, and the method further includes: In response to receiving a target processing task, read a configuration data table from the application database; Store the configuration data table in memory.
[0012] On the other hand, this application also provides an electronic device, including: A processor and a memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the data processing method.
[0013] On yet another aspect, this application provides a computer-readable storage medium storing a computer program for executing the data processing method.
[0014] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of this application, nor is it used to limit the scope of this application. Other features of this application will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] By referring to the accompanying drawings and reading the following detailed description, the above and other objects, features, and advantages of the exemplary embodiments of this application will become easily understandable. In the drawings, several embodiments of this application are shown in an exemplary rather than restrictive manner, where: In the drawings, the same or corresponding reference numerals represent the same or corresponding parts.
[0016] Figure 1 Shows a schematic structural diagram of a data processing system according to an embodiment of this application; Figure 2 Shows a schematic structural diagram of a data processing system according to another embodiment of this application; Figure 3 Shows a schematic structural diagram of a data processing system according to another embodiment of this application; Figure 4 Shows a flowchart of a data processing method according to an embodiment of this application; Figure 5 Shows a flowchart of a data processing method according to another embodiment of this application; Figure 6 Shows a flowchart of a data processing method according to another embodiment of this application; Figure 7 Shows a flowchart of a data processing method according to another embodiment of this application; Figure 8The flowchart of a data processing method according to another embodiment of the present application is shown; Figure 9 The flowchart of a data processing method according to another embodiment of the present application is shown; Figure 10 The schematic diagram of the composition structure of an electronic device according to an embodiment of the present application is shown. Detailed implementation manners
[0017] To make the objectives, features, and advantages of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present application.
[0018] To improve the batch processing efficiency, reduce the operation and maintenance costs, shorten the batch processing time, and at the same time reduce the storage pressure on the application database, an embodiment of the present application provides a data processing system, as Figure 1 shown. The system 100 includes: a data warehouse 101 and a computing engine 102; the computing engine 102 is deployed in an independent cluster.
[0019] In this embodiment, the data warehouse 101 can be selected from Hive (a data warehouse), Spark SQL (a data warehouse), etc., and the computing engine 102 can be selected from Hadoop MapReduce (a computing engine), Apache Spark (a computing engine), etc. In other implementation manners, the data warehouse 101 can be any other data warehouse, and the computing engine 102 can be any other computing engine. The device to which the computing engine 102 belongs is set in an independent cluster. Of course, in other implementation manners, the computing engine 102 can also be set in the cluster on the side of the data warehouse 101.
[0020] The computing engine 102, in response to receiving a target processing task, obtains the data to be processed from the data warehouse 101, and performs batch processing on the data to be processed based on the configuration data corresponding to the data to be processed, and obtains a batch processing result.
[0021] In this embodiment, the target processing task is issued by the user on the application database (front-end) side, and the target processing task can be a processing task that requires batch processing, such as data batch write-down processing, cross-system data fusion, AI model training, etc. In other implementation manners, the target processing task can be any other processing task that requires batch processing.
[0022] After receiving the target processing task, the computing engine 102 reads the corresponding data to be processed from the data warehouse 101 based on the target processing task, obtains the configuration data corresponding to the data to be processed, and then performs batch processing on the data to be processed based on the configuration data to obtain a batch processing result. When the target processing task is to perform batch write-down processing on data, the configuration data can be write-down parameters (including data such as general codes, write-down subject lists, and various adjustment coefficients) and computing engine parameters (including data such as judgment rule parameters, measurement rule parameters, and product measurement model configuration tables).
[0023] After the target processing task is completed, the computing engine 102 merges at least two sub-files to obtain a target file, and transmits the target file to the data warehouse 101. The sub-files are obtained by the computing engine 102 packing the preset number of batch processing results when it determines that the number of batch processing results is greater than the preset number. The merged target file adopts a structured storage format.
[0024] In this embodiment, when the number of batch processing results obtained by the computing engine 102 is greater than the preset number, these preset number of batch processing results are packed into sub-files. After the target processing task is completed, all the sub-files are merged to obtain a target file, and finally the target file is transmitted to the data warehouse 101 for storage.
[0025] For example, during the execution of the batch processing task by the computing engine 102, according to the paged batch processing results, each batch of batch processing results generates a parquet (a file format) file and stores it locally (about 7000 records per batch). After the target processing task is completed, the computing engine 102 starts a new process to merge the previously paged generated parquet files into a new standard parquet file. The computing engine 102 uploads the merged parquet file to the data warehouse 101 for storage.
[0026] The target file adopts a structured storage format.
[0027] For example, the target file adopts a binary structured storage format. The header includes fixed-length file metadata: file identification header (including information such as identifying the file type), the number of batch processing results. The main body stores a preset number of batch processing results in a preset order. The tail includes an index table. The index table stores a preset number of records in order. Each record contains the offset and data length of the corresponding batch processing result, enabling fast random access.
[0028] After receiving the target file, the data warehouse 101 stores the batch processing results in the target file.
[0029] In the above solution, by deploying the computing engine 102 in an independent cluster and constructing a data processing system 100 with the data warehouse 101, dual decoupling from the computing resources and storage resources of the application database is achieved. Since the computing engine 102 runs independently in an exclusive hardware resource environment, it can fully utilize the computing resources of the cluster to execute batch processing tasks, avoiding insufficient computing resources caused by sharing computing resources with the application database, thereby significantly improving the efficiency of data batch processing and shortening the execution duration of batch processing. At the same time, the batch processing results are directly stored in the data warehouse 101, completely eliminating the continuous occupation of the storage resources of the application database and significantly reducing the growth rate of the data stored in the application database. The architecture of the data processing system 100 systematically solves the performance bottleneck problem in the coupled deployment mode through resource isolation and data flow reconstruction, ensuring the stability of the front-end business system while improving the batch processing efficiency. And during the execution of the target processing task, the computing engine 102 no longer transmits the processing results one by one in the existing manner, but packs multiple batches of batch processing results into sub-files according to a preset quantity, and merges them into a complete target file after the target processing task ends and sends it to the data warehouse 101 uniformly. This significantly reduces the number of data transmissions, reduces the risk of network fluctuations caused by high-frequency small file transmissions, and at the same time reduces the write operation pressure on the data warehouse 101. The merged target file adopts a structured storage format, which not only preserves the integrity of the data but also avoids the problem of excessive storage space occupied by scattered files. Through the collaborative design of merged transmission and centralized storage, the system achieves a balance between processing speed and resource occupancy while ensuring data accuracy.
[0030] In an example of the present application, a data processing system is further provided, as Figure 2 shown, the system 100 further includes an application database 103.
[0031] In this embodiment, the application database 103 can be selected from Oracle (a database), MySQL (a database), SQL Server (a database), OceanBase (a database), etc. In other implementation manners, the application database 103 can be selected from any other database.
[0032] In response to receiving a target processing task, the computing engine 102 reads the configuration data table from the application database 103 and stores the configuration data table in the memory.
[0033] In this embodiment, since the application database 103 is the front-end database and the user can modify the configuration data corresponding to each processing task through front-end operations, the configuration data table is stored in the application database 103.
[0034] After receiving a target processing task, the computing engine 102 reads the configuration data table from the application database 103 and stores the configuration data table in the memory.
[0035] Since the user needs to perform front-end operations to modify the configuration data corresponding to each processing task, the configuration data table needs to be stored in the application database 103. In the prior art, usually a period (such as 1 day or 2 days) is set, and then the configuration data table is extracted from the application database 103 by means of bus data extraction and then loaded into the data warehouse 101 the next day. Therefore, the configuration data modified by the user through front-end operations cannot take effect in the data warehouse 101 in real time. In the above solution, when the computing engine 102 receives a target processing task, it reads the latest configuration data table from the application database 103 in real time and stores it in the memory, so that the configuration data modified by the user at the front end can take effect in real time. Through the above solution, the delay problem of traditional periodic synchronization can be eliminated, ensuring that the computing engine 102 always executes batch processing tasks based on the latest configuration data, and improving the system's response ability to dynamic adjustment and the real-time performance of processing tasks.
[0036] In an example of the present application, a data processing system is further provided, as Figure 2 shown. The computing engine includes a storage device 1021. When the computing engine 102 determines that it has received at least one piece of target information and the computing engine 102 is not executing a processing task, it reads the update information corresponding to the target information from the application database 103 and updates the configuration data table stored in the memory based on the update information. The target information indicates that the configuration data table in the application database 103 has changed, and the configuration data table includes a plurality of configuration data.
[0037] In this embodiment, the computing engine 102 no longer performs a full update of the configuration data table when receiving a target processing task, but performs an incremental update of the configuration data table stored in the storage device 1021 of the computing engine 102 when it receives at least one piece of target information sent by the application database 103 and the computing engine 102 is not executing a processing task.
[0038] The target information is the information generated and sent by the application database 103 to the computing engine 102 after the user modifies the configuration data table in the application database 103 through the front end, to notify the computing engine 102 that the configuration data table in the application database 103 has changed and the configuration data table in the storage device 1021 of the computing engine 102 needs to be updated.
[0039] When the computing engine 102 receives at least one piece of target information sent by the application database 103, it determines whether the computing engine 102 is currently executing a processing task. If so, it continues to wait. If not, it reads the update information corresponding to the target information from the application database 103 based on the target information, and then updates the configuration data table in the storage device 1021 of the computing engine 102 based on the update information.
[0040] It should be noted that the target information may also directly include the update information. Then, when the computing engine 102 receives at least one piece of target information sent by the application database 103 and the computing engine 102 is not executing a processing task, it can directly update the configuration data table in the storage device 1021 based on the update information in the target information without consuming additional communication time.
[0041] In the above solution, when the user modifies the configuration data at the front end, causing the configuration data table in the application database 103 to change, the application database 103 will automatically generate target information and send it to the computing engine 102. After receiving at least one piece of target information, once the computing engine 102 determines that it is in a state of not executing a processing task, it obtains the change information from the application database 103 based on the target information, and synchronously updates the configuration data table in the storage device 1021 based on the update information. This avoids the limitation of the traditional solution that must rely on a fixed period to update the configuration data table, and at the same time does not interfere with the running tasks due to real-time reading. Every adjustment of the configuration by the user at the front end can take effect quickly when the computing engine 102 is idle, which not only ensures the stability of the task execution process, but also realizes the precise connection between the configuration data change and the response of the computing engine 102, significantly improving the flexibility and efficiency of the system to dynamically adapt to business requirements.
[0042] In an example of the present application, a data processing system is also provided, as Figure 2 shown, the computing engine 102 obtains at least one target identifier from the data to be processed, and obtains the configuration data corresponding to the data to be processed from the configuration data table in the memory based on the at least one target identifier. The target identifier at least includes a service identifier, and the service identifier represents the service type to which the data to be processed belongs.
[0043] The computing engine 102 obtains at least one target identifier from the data to be processed, and obtains the configuration data corresponding to the data to be processed from the configuration data table in the memory based on the target identifier.
[0044] The target identifier at least includes a service identifier, and the service identifier represents the service type to which the data to be processed belongs. For example, the service identifier is impairment processing, real-time fusion, model training, etc. When the service identifier is impairment processing, it represents that the service type to which the data to be processed belongs is batch data impairment processing. When the service identifier is real-time fusion, it represents that the service type to which the data to be processed belongs is cross-system data fusion. When the service identifier is model training, it represents that the service type to which the data to be processed belongs is AI model training.
[0045] For different service types, the target identifier may also include other identifiers.
[0046] For example, when the service type is batch data impairment processing, the target identifier may also include an asset identifier, and the asset identifier may be corporate loans, credit card receivables, fixed assets, etc.
[0047] For another example, when the service type is cross-system data fusion, the target identifier may also include a data source identifier and a usage identifier. The data source identifier may be a core system, an external API (Application Programming Interface), a third-party database, etc., and the usage identifier may be risk rating, customer profiling, anti-fraud, etc.
[0048] For another example, when the service type is AI model training, the target identifier may also include a scenario identifier, and the scenario identifier may be credit scoring, image quality inspection, sales prediction, etc.
[0049] In the above solution, by reading the corresponding configuration data from the configuration data table pre-loaded in the memory through the target identifier carried in the data to be processed, accurate adaptation to different service scenarios can be achieved. The service identifier, as the core classification basis, enables multiple types of data to be processed simultaneously in the same batch processing task. The computing engine 102 loads the corresponding configuration data according to the service type of the data to be processed and switches the computing logic, avoiding the cumbersome operations of pre-manual classification or multiple scheduling in the traditional solution. Through the intelligent routing mechanism based on the target identifier, the system can still maintain high-efficiency and stable processing efficiency when facing complex service mixed processing requirements.
[0050] In an example of the present application, a data processing system is also provided, as Figure 3 shown, the data warehouse 101 further includes at least one storage node 1011.
[0051] The computing engine 102 transfers the target file to the target storage node, and the target storage node is determined from the at least one storage node 1011.
[0052] The computing engine 102 can determine the target storage node based on the current load status of each storage node 1011 in the data warehouse.
[0053] After receiving the target file, the target storage node imports the batch processing result in the target file into the database, updates the data in the target data table of the database, and performs data refreshing to make the updated data in the target data table of the database take effect.
[0054] For example, the computing engine 102 establishes communication with Hive through the Hive connection tool, executes the "Load Data Inpath" command to import the parquet file on the target storage node into the specified table of the Hive database. Then, the computing engine establishes communication with Impala (the Impala process running on the query node) through the Impala connection tool and executes the "data refreshing" command to make the newly added or updated data in the Hive database take effect immediately in the Impala node (query node).
[0055] In the above solution, by introducing a dynamic storage node allocation mechanism, the efficiency and reliability of the data storage and effect-taking process are further optimized. The system can determine the optimal storage node to receive the packaged file of the batch processing result according to the current load status, avoiding processing delays caused by excessive pressure on a single storage node. After receiving the file, the target storage node writes the result into the database in batches, significantly reducing the frequency of database operations compared with updating one by one, while maintaining the stability of the data table structure. After the writing is completed, through the data refreshing mechanism, it is ensured that the update result takes effect immediately and the external service is unaware, effectively shortening the interval from data storage to business availability. Through node selection and batch processing, while maintaining data consistency, the execution efficiency of large-scale data updates is improved.
[0056] In an example of the present application, a data processing system is further provided. As Figure 2 shown, after the target processing task is completed, the computing engine 102 generates a report based on the batch processing result and sends it to the application database 103.
[0057] After the target processing task is completed, the computing engine 102 generates a report based on all the obtained batch processing results and sends it to the application database 103. The report can represent the overall information of the business related to the target processing task.
[0058] For example, when the business type is batch data impairment processing, the report may include data such as impairment detail information (e.g., impairment amount, impairment ratio of assets, etc.), risk distribution information (e.g., risk level of assets), execution verification information (e.g., deviation between execution result and expectation), and impact assessment information (e.g., impact value of impairment on departments and products).
[0059] For another example, when the business type is cross-system data fusion, the report may include data such as data quality information (e.g., integrity, consistency, accuracy, etc. of the fused data), field mapping information (e.g., field mapping table between different systems), correlation analysis information (e.g., impact of newly added data dimensions after fusion on the business), and processing efficiency information (e.g., data loading duration, data fusion duration, etc.).
[0060] For another example, when the business type is AI model training, the report may include data such as performance metric information (e.g., accuracy, recall, etc. of the training set and validation set) and training monitoring information (e.g., loss function convergence curve, training duration, resource utilization, etc.).
[0061] After receiving the report, the application database 103 stores the report and presents it to the user.
[0062] After receiving the report, the application database 103 stores the report and visualizes it through a preset template or other means, and presents it to the user.
[0063] In the above solution, after the target processing task is completed, the computing engine 102 generates a report based on all batch processing results and transmits it to the application database 103 for storage, without manual export or secondary processing. After receiving the report, the application database 103 can not only retain the original records for a long time, but also visualize them through a preset template or other means and present them to the user. The user can immediately view the report to understand the relevant information of the target processing task. This ensures the coherence of data from processing to presentation, reduces the delay and error in the intermediate links, and avoids the performance loss caused by frequent access of external systems to the data warehouse 101. While ensuring data security, it significantly improves the timeliness and usability of result delivery.
[0064] To improve the batch processing efficiency, reduce the operation and maintenance cost, shorten the batch processing time, and at the same time reduce the storage pressure on the application database, an embodiment of the present application provides a data processing method, which is applied to a data processing system, and the system includes a data warehouse and a computing engine, and the computing engine is deployed in an independent cluster, as Figure 4 shown, and the method includes: Step 201, in response to receiving a target processing task, obtain the data to be processed from the data warehouse.
[0065] After receiving the target processing task, the computing engine reads the corresponding data to be processed from the data warehouse based on the target processing task.
[0066] Step 202: Perform batch processing on the data to be processed based on the configuration data corresponding to the data to be processed, and obtain a batch processing result.
[0067] Obtain the configuration data corresponding to the data to be processed, and then perform batch processing on the data to be processed based on the configuration data to obtain a batch processing result.
[0068] Step 203: After the target processing task is completed, merge at least two sub-files to obtain a target file, and transmit the target file to the data warehouse, so that after receiving the target file, the data warehouse stores the batch processing result in the target file. The sub-files are obtained by packing the preset number of batch processing results when it is determined that the number of batch processing results is greater than the preset number. The merged target file adopts a structured storage format.
[0069] In an example of the present application, a data processing method is further provided. The system further includes an application database. As Figure 5 shown, the method further includes: Step 301: In response to receiving a target processing task, read a configuration data table from the application database.
[0070] The user can issue a target processing task at the front end. After the computing engine receives the target processing task issued by the user, it reads the configuration data table from the application database.
[0071] Step 302: Store the configuration data table in the memory.
[0072] Store the configuration data table in the memory, and then the corresponding configuration data can be read from the memory to execute the target processing task.
[0073] In an example of the present application, a data processing method is further provided. The computing engine includes a storage device. As Figure 6 shown, the method further includes: Step 401: Determine that at least one piece of target information is received and the computing engine is not executing a processing task, then read the update information corresponding to the target information from the application database. The target information indicates that the configuration data table in the application database has changed.
[0074] When the user modifies the configuration data table in the application database through front-end operations, the application database sends target information to the computing engine. After receiving at least one piece of target information, the computing engine determines whether the computing engine is currently executing a processing task. If so, it continues to wait. If not, it reads the update information corresponding to the target information from the application database.
[0075] Step 402: Update the configuration data table stored in the storage device based on the update information. The configuration data table includes multiple pieces of configuration data.
[0076] Perform an incremental update on the configuration data table stored in the memory based on the update information.
[0077] In an example of the present application, a data processing method is also provided. As Figure 7 shown, the method further includes: Step 501: Obtain at least one target identifier from the data to be processed. The target identifier includes at least a service identifier, and the service identifier represents the service type to which the data to be processed belongs.
[0078] Obtain at least one target identifier from the data to be processed.
[0079] Step 502: Obtain the configuration data corresponding to the data to be processed from the configuration data table in the memory based on the at least one target identifier.
[0080] Obtain the configuration data corresponding to the data to be processed from the configuration data table in the memory based on the target identifier.
[0081] In an example of the present application, a data processing method is also provided. As Figure 8 shown, the method further includes: Step 601: Transmit the target file to the target storage node, so that after receiving the target file, the target storage node imports the batch processing result in the target file into the database. The target storage node is determined from the at least one storage node.
[0082] Obtain the current load status of multiple storage nodes in the data warehouse, and determine the optimal storage node as the target storage node from these storage nodes based on the current load status. The optimal storage node can be the storage node with the lowest load or not executing other tasks.
[0083] After determining the target storage node, transmit the target file to the target storage node.
[0084] After the target storage node receives the target file, import the batch processing result in the target file into the database for storage.
[0085] Step 602: Send a refresh instruction to the target storage node so that the target storage node refreshes data based on the data in the target data table of the database.
[0086] The computing engine sends a refresh instruction to the target storage node.
[0087] After receiving the refresh instruction, the target storage node refreshes data based on the data in the target data table of the database. The target data table is the data table into which the batch processing result in the target file is imported.
[0088] In an example of the present application, a data processing method is further provided. As Figure 9 shown, the method further includes: Step 701: After the target processing task is completed, generate a report based on the batch processing result.
[0089] After the target processing task is completed, the computing engine generates a report corresponding to the service type of the target processing task based on all the obtained batch processing results.
[0090] Step 702: Send the report to the application database so that the application database stores the report after receiving it and displays it to the user.
[0091] Send the report to the application database.
[0092] After receiving the report, the application database stores the report and visualizes the report through a preset template or other means and displays it to the user.
[0093] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device and a readable storage medium.
[0094] Figure 10 FIG. shows a schematic block diagram of an example electronic device 800 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0095] As Figure 10As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 802 or computer programs loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0096] Multiple components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0097] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as a data processing method. For example, in some embodiments, the data processing method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the data processing method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the data processing method in any other appropriate manner (e.g., by means of firmware).
[0098] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0099] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0100] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0101] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0102] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0103] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client - server relationship is generated by computer programs running on the respective computers and having a client - server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating blockchain.
[0104] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. No limitation is imposed herein.
[0105] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" can explicitly or implicitly include at least one such feature. In the description of this disclosure, "a plurality" means two or more unless otherwise specifically defined.
[0106] As described above, it is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claims described above.
Claims
1. A data processing system, the system comprising: Data warehouse and computing engine; The computing engine is deployed in an independent cluster; In response to receiving a target processing task, the computing engine obtains data to be processed from the data warehouse, and performs batch processing on the data to be processed based on configuration data corresponding to the data to be processed, to obtain a batch processing result; After the target processing task is completed, the computing engine merges at least two sub-files to obtain a target file, and transmits the target file to the data warehouse. The sub-files are obtained by the computing engine by packing a preset number of batch processing results when it is determined that the number of batch processing results is greater than the preset number. The target file adopts a structured storage format; After receiving the target file, the data warehouse stores the batch processing results in the target file.
2. The system according to claim 1, the system further includes an application database: In response to receiving a target processing task, the computing engine reads a configuration data table from the application database, and stores the configuration data table in memory.
3. The system according to claim 2, the computing engine includes a storage device. When the computing engine determines that it has received at least one piece of target information and the computing engine is not executing a processing task, it reads update information corresponding to the target information from the application database, and updates the configuration data table stored in the storage device based on the update information. The target information indicates that the configuration data table in the application database has changed, and the configuration data table includes a plurality of configuration data.
4. The system according to claim 3, the computing engine obtains at least one target identifier from the data to be processed, and obtains configuration data corresponding to the data to be processed from the configuration data table in memory based on the at least one target identifier. The target identifier at least includes a service identifier, and the service identifier represents the service type to which the data to be processed belongs.
5. The system according to claim 1, the data warehouse further includes at least one storage node; The computing engine transmits the target file to a target storage node, and the target storage node is determined from the at least one storage node; After receiving the target file, the target storage node imports the batch processing results in the target file into a database, updates the data in the target data table of the database, and performs data refreshing to make the updated data in the target data table of the database take effect.
6. The system according to claim 2, after the target processing task is completed, the computing engine generates a report based on the batch processing result, and sends the report to the application database; After receiving the report, the application database stores the report and displays it to the user.
7. A data processing method, the method is applied to a data processing system, the system includes a data warehouse and a computing engine, the computing engine is deployed in an independent cluster, and the method includes: In response to receiving a target processing task, obtain data to be processed from the data warehouse; Perform batch processing on the data to be processed based on the configuration data corresponding to the data to be processed, and obtain a batch processing result; After the target processing task is completed, merge at least two sub-files to obtain a target file, and transmit the target file to the data warehouse, so that after receiving the target file, the data warehouse stores the batch processing result in the target file. The sub-files are obtained by packing the preset number of batch processing results when it is determined that the number of batch processing results is greater than the preset number, and the target file adopts a structured storage format.
8. The method according to claim 7, wherein the system further comprises an application database, and the method further comprises: In response to receiving a target processing task, read a configuration data table from the application database; Store the configuration data table in memory.
9. An electronic device, comprising: A processor and a memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the data processing method according to claims 7-8.
10. A computer-readable storage medium, wherein the storage medium stores a computer program for executing the data processing method according to claims 7-8.
Citation Information
Patent Citations
Data warehouse based medical data integration method and system
CN104462082A
Cloud storage method, cloud platform and computer readable storage medium
CN108712483A
Calling external functions from a data warehouse
CN113490918A
Data processing method and device and computer readable storage medium
CN115544182A
Stream batch integrated data processing method and device, equipment and storage medium
CN117271012A