Data processing method, electronic device, and storage medium
By configuring the data processing and write processing functions of the batch node on multiple child nodes, the problem that batch jobs cannot utilize multi-server computing resources is solved, efficient batch execution is achieved, and development costs and system intrusion is reduced.
Patent Information
- Application Number
- PCT/CN2024/084087
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-15
- Filing Date
- 2024-03-27
- Publication Date
- 2025-06-19
AI Technical Summary
In the prior art, batch jobs cannot synchronously utilize the computing resources of multiple servers, resulting in low execution efficiency, and the distributed parallel computing framework is complex and difficult to use. It is necessary to make major changes to the original batch job code structure to increase development costs.
By configuring the data processing and write processing functions of the batch node to multiple different child nodes, each child node is set on another server other than the batch node, the master node can read the pending data and generate batch tasks, distribute it to the child node for execution through the interface, and write the processed data to the database.
It realizes that the batch node effectively calls the operation resources of multiple servers, improves execution efficiency, reduces the computing load of the master node, and reduces system intrusion and development costs.
Smart Images

Figure CN2024084087_19062025_PF_FP_ABST
Abstract
Description
Data processing method, electronic device and storage medium
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 15, 2023, with application number 202311734745.6 and application name “Data processing method, electronic device and storage medium”. The entire contents of the above application are incorporated by reference into this application. Technical Field
[0002] The present invention relates to the field of computer technology, and in particular to a data processing method, an electronic device, and a computer-readable storage medium. Background Art
[0003] Batch processing is a processing model in which a computer program executes a series of tasks based on a batch of inputs. Batch processing can process large amounts of data in a single execution, takes a long time, and does not require a connection to the caller. Results are typically communicated to the caller via reports, so real-time and interactivity requirements are low. Batch processing is suitable for batch data processing in various business systems, such as reconciliation between insurance systems and banks.
[0004] In existing technologies, batch processing jobs are typically executed sequentially on a single server, i.e., batch processing is performed linearly on a single server. However, batch processing of data cannot synchronously utilize the computing resources of multiple servers, resulting in low execution efficiency.
[0005] To enable batch processing jobs to be executed in parallel across multiple servers, distributed parallel computing frameworks have been proposed. These frameworks can concurrently process multiple data processing tasks and distribute them across multiple servers to leverage their computing resources, such as MapReduce and Spark. However, distributed parallel computing frameworks are primarily used for large-scale distributed computing and are not suitable for batch processing jobs. Furthermore, many functions within distributed parallel computing frameworks require the development of corresponding functional logic code, making functional expansion more complex. For example, data dependencies and configurations must be manually managed. Consequently, distributed parallel computing frameworks are difficult for users to use.
[0006] To avoid these issues, batch schedulers have been proposed that split batch jobs into multiple concurrent tasks, providing technical support for distributed execution. However, this task splitting often requires significant changes to the existing batch job code structure, resulting in high development costs.
[0007] Summary of the Invention
[0008] An embodiment of the present application proposes a data processing method, an electronic device, and a computer-readable storage medium. The method configures the data processing function and the write processing function in a batch processing node to one or more child nodes, each of which is set on a server other than the batch processing node. When the batch processing node reads the data to be processed, the child nodes on other servers can be called to process the converted data, and the processed data can be written to the database, so that the batch processing node can effectively call the operating resources of multiple servers to perform data batch processing, thereby improving execution efficiency.
[0009] In a first aspect, an embodiment of the present application provides a data processing method, which is applied to a data processing system, wherein the data processing system includes a master node and at least one child node, wherein the master node and the at least one child node are deployed on different electronic devices, and the master node calls at least one child node through an interface, and the method includes: reading the data to be processed through the master node, and generating batch processing tasks based on the data to be processed; distributing the batch processing tasks generated by the master node to at least one child node through the interface; executing the batch processing tasks through at least one child node to obtain processed data, and writing the processed data into a database, and at least one child node cannot generate batch processing tasks based on the processed data or the data to be processed.
[0010] It should be understood that the aforementioned batch processing tasks are tasks that perform various batch processing operations on data. Here, the master node can generate batch processing tasks, while the child nodes cannot. Consequently, the master node is configured with batch processing code, while the child nodes do not need to be configured with batch processing code. However, the child nodes can share the batch processing computational tasks performed by the master node, reducing the computational load on the master node.
[0011] Specifically, by assigning the data processing and write processing functions of a batch processing node to multiple different subnodes, each subnode can provide data processing and write services to the batch processing node. Each subnode is located on a server other than the batch processing node. Once the batch processing node acquires and converts the data to be processed, it can call on subnodes on other servers to process the converted data and write the processed data to the database. This allows the batch processing node to effectively utilize the operating resources of multiple servers for data batch processing, improving execution efficiency.
[0012] In a possible implementation of the first aspect above, the batch processing task includes a batch execution task, and generating the batch processing task based on the data to be processed further includes: generating a batch execution task based on the processed data, wherein the batch execution task includes batch processing and batch writing.
[0013] It should be understood that batch processing tasks may include batch execution tasks. In some embodiments, the above-mentioned batch execution tasks may at least include batch processing and batch data writing. For example, the converted data can be processed according to the execution scenario to obtain processed data, and the processed data can be written into the database in batches.
[0014] In a possible implementation of the first aspect, distributing the batch processing tasks generated by the master node to at least one child node through an interface includes distributing the batch execution tasks to at least one child node through an interface.
[0015] It should be understood that the batch execution tasks that can be executed by the child nodes in the batch processing tasks will be distributed to at least one child node for execution through the interface.
[0016] In a possible implementation of the first aspect above, the system also includes a load balancing unit, and distributing batch execution tasks to at least one child node includes: distributing batch execution tasks to at least one child node through the load balancing unit, wherein the load balancing unit is used to adjust the operating load of at least one child node.
[0017] It should be understood that the load balancing unit connects the master node and the sub-node layer, and is used to distribute the batch execution tasks obtained from the execution unit to the relatively idle sub-nodes. This can better balance the computing load between different sub-nodes and avoid repeatedly distributing batch execution tasks to the same sub-node, thereby effectively improving the computing efficiency of the master node.
[0018] In some embodiments, the master node detects that a processing function and / or a write function is about to be run on the converted data and can send the converted data to the load balancing unit without passing through the execution unit.
[0019] In a possible implementation of the first aspect above, the master node further includes an execution unit, and the method further includes: configuring execution capability for at least one child node so that at least one child node can execute batch execution tasks, wherein the execution capability includes data processing capability and writing capability.
[0020] It should be understood that the execution unit can refer to the batch processing code set in the master node, process data and write the processed data in batches, and generate batch execution tasks in batches.
[0021] In a possible implementation of the first aspect above, batch processing tasks generated by the master node are distributed to at least one child node through an interface, including: using the master node to call the interface based on the RESTful protocol or HTTP protocol to call at least one child node.
[0022] In some embodiments, the architectural style of the interface may be representational state transfer (REST), and the interface may be a RESTful API.
[0023] In a possible implementation of the first aspect above, the load balancing unit, the interface, and the at least one sub-node are all deployed through a preset microservice architecture.
[0024] In some embodiments, the load balancing unit, interface, and sub-node layer can be provided by a microservice architecture. The load balancing unit can be a load balancer in the microservice architecture, which can be used to distribute access to various computing resources, such as multiple sub-nodes. The load balancer can prevent batch execution tasks from being sent to servers that cannot operate normally, prevent resource overloading, and eliminate single points of failure. In a possible implementation of the first aspect above, the method further includes: monitoring the execution process of the batch tasks by presetting the microservice architecture and recording monitoring information.
[0025] In the second aspect, an embodiment of the present application also provides a data processing system, which includes a main node and at least one child node, wherein the main node and the at least one child node are deployed on different electronic devices, the main node calls at least one child node through an interface, and the main node is used to read the data to be processed and generate batch processing tasks based on the data to be processed; the interface is used to distribute the batch processing tasks generated by the main node to at least one child node; the at least one child node is used to execute the batch processing tasks, obtain the processed data, and write the processed data into a database, and at least one child node cannot generate batch processing tasks based on the processed data or the data to be processed.
[0026] In a possible implementation of the second aspect above, the data processing system further includes a load balancing unit, wherein the load balancing unit is configured to adjust an operating load of at least one subnode.
[0027] In a possible implementation of the second aspect above, the master node further includes an execution unit, and the execution unit is used to configure execution capabilities for at least one child node so that at least one child node can execute batch processing tasks, wherein the execution capabilities include data processing capabilities and writing capabilities.
[0028] In a possible implementation of the second aspect above, the load balancing unit, the interface, and the at least one sub-node are all deployed within a preset microservice architecture.
[0029] In a possible implementation of the second aspect above, the preset microservice architecture is further used to monitor the execution process of the batch task through the preset microservice architecture and record the monitoring information.
[0030] In a third aspect, an embodiment of the present application further provides a computer-readable storage medium, characterized in that instructions are stored on the storage medium, which, when executed on a computer, enable the computer to execute the data processing method provided by the above-mentioned first aspect and various possible implementations.
[0031] In a fourth aspect, an embodiment of the present application further provides a computer program product, characterized in that it includes a computer program / instruction, which, when executed by a processor, implements the data processing method provided by the above-mentioned first aspect and various possible implementations.
[0032] The beneficial effects of the second to fourth aspects mentioned above can be found in the relevant descriptions of the first aspect and various possible implementations of the first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] FIG1 shows a schematic diagram of a batch processing scenario according to some embodiments of the present application;
[0034] FIG2 shows a structural block diagram of a batch processing node according to some embodiments of the present application;
[0035] FIG3 shows a block diagram of a node structure of a batch process according to some other embodiments of the present application;
[0036] FIG4 shows a flow chart of a data processing method according to some embodiments of the present application;
[0037] FIG5 shows a block diagram of a node structure of a batch process according to some other embodiments of the present application;
[0038] FIG6 shows a structural block diagram of a method for distributing tasks through a message middleware according to some embodiments of the present application;
[0039] FIG7 shows a schematic flow chart of a data processing method according to other embodiments of the present application;
[0040] FIG8 shows a block diagram of a server 200 according to an embodiment of the present application. DETAILED DESCRIPTION
[0041] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described in detail below with reference to the accompanying drawings and specific implementation methods.
[0042] Illustrative embodiments of the present application include, but are not limited to, data processing methods, electronic devices, and computer-readable storage media.
[0043] It can be understood that the electronic device applicable to the present application can be a server, wherein the applicable server can be a cloud server, a physical server, a large bandwidth server, a high-defense server, a dedicated line server or a group server and other rental type servers. In addition, the applicable server can be a complex instruction set computing (CISC) architecture server or a reduced instruction set computing (RISC) architecture server, and there is no limitation here.
[0044] In order to facilitate understanding of the technical solutions provided by the embodiments of the present application, the meanings of some related field terms involved in the embodiments of the present application are explained below.
[0045] (1) Spring Batch is a lightweight batch processing framework provided by the Spring Framework. It can efficiently process data in batches without user interaction. It is suitable for regular application scenarios where complex business rules are repeatedly processed in large data sets. For example, it can be used in application scenarios such as policy renewal payment processing, batch reconciliation between insurance business systems and banks, regulatory data reporting processing, and reinsurance processing.
[0046] The following describes the batch processing scenario in detail with reference to Figure 1.
[0047] Figure 1 shows a schematic diagram of a batch processing scenario according to some embodiments of the present application. Figure 2 shows a structural block diagram of a batch processing node according to some embodiments of the present application.
[0048] Referring to Figure 1 , a user may be an administrative user of a terminal 100. Terminal 100 may be connected to a server 200, which may be a cloud server. Server 200 may provide a storage environment 201. This storage environment 201 may provide a database cluster storage environment and may include multiple batch processing nodes 300 for performing read and write operations on the database cluster. Multiple batch processing nodes 300 may be used to run multiple batch processing processes simultaneously, thereby enabling asynchronous read and write operations on the database cluster and improving batch processing efficiency. It should be understood that the batch processing nodes 300 may include multiple servers to enable data processing such as read and write operations on the database cluster.
[0049] In some embodiments, the server 200 may be a server cluster, and the server cluster may provide a microservice framework structure.
[0050] As you can understand, a server cluster allows multiple servers to centrally provide the same service. Clients can treat the server cluster as a single server when interacting with it. Compared to clients invoking services provided by multiple servers, using a server cluster to provide services to clients can achieve faster service responses. Furthermore, a server cluster can leverage multiple computers for parallel computing, achieving higher computation speeds and improving batch processing efficiency.
[0051] The microservices architecture is a method for partitioning a monolithic application into a set of smaller services. Each service runs in its own independent process, and services communicate with each other through well-defined application programming interfaces (APIs). This architectural design allows teams to independently develop and deploy their services, enhancing system scalability.
[0052] The microservice can run in its own program and communicate via a lightweight HTTP API. In a microservices architecture, it's possible to add functionality to a specific service without affecting the overall process architecture, which facilitates efficient expansion of system functionality.
[0053] 2 , in some embodiments, a batch processing node 300 may be used to perform batch processing on data to be processed. The batch processing node 300 includes a job step component 301, which is used to perform data processing tasks. The job step component 301 may generally include a reader unit for reading data to be processed, a processor unit for processing the data to be processed, and a writer unit for writing processed data. The writer unit may write the processed data to a database.
[0054] It should be understood that, to facilitate batch processing of data to be processed, batch processing node 300 can process large amounts of data to be processed using multiple job step components 301 and write the data into a database in batches. However, batch processing node 300 is typically located on a single server. Therefore, using batch processing node 300 to perform data batch processing makes it difficult to call upon multiple servers to complete the batch processing tasks, and the computing resources of multiple servers cannot be synchronously utilized, resulting in low execution efficiency.
[0055] In order to solve the above problems, the present application proposes a data processing method, which configures the data processing function and write processing function of the batch processing node (such as the data processing function of the processing unit shown in Figure 2 above and the write processing function that can be provided by the write unit) to multiple different sub-nodes, wherein each sub-node is set on a server other than the batch processing node (for example, a server other than the batch processing node in the server). The batch processing node obtains and converts the data to be processed, and then calls the sub-nodes on other servers to process the converted data and write the processed data into the database. In this way, the batch processing node can effectively call the operating resources of multiple servers to perform data batch processing, thereby improving execution efficiency.
[0056] In some embodiments, the above-mentioned sub-nodes can be deployed through a microservice framework, and the microservice framework can set a corresponding calling interface for each sub-node. When the reading unit calls the sub-node, it only needs to call the calling interface exposed by the sub-node to perform batch processing on the read data to be processed through the sub-node.
[0057] It can be understood that the above-mentioned database is an example of a storage location designated for processed data. The processed data can also be stored in other storage locations, for example, in a designated storage location on a cloud server or in a designated storage location on a terminal. No specific restrictions are made here.
[0058] The specific implementation process of a data processing method proposed according to an embodiment of the present application is described in detail below with reference to relevant drawings.
[0059] Example 1
[0060] FIG3 shows a block diagram of a node structure of batch processing according to some other embodiments of the present application.
[0061] 3 , batch processing node cluster 300A may include a master node 302 and a subnode layer 303, wherein master node 302 includes at least a read unit 021, a conversion unit 022, and an execution unit 023. The interface may include one or more interfaces, and the subnode layer 303 may include one or more subnodes.
[0062] Among them, the reading unit 021 is used to read the data to be processed, the conversion unit 022 is used to convert the read data to be processed, and the execution unit 023 is used to provide processing functions and writing functions, such as processing the data converted by the conversion unit 022 and writing the processed data into the database.
[0063] The interface can be used to provide the master node 302 with the functions provided by each sub-node in the sub-node layer 303, such as processing functions and writing functions. In some embodiments, the interface within the sub-node layer 303 can distribute the batch processing tasks generated by the master node 302 to each sub-node in the sub-node layer 303, so that each sub-node can execute the batch execution tasks in the batch processing tasks. In some embodiments, the master node 302 can run the functions of the sub-node corresponding to the interface by calling the above interface and perform the corresponding tasks, such as executing the batch execution tasks to obtain processed data and writing the processed data into a database.
[0064] In some embodiments, the interface can be provided as part of the software corresponding to the microservice architecture on the electronic device where one or more sub-nodes are located.
[0065] In other embodiments, the interface may be independently deployed on an electronic device other than the electronic device (eg, a server) where the master node and the sub-node are located.
[0066] The sub-node layer 303 includes sub-nodes 1 to n. Each sub-node exposes an access address through an interface to provide a write function for the master node 302 through interfaces 1 to n respectively.
[0067] It should be understood that the above-mentioned master node 302 may also include other units for executing different batch processing logics, for example, performing asynchronous reading, asynchronous processing and asynchronous writing on the data to be processed, and for example, performing batch processing when the number of read data to be processed meets the first value, etc., which is not specifically limited here.
[0068] It can be understood that the execution unit 023 in the batch processing node cluster 300A sends the converted data to the child node layer 303 for processing by calling the interface. During this process, the master node 302 can create a batch processing task and distribute the batch processing task to the child node layer 303 by calling the interface. The child nodes 1 to n in the child node layer 303 then execute part of the functions of the batch processing task and ultimately write the processed data into the database.
[0069] In some embodiments, the execution function of the execution unit 023 in the master node 302 can be distributed to one or more sub-nodes within the sub-node layer 303, so that the execution unit 023 does not need to perform batch processing of data and batch writing of data to the database. Instead, the sub-nodes within the sub-node layer 303 batch process the converted data and write the processed data in batches to the database, thereby reducing the operating load of the master node 302.
[0070] It should be understood that master node 302 can provide batch processing capabilities and distribute the above-mentioned execution functions to the sub-node layer 303. Therefore, the batch processing code only needs to be deployed in master node 302, and there is no need to deploy batch processing code in each sub-node. Furthermore, the batch processing node cluster 300A is less invasive and easier to implement for the microservice architecture.
[0071] It should be noted that the above-mentioned node may be a server included in the server.
[0072] The specific implementation method flow of the batch processing node architecture shown in FIG3 is described in detail below in conjunction with FIG4 .
[0073] Figure 4 shows a flow chart of a data processing method according to some embodiments of the present application. It is understood that the execution entity of each step of the process shown in Figure 4 can be the server 200, and the execution entity of a single step will not be described in detail.
[0074] Referring to Figure 4, the specific process steps include:
[0075] S401, reading the data to be processed through the master node.
[0076] It can be understood that the master node can be the master node 302 shown in FIG. 3 above, and the server 200 can read the data to be processed through the master node 302 and generate corresponding batch processing tasks.
[0077] In some embodiments, the master node 302 can be connected to a microservice framework, which can provide execution services to the master node 302 by providing an interface. The execution services include processing capabilities and write services, wherein the processing capabilities and write capabilities are allocated to each subnode in the subnode layer 303. Therefore, when the master node 302 uses the microservice framework to call the execution service to perform batch execution processing on the read data to be processed, a large number of batch execution tasks can be distributed to the subnodes corresponding to the interface by calling the interface. The subnodes execute the batch execution tasks to process the data and write the processed data into the database, thereby reducing the computing load of the master node 302.
[0078] For example, in the policy renewal payment processing scenario, the server 200 can distribute the renewal records to each child node through the interface through the main node 302 for processing and writing into the database; in the scenario of batch reconciliation processing between the insurance business system and the bank, the server 200 can distribute the reconciliation statement information to each child node through the interface through the main node 302 for processing and writing into the database; in the regulatory data reporting processing scenario, the server 200 can distribute the regulatory data and information reporting records to be reported to each child node through the interface through the main node 302 for processing and writing into the database; in the reinsurance processing scenario, the server 200 can distribute the reinsurance policy to each child node through the interface through the main node 302 for processing and writing into the database.
[0079] It should be noted that the aforementioned batch write task can be part of a batch processing task. The master node 302 can include batch processing code to generate a batch processing task including the batch write task. The master node 302 can call an interface provided by the microservice framework to distribute the batch write task to each child node configured with an execution service. The child nodes then execute the batch write task, reducing the computational load on the master node 302.
[0080] In some embodiments, the architectural style of the interface may be representational state transfer (REST), and the interface may be a RESTful API.
[0081] S402: Convert the data to be processed through the master node to obtain converted data.
[0082] It can be understood that the master node can provide a batch conversion function, and the server 200 can batch convert the data to be processed through the master node to obtain the converted data.
[0083] S403: Call the target child node through the first interface.
[0084] It can be understood that the server 200 can call the corresponding child node through the main node to specify one of all interfaces to distribute the batch execution task corresponding to writing the processed data into the database to the target child node, and execute the batch execution task through the target child node to process the converted data and write the processed data into the database.
[0085] It should be understood that the target child node does not need to deploy batch processing related code, and can directly execute the acquired batch execution task to process the converted data and write the processed data into the database. For example, corresponding to the scenario of insurance policy reconciliation processing, the server 200 can generate a batch processing task through the master node 302, and the batch processing task includes a batch execution task for reconciling the large amount of insurance policy data acquired, obtaining the reconciled insurance policy data and reconciliation result data, and writing the reconciled insurance policy data and reconciliation result data into the database. The target child node can then directly execute the batch write task to write the reconciled insurance policy data and reconciliation result data into the database.
[0086] S404: Process the converted data through the target child node and write the processed data into the database to complete the batch processing.
[0087] It should be understood that each sub-node may correspond to a server, and each sub-node may provide computing resources.
[0088] Therefore, the server 200 can process the converted data through the target child node and write the processed data into the database, reducing the operating load of the main node, effectively allocating the resources of multiple servers, and thus improving the batch processing efficiency.
[0089] For example, in a policy reconciliation scenario, the server acting as the target child node can directly execute the batch execution task, reconciling the large batch of policy data obtained, obtaining the reconciled policy data and reconciliation results, and writing the reconciled policy data and reconciliation results to the database. This eliminates the need to return the processed data to the master node or execute the batch execution task, reducing the master node's operational load and effectively allocating resources across multiple servers, thereby improving batch processing efficiency.
[0090] It can be understood that through the specific implementation process of steps S401 to S404 above, the embodiment of the present application can, on the one hand, set the batch processing-related code only on the master node, without writing the batch processing code in the child nodes. Instead, the execution function of the master node is distributed to multiple child nodes, which is less invasive to the system and easier to use. On the other hand, the data converted by the master node can be processed and written to the database by multiple child nodes respectively, reducing the operating load of the master node and effectively improving the efficiency of batch processing.
[0091] Example 2
[0092] In some embodiments, in order to better distribute the batch write tasks, the server 200 may utilize the load balancing unit 304 to distribute the batch write tasks of the master node 302 to each child node.
[0093] FIG5 shows a block diagram of a node structure of batch processing according to some further embodiments of the present application.
[0094] 5 and 3 , a batch processing node cluster 300B may include a master node 302, a load balancing unit 304, and a subnode layer 303. The master node 302 includes at least a read unit 021, a conversion unit 022, and an execution unit 023. The subnode layer 303 may include one or more subnodes, each of which may correspond to an interface.
[0095] Among them, the specific implementation of the reading unit 021, the conversion unit 022 and the execution unit 023 can refer to the specific implementation of the reading unit 021, the conversion unit 022 and the execution unit 023 in Figure 3 above. For the sake of simplicity, they are not repeated here.
[0096] The load balancing unit 304 is connected to the master node 302 and the sub-node layer 303, and is used to distribute the batch execution tasks obtained from the execution unit 023 to the relatively idle sub-nodes. Compared with the above embodiment 1, the computing load between different sub-nodes can be better balanced, and the repeated distribution of batch execution tasks to the same sub-node can be avoided, thereby effectively improving the computing efficiency of the batch processing node group 300B.
[0097] In some implementations, batch processing of data can be implemented using a batch processing component within the Spring framework (i.e., Spring Batch). Spring Batch itself also has certain task distribution capabilities. For example, referring to FIG6 , Spring Batch can provide a message middleware 306 having a master node 302′. Master node 302′ can include a read unit 021, a processing unit 022′, and a write unit 023′. Furthermore, batch processing tasks in master node 302′ can be distributed to child nodes 1′ through n′ via message middleware 306. However, using message middleware 306 to distribute batch processing tasks requires deploying batch processing code in each child node, allowing each child node to access the batch processing tasks distributed by the message middleware via the message model and generate corresponding batch processing tasks. Compared to the structure shown in FIG5 , the method shown in FIG6 using message middleware 306 to distribute batch processing tasks requires extensive modifications to the system code, which is invasive to the system. Furthermore, the monitoring information obtained by monitoring the batch processing process through the message middleware is relatively limited, resulting in relatively low data security for batch processing.
[0098] Therefore, continuing to refer to Figure 5, in an embodiment of the present application, a load balancing unit 304 and a sub-node layer 303 can be deployed using a microservice architecture, and batch processing code can be configured for the master node 302 using spring batch. Furthermore, the processing capability corresponding to the processing unit 022' and the writing capability corresponding to the writing unit 023' in the master node 302' shown in Figure 6 are configured as the processing service and writing service that can be provided by the microservice architecture, thereby obtaining the framework structure shown in Figure 5. Here, each sub-node in the sub-node layer 303 can provide the above-mentioned processing service and writing service to the execution unit 023. When the master node 302 calls the execution unit 023, the processing service and writing service on each sub-node can be called, and the load balancing unit 304 can balance the computing load between different sub-nodes, avoiding the repeated distribution of batch execution tasks to the same sub-node. On the one hand, the batch processing code can be configured only in the master node 302, avoiding invasive modifications to the system; on the other hand, the computing efficiency of the batch processing node group 300B can be effectively improved.
[0099] The specific implementation process of the node structure of the batch processing shown in FIG5 above is described in detail below with reference to FIG7 .
[0100] FIG7 shows a flow chart of a data processing method according to other embodiments of the present application.
[0101] Referring to Figure 7, the specific process steps include:
[0102] S701, the reading unit 021 reads the data to be processed.
[0103] It can be understood that the reading unit 021 can provide a reading function and can read data to be processed for batch processing. For example, in a policy renewal payment processing scenario, the policy renewal payment record can be read; in a batch reconciliation processing scenario between the insurance business system and the bank, the reconciliation record between the insurance business system and the bank can be read; in a regulatory data reporting processing scenario, the regulatory data to be reported can be read; in a reinsurance processing scenario, historical policy information and newly added reinsurance application information can be read, etc.
[0104] In some embodiments, referring to Figure 7, the master node 302 may include a reading unit 021, a conversion unit 022, and an execution unit 023. The master node 302 may execute a spring batch job step, and the job step may include the complete execution logic of a reading component (reader), a processing component (processor), and a writing component (writer).
[0105] It should be understood that in the process of using Spring Batch to process data in batches, the execution logic of specific job steps can be defined through small tasks (tasklets). For example, you can configure how the reader component (reader) reads, how the processor component (processor) processes, and how the writer component (writer) writes.
[0106] For example, small tasks can create transactions to execute batch processing tasks in batches. For example, to batch process 1,000 data items, you need to read 1,000 data items, process 1,000 data items, and write 1,000 data items. You can read, process, and write 10 data items in one transaction, commit the transaction, and then repeat the above processing on the next 10 data items in the next transaction.
[0107] Therefore, batch processing code is required to support the execution of the above-mentioned small tasks. Batch processing code can support the processing logic of batch execution. Compared with the method shown in Figure 6 where batch processing code is configured in each child node, the embodiment of the present application only needs to configure batch processing code for the master node 302, and then the load balancing unit 304 can be used to call each child node to implement batch processing and batch writing functions, which not only improves batch processing efficiency but also effectively reduces the intrusion of code modifications on the system.
[0108] In some embodiments, the processing functions of the processing component and the writing functions of the writing component in the aforementioned Spring Batch can be configured as the processing services and writing services provided by a pre-defined microservices architecture, so that each child node can provide data processing and data writing capabilities. When a child node receives a batch execution task, it can process the converted data based on its data processing capabilities to obtain processed data, and then write the processed data to the repository based on its data writing capabilities.
[0109] S702 , the reading unit 021 sends the data to be processed to the converting unit 022 .
[0110] Exemplarily, the reading unit 021 sends the read data to be processed to the conversion unit 022, so that the conversion unit 022 can perform batch data conversion based on the read data to be processed to obtain the required data after conversion. For example, the format of the insurance policy data can be converted into the required specified format.
[0111] S703, the conversion unit 022 converts the data to be processed to obtain converted data.
[0112] Exemplarily, the conversion unit 022 may convert the data to be processed according to a specific application scenario to obtain converted data.
[0113] For example, in the policy renewal payment processing scenario, the format of the read policy renewal payment related information can be converted into the required specified format; in the batch reconciliation processing scenario between the insurance business system and the bank, the read insurance business system bills and the bank's premium records can be converted into data of the same format to facilitate reconciliation; in the regulatory data reporting processing scenario, the read regulatory data to be reported can be converted into the data format required for reporting, which facilitates regulatory data reporting; in the reinsurance processing scenario, historical policy information and new reinsurance application information can be converted into data of the same format to facilitate the execution of reinsurance applications or reinsurance data updates.
[0114] S704 , the conversion unit 022 sends the converted data to the execution unit 023 .
[0115] Exemplarily, the conversion unit 022 sends the converted data to the execution unit 023 so that the execution unit 023 can generate batch execution tasks based on the converted data and distribute them to each child node for batch execution, such as writing them into a database after batch processing.
[0116] In some embodiments, the conversion unit 022 sends the processed data to the execution unit 023. This sending operation correspondingly calls the execution interface corresponding to the execution unit 023 to perform the data batch processing and data batch writing configured by the batch processing code. In this case, the processing logic of the processing interface of the original execution unit 023 and the writing logic of the writing interface can be exposed as a remote call interface in the form of a Restful API, so that when the execution unit 023 is called within the master node 302, the processing service and writing service provided by the microservice architecture can be remotely called. In this case, the method of calling the execution unit 023 corresponds to the execution method of the batch processing task. For example, the batch processing task has defined how to call the processing interface corresponding to the processing logic and the writing interface corresponding to the writing logic within the execution unit 023 in batches, and the number of records to be processed by a single call to the processing interface and the writing interface. Therefore, there is no need to configure the batch processing code in each child node. The distributed execution of the batch execution task can be achieved by using the child nodes provided by the microservice architecture, effectively improving the batch processing efficiency.
[0117] In some embodiments, the processing logic of the processing interface and the writing logic of the writing interface can be encapsulated as a remote Restful API based on the model-view-controller (MVC) framework provided by the Spring framework. For example, MVC can add annotations for exposing the remote call interface on the processing interface and the writing interface. When the processing service and / or the writing service of the execution unit 023 is started (that is, when the writing interface of the execution unit 023 is called), a Restful API can be automatically exposed. The main node 302 only needs to call the Restful API to call the processing service and / or the writing service in the microservice framework, and then call one or more subnodes in the subnode layer 303.
[0118] S705, the execution unit 023 generates a batch execution task according to the converted data.
[0119] It should be understood that the execution unit 023 may not directly process and write the converted data, but instead distribute the converted data as batch execution tasks to each child node. For example, the data may be distributed to each child node within the child node layer 303 via the load balancing unit 304, with each child node processing the converted data in batches and writing the data to the database in batches. This effectively reduces the operating load of the master node and improves batch processing efficiency.
[0120] In some embodiments, when the master node 302 detects that a processing function and / or a write function is about to be performed on the converted data, the master node 302 may send the converted data to the load balancing unit 304 without passing through the execution unit 023 .
[0121] S706 , the execution unit 023 sends the batch execution task to the load balancing unit 304 .
[0122] Exemplarily, the execution unit 023 can send batch execution tasks in batches to the load balancing unit 304, for example, referring to the batch processing code set in the master node 203, batch processing data and writing processed data, and generating batch execution tasks in batches and sending them to the load balancing unit 304.
[0123] In some embodiments, the master node 302 may generate batch execution tasks without going through the execution unit 023 and send the batch execution tasks to the load balancing unit 304 in batches.
[0124] S707 , the load balancing unit 304 calls an interface to distribute batch execution tasks to target child nodes in the child node layer 303 .
[0125] Exemplarily, the load balancing unit 304 may call an interface, such as a RESTful API, to determine a target subnode in the subnode layer 303 and distribute the batch execution task to the target subnode.
[0126] In some embodiments, the load balancing unit 304 and the sub-node layer 303 can be provided by a microservice architecture. The load balancing unit 304 can be a load balancer in the microservice architecture, which can be used to distribute access to various computing resources, such as multiple sub-nodes (i.e., individual servers). The load balancer can prevent batch execution tasks from being sent to servers that are not functioning properly, prevent resource overloading, and eliminate single points of failure.
[0127] In some embodiments, the protocol for calling the interface may be a standard RESTful protocol or a hypertext transfer protocol (HTTP).
[0128] S708 , the target sub-node in the sub-node layer 303 performs data processing operations according to the batch execution task, and writes the processed data into the database in batches.
[0129] For example, the target subnodes within subnode layer 303 can perform data processing operations on the converted data based on the batch execution tasks obtained in batches, and write the processed data in batches to the database to complete the update of the data in the database. In some embodiments, each subnode can be a server. By having multiple servers execute batch execution tasks separately, the efficiency of batch processing and batch writing of data can be effectively improved.
[0130] In some embodiments, the target sub-node may process the converted data according to a specific application scenario to obtain processed data.
[0131] For example, in the policy renewal payment processing scenario, the insured can be renewed based on the read policy renewal payment related information to obtain a renewal record; in the batch reconciliation processing scenario between the insurance business system and the bank, the reconciliation can be performed based on the read insurance business system bills and the bank's premium records, and reconciliation information can be generated; in the regulatory data reporting processing scenario, the read regulatory data to be reported can be reported, and the information reporting record can be obtained; in the reinsurance processing scenario, reinsurance processing can be performed based on historical policy information and newly added reinsurance application information to obtain a reinsurance policy.
[0132] Furthermore, the target subnode can write the processed data into the database. For example, in the policy renewal payment processing scenario, the target subnode in subnode layer 303 can write the renewal record into the database; in the scenario of batch reconciliation between the insurance business system and the bank, the target subnode in subnode layer 303 can write the reconciliation statement information into the database; in the regulatory data reporting processing scenario, the target subnode in subnode layer 303 can write the information reporting record and updated regulatory data into the database; in the reinsurance processing scenario, the target subnode in subnode layer 303 can write the reinsurance policy into the database.
[0133] It can be understood that through the specific implementation of the above steps S701 to S708, the converted data is distributed to each child node through the load balancing unit 304, thereby preventing batch execution tasks from being sent to servers that cannot operate normally, preventing resource overload, and eliminating single point failures.
[0134] In some embodiments, the multiple sub-nodes within the above-mentioned sub-node layer 303 can be composed of multiple micro-services in a micro-service architecture, and each micro-service can be deployed on an independent server. Furthermore, when the master node performs a batch processing task of data, the standard micro-service RESTful API interface in the micro-service architecture can be called to distribute the converted data to each micro-service, and each micro-service writes the processed data into the database. On the one hand, the operating load of the master node can be reduced, and on the other hand, the standard RESTful API interface provided by the micro-service architecture can be used to implement monitoring and management of batch processing. Compared with the small amount of monitoring information provided by the message middleware shown in Figure 6, more complete monitoring and management can be achieved through the standard RESTful API interface.
[0135] FIG8 shows a block diagram of a server 200 according to an embodiment of the present application. In some embodiments, the server 200 may include one or more processors 804, a system control logic 808 connected to at least one of the processors 804, a system memory 812 connected to the system control logic 808, a non-volatile memory (NVM) 816 connected to the system control logic 808, and a network interface 820 connected to the system control logic 808.
[0136] In some embodiments, the processor 804 may include one or more single-core or multi-core processors. In some embodiments, the processor 804 may include any combination of general-purpose processors and specialized processors (e.g., graphics processors, application processors, baseband processors, etc.). In embodiments where the server 200 employs an enhanced base station (evolved node b, eNB) 101 or a radio access network (RAN) controller 102, the processor 804 may be configured to execute various embodiments.
[0137] In some embodiments, system control logic 808 may include any suitable interface controller to provide any suitable interface to at least one of processors 804 and / or any suitable device or component in communication with system control logic 808 .
[0138] In some embodiments, the system control logic 808 may include one or more memory controllers to provide an interface to the system memory 812. The system memory 812 may be used to load and store data and / or instructions. In some embodiments, the memory 812 of the server 200 may include any suitable volatile memory, such as a suitable dynamic random access memory (DRAM).
[0139] NVM / memory 816 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, NVM / memory 816 may include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device, such as at least one of a hard disk drive (HDD), a compact disc (CD) drive, and a digital versatile disc (DVD) drive.
[0140] NVM / storage 816 may include a portion of storage resources on the device where server 200 is installed, or it may be accessible to the device but not necessarily part of the device. For example, NVM / storage 816 may be accessed over a network via network interface 820.
[0141] In particular, system memory 812 and NVM / storage 816 may respectively include a temporary copy and a permanent copy of instructions 824. Instructions 824 may include instructions that, when executed by at least one of processors 804, cause server 200 to implement the aforementioned data processing method. In some embodiments, instructions 824, hardware, firmware, and / or software components thereof may additionally or alternatively be located in system control logic 808, network interface 820, and / or processor 804.
[0142] The network interface 820 may include a transceiver for providing a radio interface for the server 200, thereby communicating with any other suitable devices (such as a front-end module, an antenna, etc.) via one or more networks. In some embodiments, the network interface 820 may be integrated with other components of the server 200. For example, the network interface 820 may be integrated with at least one of the processor 804, the system memory 812, the NVM / storage 816, and a firmware device (not shown) having instructions. When at least one of the processors 804 executes the instructions, the server 200 implements the above-described data processing method.
[0143] The network interface 820 may further include any suitable hardware and / or firmware to provide a multiple-input multiple-output radio interface. For example, the network interface 820 may be a network adapter, a wireless network adapter, a telephone modem, and / or a wireless modem.
[0144] In one embodiment, at least one of the processors 804 may be packaged together with logic for one or more controllers of the system control logic 808 to form a system-in-package (SiP). In one embodiment, at least one of the processors 804 may be integrated on the same die with logic for one or more controllers of the system control logic 808 to form a system-on-chip (SoC).
[0145] Server 200 may further include input / output (I / O) devices 832. I / O devices 832 may include a user interface to enable a user to interact with server 200, and peripheral component interfaces to enable peripheral components to interact with server 200. In some embodiments, server 200 may also include sensors for determining at least one of environmental conditions and location information related to server 200.
[0146] In some embodiments, the user interface may include, but is not limited to, a display (e.g., an LCD display, a touch screen display, etc.), a speaker, a microphone, one or more cameras (e.g., a still image camera and / or a video camera), a flashlight (e.g., an LED flash), and a keyboard.
[0147] In some embodiments, the peripheral component interface may include, but is not limited to, a non-volatile memory port, an audio jack, and a power interface.
[0148] In some embodiments, the sensors may include, but are not limited to, a gyroscope sensor, an accelerometer, a proximity sensor, an ambient light sensor, and a positioning unit. The positioning unit may also be part of or interact with the network interface 820 to communicate with components of a positioning network (e.g., Global Positioning System (GPS) satellites).
[0149] According to the method provided in the embodiments of the present application, the present application also provides a computer program product, which includes: computer program code, which, when executed on a computer, enables the computer to implement the steps performed by the server 200 in any one of the above embodiments.
[0150] According to the method provided in the embodiments of the present application, the present application also provides a computer-readable medium, which stores program code. When the program code runs on a computer, the computer implements the steps performed by the server 200 in any of the above embodiments.
[0151] The various embodiments disclosed in this application can be implemented in hardware, software, firmware, or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0152] Program code can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.
[0153] Program code can be implemented with a high-level programming language or an object-oriented programming language to communicate with the processing system. Where necessary, program code can also be implemented in assembly language or machine language. In fact, the mechanism described in this application is not limited to the scope of any particular programming language. In either case, the language can be a compiled language or an interpreted language.
[0154] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed over a network or through other computer-readable media. Therefore, a machine-readable medium may include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including but not limited to floppy disks, optical disks, optical discs, read-only memories (CD-ROMs), magneto-optical disks, read-only memories (ROMs), random access memories (RAMs), erasable programmable read-only memories (EPROMs), electrically erasable programmable read-only memories (EEPROMs), magnetic or optical cards, flash memory, or a tangible machine-readable memory for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in electrical, optical, acoustic, or other forms of propagation signals. Therefore, a machine-readable medium includes any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).
[0155] In the accompanying drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or order may not be required. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. In addition, the inclusion of a structural or method feature in a particular figure does not imply that such feature is required in all embodiments, and in some embodiments, such features may not be included or may be combined with other features.
[0156] It should be noted that the units / modules mentioned in the various device embodiments of the present application are all logical units / modules. Physically, a logical unit / module can be a physical unit / module, or a part of a physical unit / module, or can be implemented as a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important. The combination of functions implemented by these logical units / modules is the key to solving the technical problems raised by this application. In addition, in order to highlight the innovative part of this application, the above-mentioned device embodiments of this application do not introduce units / modules that are not closely related to solving the technical problems raised by this application. This does not mean that other units / modules do not exist in the above-mentioned device embodiments.
[0157] It should be noted that in the examples and description of this patent, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "including a" does not exclude the presence of other identical elements in the process, method, article or device that includes the element.
[0158] While the present application has been shown and described with reference to certain embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the scope of the present application.
Claims
1. A data processing method, applied to a data processing system, characterized in that: The data processing system includes a main node and at least one sub-node, wherein the main node and the at least one sub-node are deployed in different electronic devices, the main node calls the at least one sub-node through an interface, and the method includes: Reading the data to be processed through the master node, and generating batch processing tasks according to the data to be processed; Distributing the batch processing tasks generated by the master node to the at least one child node through the interface; The batch processing task is executed by the at least one child node to obtain processed data, and the processed data is written into a database, and the at least one child node cannot generate the batch processing task according to the processed data or the data to be processed.
2. The method according to claim 1, characterized in that The batch processing task includes a batch execution task, and, The step of generating a batch processing task according to the data to be processed further includes: A batch execution task is generated according to the processed data, wherein the batch execution task includes batch processing and batch writing.
3. The method according to claim 2, characterized in that The distributing the batch processing tasks generated by the master node to the at least one child node through the interface includes: The batch execution task is distributed to the at least one child node through the interface.
4. The method according to claim 3, characterized in that The system further includes a load balancing unit, and distributing the batch execution task to the at least one child node includes: The batch execution tasks are distributed to the at least one child node through the load balancing unit, wherein the load balancing unit is used to adjust the operating load of the at least one child node.
5. The method according to claim 2, characterized in that: The master node further includes an execution unit, and the method further includes: An execution capability is configured for the at least one child node so that the at least one child node can execute the batch execution task, wherein the execution capability includes a data processing capability and a writing capability.
6. The method according to claim 1, characterized in that The distributing the batch processing tasks generated by the master node to the at least one child node through the interface includes: The main node is used to call the interface based on the RESTful protocol or the HTTP protocol to call the at least one child node.
7. The method according to claim 4, characterized in that The load balancing unit, the interface and the at least one sub-node are all deployed through a preset microservice architecture.
8. The method according to claim 7, characterized in that The method further comprises: The execution process of the batch processing task is monitored through the preset microservice architecture, and the monitoring information is recorded.
9. A data processing system, characterized in that: The data processing system includes a main node and at least one sub-node, wherein the main node and the at least one sub-node are deployed in different electronic devices, the main node calls the at least one sub-node through an interface, and, The master node is used to read the data to be processed and generate batch processing tasks according to the data to be processed; The interface is used to distribute the batch processing tasks generated by the master node to the at least one child node; The at least one child node is used to execute the batch processing task, obtain the processed data, and The processed data is written into the database, and the at least one child node cannot generate the batch processing task according to the processed data or the data to be processed.
10. The system according to claim 9, characterized in that The data processing system further comprises a load balancing unit, wherein the load balancing unit is used to adjust the operating load of the at least one sub-node.
11. The system according to claim 9, characterized in that The master node further includes an execution unit, and the execution unit is used to configure execution capability for the at least one child node so that the at least one child node can execute the batch processing task, wherein the execution capability includes data processing capability and writing capability.
12. The system according to claim 10, characterized in that The load balancing unit, the interface and the at least one sub-node are all deployed in a preset microservice architecture.
13. The system according to claim 12, characterized in that The preset microservice architecture is also used to monitor the execution process of the batch processing task through the preset microservice architecture and record monitoring information.
14. A computer-readable storage medium, characterized in that: The storage medium stores instructions, which, when executed on a computer, cause the computer to execute the data processing method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Distributed real-time calculation system and data processing method thereof
CN103701906A
Batch processing cluster system and method
CN105589756A
Massive data processing method and system
CN108446352A
Task processing system and task processing method
CN116225662A
Mechanism to Enable and Ensure Failover Integrity and High Availability of Batch Processing
US20090265710A1