A data processing method and apparatus
By creating multiple threads in the business nodes of the financial system, processing unlabeled shard information to improve data processing performance, the upper limit of the existing system's concurrent processing capability when processing large amounts of data is solved, and more efficient data processing and horizontal scaling is achieved.
Patent Information
- Application Number
- CN202010607664.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-06-29
AI Technical Summary
In existing financial systems deal with bulk transactions of large amounts of customer assets and liabilities, there is an upper limit of concurrent processing capabilities, resulting in frequent CPU context switching, increasing time-consuming and degrading performance.
By creating multiple threads in the business node, the thread looks for unmarked shard information from the shard information, locks and marks it, obtains the corresponding shard data from the database for processing, and stores the processed data into the database.
By increasing the number of business nodes and threads, the concurrency capability and performance of data processing are improved, the time-consuming process of processing large amounts of data is shortened, and the horizontal scalability of processing is improved.
Smart Images

Figure CN111752961B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing in financial technology (Fintech), and particularly to a data processing method and apparatus. Background Art
[0002] With the continuous development of financial technology, especially Internet technology finance, more and more technologies (such as distributed, blockchain, artificial intelligence, etc.) are applied in the financial field. However, the financial industry also poses higher requirements for technologies, such as for the data processing process.
[0003] Banks and financial systems perform batch processing such as clearing, interest calculation, and deduction of a large amount of customer assets and liabilities every day. Currently, most banks and financial systems still use single-database storage for data storage. There are complex pre-dependency relationships between batch transactions, and most batches are only allowed to be executed once. Therefore, batch transactions are often executed on a single node. For some time-consuming transactions, multi-threaded concurrent execution is generally used to improve the concurrent execution efficiency.
[0004] In the prior art, the concurrent processing ability of batch transactions is increased by increasing the number of threads. However, this expansion ability always has an upper limit. When the number of threads increases beyond a certain number, it will instead increase the time consumption due to frequent CPU context switching, thus reducing the performance. Summary of the Invention
[0005] This application provides a data processing method and apparatus to improve data processing performance and shorten the time consumption for processing a large amount of data.
[0006] A data processing method provided by an embodiment of the present invention is applied to a business system. The business system includes a database and at least two business nodes. The method includes:
[0007] A business node creates N threads, where N is an integer greater than 1;
[0008] For each thread of the business node, the thread searches for unmarked shard information from all shard information. The shard information is data information corresponding to shard data obtained by sharding business data;
[0009] The thread locks an unmarked shard information and marks it;
[0010] The thread obtains corresponding shard data from the database according to the shard information;
[0011] The thread processes the shard data and stores the processed shard data into the database.
[0012] In an alternative embodiment, the thread locks an unmarked shard information and marks it, including:
[0013] The thread locks the unmarked shard information using a pessimistic lock;
[0014] The thread marks the shard information using the thread identifier.
[0015] In an alternative embodiment, after the thread marks the shard information using the thread identifier, it further includes:
[0016] The thread determines whether the thread identifier corresponding to the shard information is the thread identifier of the thread;
[0017] If so, the thread executes the step of obtaining the corresponding shard data to be processed from the database according to the shard information;
[0018] If not, the thread executes the step of finding unmarked shard information from all shard information.
[0019] In an alternative embodiment, the shard information includes the processing status of the corresponding shard data; before the thread locks an unmarked shard information and marks it, the processing status of the shard information is unprocessed;
[0020] After the thread locks an unmarked shard information and marks it, it further includes:
[0021] The thread changes the processing status of the shard information to be processed to processing;
[0022] After the thread processes the shard data to be processed and stores the processed shard data into the database, it further includes:
[0023] The thread changes the processing status of the shard information to be processed to processed.
[0024] In an alternative embodiment, after the thread finds unmarked shard information from all shard information, it further includes:
[0025] If all shard information has been marked, the service node cancels the thread.
[0026] In an alternative embodiment, after the thread processes the shard data, it further includes:
[0027] The service node sends a heartbeat message to the database at a set frequency;
[0028] If the database does not receive a heartbeat message from the first service node within a set time, query the sharding information corresponding to the first service node, where the first service node is any service node in the service system;
[0029] If the process identifier of the sharding information does not correspond to the first service node, the database determines the processing status of the sharding information of the sharded data;
[0030] If the processing status of the sharding information is processing failure, send an alarm indication;
[0031] If the processing status of the sharding information is in processing, change the processing status of the sharding information to unprocessed.
[0032] An embodiment of the present invention further provides a data processing device, which is applied to a service system. The service system includes a database and at least two service nodes. The device includes:
[0033] A creation unit, configured to create N threads, where N is an integer greater than 1;
[0034] A search unit, configured to search for unmarked sharding information from all sharding information for each thread of the service node, where the sharding information is data information corresponding to sharded data obtained by sharding service data;
[0035] A locking unit, configured to lock an unmarked sharding information and mark it;
[0036] An obtaining unit, configured to obtain corresponding sharded data from the database according to the sharding information;
[0037] A processing unit, configured to process the sharded data and store the processed sharded data in the database.
[0038] In an alternative embodiment, the locking unit is specifically configured to:
[0039] Lock the unmarked sharding information by using a pessimistic lock;
[0040] Mark the sharding information by using a thread identifier.
[0041] In an alternative embodiment, the locking unit is further configured to:
[0042] Determine whether the thread identifier corresponding to the sharding information is the thread identifier of the thread;
[0043] If so, execute the step of obtaining corresponding sharded data to be processed from the database according to the sharding information;
[0044] Otherwise, perform the step of finding unmarked shard information from all shard information.
[0045] In an alternative embodiment, the shard information includes the processing status of the corresponding shard data; before the thread locks an unmarked shard information and marks it, the processing status of the shard information is unprocessed;
[0046] The locking unit is further configured to:
[0047] After the thread locks an unmarked shard information and marks it, change the processing status of the to-be-processed shard information to processing;
[0048] After the thread processes the to-be-processed shard data and stores the processed processed shard data into the database, change the processing status of the to-be-processed shard information to processed.
[0049] In an alternative embodiment, the creating unit is further configured to:
[0050] If all shard information has been marked, cancel the thread.
[0051] An embodiment of the present invention further provides an electronic device, including:
[0052] At least one processor; and,
[0053] A memory communicatively connected to the at least one processor; wherein,
[0054] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.
[0055] An embodiment of the present invention further provides a non-transitory computer-readable storage medium, and the non-transitory computer-readable storage medium stores computer instructions for causing the computer to execute the method as described above.
[0056] An embodiment of the present invention provides a service system, which includes a database and at least two service nodes. For each service node, the service node creates N threads, where N is an integer greater than 1. For each thread, the thread searches for unmarked shard information from all shard information, where the shard information is the data information corresponding to the shard data obtained by sharding the service data. The thread locks an unmarked shard information and marks it, and the thread obtains the corresponding shard data from the database according to the shard information. Then, the thread processes the shard data and stores the processed shard data into the database. By adding service nodes in the embodiment of the present invention, multiple service nodes jointly execute the service data processing task, and each service node creates multiple threads, thereby increasing the upper limit of the number of threads, and further making full use of the computing resources of multiple nodes, improving the processing performance, and also improving the horizontal scalability of the processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0058] Figure 1 It is a schematic structural diagram of a possible system architecture provided by an embodiment of the present invention;
[0059] Figure 2 It is a schematic flow diagram of a data processing method provided by an embodiment of the present invention;
[0060] Figure 3 It is a schematic structural diagram of a data processing device provided by an embodiment of the present invention;
[0061] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0063] For ease of understanding, the following defines and explains the nouns that may be involved in the embodiments of the present invention.
[0064] Loan Core System: A system program that provides services such as credit limit control, loan disbursement, repayment, accounting, and ledger posting for loan products (such as Micro Business Loan).
[0065] Promissory Note: A data storage for a certain creditor - debtor relationship between a bank and an enterprise (individual), mainly storing information such as the debt amount, debtor, debt execution interest rate, term, repayment method, etc.
[0066] Interest Accrual: Calculate the interest on a certain loan (or promissory note) daily according to the established execution interest rate and the specified base (usually the loan balance), and record it in the interest account.
[0067] Multi - instance: An application process that executes business logic is an instance. In actual enterprise - level applications, to ensure high - concurrency and high - availability services, multiple instances are usually deployed in multiple computer rooms of different machines.
[0068] Data Sharding Pool: For a large amount of data, it is sliced according to a certain algorithm in a certain dimension, sliced into several shards to form a sharding pool, so that the application program can schedule and process the data in units of shards. In this article, for a large number of promissory notes in the business system, modulo 1000 is taken according to the customer number to form 1000 shards, and the shard number is the modulo value.
[0069] JVM: An abbreviation for Java Virtual Machine. The JVM is a specification for computing devices. It is a fictional computer that is implemented by emulating various computer functions on an actual computer. After introducing the Java language virtual machine, the Java language does not need to be recompiled when running on different operating system platforms.
[0070] As Figure 1 shown, a system architecture applicable to the embodiments of the present invention includes at least two service nodes 101 and a database 102.
[0071] Among them, the service node 101 can be an electronic device with wireless communication functions such as a mobile phone, a tablet computer, or a dedicated handheld device, or can be a device connected to the Internet through a wired access method such as a personal computer (abbreviated as PC), a laptop computer, a server, etc. The service node 101 can also be a network device such as a computer, which can be an independent device or a server cluster formed by multiple servers. Preferably, the service node 101 can adopt cloud computing technology for information processing.
[0072] The database 102 can be various types of databases. Preferably, it is a database cluster and can adopt cloud computing technology for information processing.
[0073] The service node 101 can communicate with the database 102 through the INTERNET network, or can also communicate with the database 102 through mobile communication systems such as the Global System for Mobile Communications (GSM) and the Long Term Evolution (LTE) system.
[0074] Based on the above architecture, an embodiment of the present invention provides a data processing method. As Figure 2 shown, the data processing method provided by the embodiment of the present invention is applied to a service system. The service system includes a database and at least two service nodes. The method includes the following steps:
[0075] Step 201: The service node creates N threads, where N is an integer greater than 1.
[0076] Step 202: For each thread of the service node, the thread searches for unmarked shard information from all shard information. The shard information is the data information corresponding to the shard data obtained by sharding service data.
[0077] Step 203: The thread locks an unmarked shard information and marks it.
[0078] Step 204: The thread obtains the corresponding shard data from the database according to the shard information.
[0079] Step 205: The thread processes the shard data and stores the processed shard data into the database.
[0080] An embodiment of the present invention provides a service system, which includes a database and at least two service nodes. For each service node, the service node creates N threads, where N is an integer greater than 1. For each thread, the thread searches for unmarked shard information from all shard information, where the shard information is the data information corresponding to the shard data obtained by sharding service data. The thread locks an unmarked shard information and marks it. The thread obtains the corresponding shard data from the database according to the shard information. Then, the thread processes the shard data and stores the processed shard data into the database. By adding service nodes in the embodiment of the present invention, multiple service nodes jointly execute the service data processing task, and each service node creates multiple threads, thereby increasing the upper limit of the number of threads, further making full use of the computing resources of multiple nodes, improving the processing performance, and also improving the horizontal scalability of the processing.
[0081] In the embodiments of the present invention, in order to facilitate the storage of shard information, the shard information is stored in the form of a table through a shard pool. For a large amount of service data of a service, it is sliced according to a certain algorithm in a certain dimension into several shards to form a shard pool, so that the application program can schedule and process it in units of shards.
[0082] Suppose there are M service nodes deployed in a certain service system, and there is one instance in each service node. If one instance creates N threads, then the total number of threads that can process batch tasks is M * N. The service data is sliced into shard data with a finer granularity of quantity, usually 10 - 50 times the total number of threads, and generally an integer is taken, such as 1000.
[0083] The reason for the multi-instance batch scheduling algorithm to slice the batch-processed data into finer granularity is to shield the problems caused by some instances being busy with resources such as CPU, or a certain shard having too much data resulting in too long processing time for that shard. Only when the shards are sliced into a finer granularity can the processing time of each instance for shard data be more balanced.
[0084] For the selection and implementation of the shard slicing algorithm, the embodiments of the present invention include but are not limited to modulo algorithm, number segment algorithm, custom slicing function algorithm, etc. Preferably, in the embodiments of the present invention, the modulo algorithm is taken as an example to illustrate the setting of the shard pool. For example, in the evening interest accrual batch of a certain loan core system, when processing the existing loan receipts of 1 million customers, it is expected to enable 5 instances, with 10 threads in each instance, so there are a total of 50 threads, and the service data is divided into 1000 shard data.
[0085] Further, the thread locks an unmarked shard information and marks it, including:
[0086] The thread locks the unmarked shard information using a pessimistic lock;
[0087] The thread marks the shard information using the thread identifier.
[0088] In the specific implementation process, taking a relational database as an example, the shard pool in the embodiments of the present invention uses the date, shard pool ID, and shard ID as the composite primary key to identify a certain shard data of a batch task. The shard ID is an integer from 0 to 999, the default shard execution status is 0 (unprocessed), and the preemptive instance IP field is initially empty. The specific manifestation form of the shard pool can be a table structure, as shown in Table 1.
[0089] Table 1
[0090]
[0091] Specifically, for a certain task data, the sharding pool table batch generates 1000 shards. The shard IDs are integers from 0 to 999, and the initial state is 0 (unprocessed). The preemptive instance IP field is initially empty, and the data is stored as shown in Table 2. Among them, the preemptive thread IP is the thread identifier.
[0092] Table 2
[0093] Date Fragment Pool ID Fragment ID Status Description Preemption Thread IP 2019 / 10 / 08 loan_end_batch 0 2 Processing Successful 10.168.8.1 2019 / 10 / 08 loan_end_batch 1 2 Processing Successful 10.168.8.2 2019 / 10 / 08 loan_end_batch 1 1 Processing 10.168.8.3 2019 / 10 / 08 loan_end_batch … 0 Not Processed 2019 / 10 / 08 loan_end_batch n-2 0 Not Processed 2019 / 10 / 08 loan_end_batch n-1 0 Not Processed
[0094] Furthermore, the shard information includes the processing status of the corresponding shard data; before the thread locks an unmarked shard information and marks it, the processing status of the shard information is unprocessed;
[0095] After the thread locks an unmarked shard information and marks it, it further includes:
[0096] The thread changes the processing status of the to-be-processed shard information to processing;
[0097] After the thread processes the to-be-processed shard data and stores the processed shard data into the database, it further includes:
[0098] The thread changes the processing status of the to-be-processed shard information to processed.
[0099] In the specific implementation process, after a task marked as multi-instance is scheduled, it first undergoes instance executability verification. Only for an instance that does not exceed the task's executable instance quantity limit and has no task being executed can it process the task.
[0100] In the embodiment of the present invention, the instance process inserts a batch task execution record R of the instance into the relational database, avoiding the instance from being rescheduled and executed again due to a new round of timed tasks or manually triggered tasks. The key fields of the instance execution record include: task ID, execution date, execution instance IP, execution status, execution time, etc.
[0101] In the embodiment of the present invention, multiple service nodes jointly process task data, and each service node corresponds to an instance. For any instance, N threads are created, and all threads jointly preempt and process the shard data.
[0102] The specific process for any thread is as follows:
[0103] The thread lock preempts shard information with an unprocessed state and an empty occupied instance IP field.
[0104] If no processable shard information is preempted, the thread is cancelled.
[0105] If the shard information that can be processed is preempted, update the status of the shard to being processed, set the occupied instance IP field to the current instance IP @ thread number, and process the shard data. After the shard data processing is completed, update the processing status of the shard information.
[0106] Further, after the thread finds the unmarked shard information from all the shard information, it also includes:
[0107] If all the shard information has been marked, the service node cancels the thread.
[0108] In the specific implementation process, the above process is looped until there is no shard data that can be processed, and then the thread is cancelled.
[0109] After all the N threads created above are cancelled, according to whether there are processing errors in the N threads, if there is no error, update the status of instance task R to success; if there is an error, update the status of instance task R to failure, and store the error information thrown by the thread.
[0110] Further, after the thread marks the shard information using the thread identifier, it also includes:
[0111] The thread determines whether the thread identifier corresponding to the shard information is the thread identifier of the thread;
[0112] If so, the thread executes the step of obtaining the corresponding shard data to be processed from the database according to the shard information;
[0113] If not, the thread executes the step of finding the unmarked shard information from all the shard information.
[0114] In the specific implementation process, thread n of instance P queries the number K of shards with the status of unprocessed in the shard pool. If there is no unprocessed shard, the thread is cancelled. If there is an unprocessed shard, use a pessimistic lock to query and lock 1 unprocessed shard A, and at the same time occupy shard A, update the status of the shard to being processed, and the occupied instance information field is instance P @ thread n. At the same time, to prevent multiple threads from concurrently preempting the shard during shard preemption, resulting in the shard A returned by the query being occupied by other threads and causing the shard A to be processed repeatedly, the present invention designs a shard conflict detection mechanism to ensure that the shard is uniquely occupied and processed.
[0115] The shard preemption conflict detection mechanism is as follows: After thread n of instance P preempts and updates shard A, query the occupied instance information of shard A again. If the occupied instance information is the same as that of thread n of the current instance P, it means that shard A is correctly preempted by thread n of instance P. Otherwise, it means that during the process of thread n of instance P querying and occupying shard A, other instance threads preempted first, and then thread n of instance P returns null in this preemption and occupation of the shard.
[0116] After thread n of instance P preempts shard A, it processes the data records in the service table with a modulus value of A. The modulo operation logic is: MOD(modulo field, 1000) = shard A. After processing the shard data, it loops to preempt and process the remaining shards to be processed until there are no more shards to be processed, and then the thread ends.
[0117] Further, after the thread processes the shard data, it further includes:
[0118] The service node sends a heartbeat message to the database at a set frequency;
[0119] If the database does not receive the heartbeat message of the first service node within a set time, it queries the shard information corresponding to the first service node, where the first service node is any service node in the service system;
[0120] If the process identifier of the shard information does not correspond to the first service node, the database determines the processing status of the shard information of the shard data;
[0121] If the processing status of the shard information is processing failure, an alarm indication is sent;
[0122] If the processing status of the shard information is in processing, the processing status of the shard information is changed to unprocessed.
[0123] In the specific implementation process, when n threads of instance P preempt shards and are processing, instance P may cause the processing of shard data to fail due to downtime or program logic errors. The shard status in the shard pool may be in processing or processing failure. For example, when the machine fails and causes downtime, the shard status is not updated to the final state (success or failure) in time, resulting in the status being in processing; another example is when the application program logic is abnormal and the shard status is processing failure. Since some shards are not processed successfully, the overall task is not executed completely.
[0124] To confirm whether the shard data in instance P is processed successfully, a liveness detection mechanism is set for instance P. The liveness detection mechanism is: each instance reports a heartbeat to the cluster center (such as the database) at regular intervals (such as 1 minute). If the heartbeat is not reported within the timeout period, the cluster center actively or passively queries the process ID information of instance P. If the process ID of the instance P that preempts the shard does not exist or is inconsistent with the process ID of this instance P, it is determined that instance P is inactive.
[0125] In the inactive instance P detected by the liveness detection, there are shards with a status of in processing and shards with a status of processing failure. For the shards with a status of in processing, their status can be changed to unprocessed, and other instances can automatically preempt and process this shard to complete the continuation of the shard task data.
[0126] For the shards with a processing failure status, an alarm indication can be sent. The shard scheduling algorithm in the embodiments of the present invention can also allocate the shards with a processing failure status to other instances for scheduling and execution. For the shards with a processing failure status, after all other shards are processed, batch tasks can be executed. To understand the present invention more clearly, the implementation effects of the above process are described below with specific embodiments.
[0127] It is implemented in the JAVA language in a certain financial business system, and a certain batch task is stress-tested. The stress test resource parameters are as follows:
[0128] Application instances: 4;
[0129] CPU: 6 cores;
[0130] Memory: 15G;
[0131] JVM memory: Xmx = 2048M;
[0132] Data volume processed by batch A: 400,000;
[0133] Number of shards per single task in the shard pool: 1000.
[0134] Judging from the stress test data, with the same total number of threads, when adding 1 instance number, the number of threads decreases inversely, the batch execution time is shortened by nearly 1 time, and the CPU usage rate and load will also be reduced to a certain extent. The stress test data is shown in Table 3. At the same time, in the execution of batch tasks, horizontal expansion of the application layer is achieved.
[0135] Table 3
[0136] Number of Instances Number of Threads per Instance Batch Duration CPU Usage Percentage Load 1 20 56 87% 11 2 10 30 77% 5 4 10 14 72% 2 2 20 31 88% 11 2 50 34 90% 17 2 100 35 95% 20
[0137] In addition, in the scenario of multiple instances preempting shards for processing, when a certain instance crashes during batch execution, data failures can be minimized, and the backend management console can be used to continue processing the data of the faulty shards, and the data of other shards will not be affected at all.
[0138] The embodiments of the present invention also provide a data processing device, which is applied to a business system. The business system includes a database and at least 2 business nodes. The device is as Figure 3 shown and includes:
[0139] A creation unit 301, configured to create N threads, where N is an integer greater than 1;
[0140] A search unit 302, configured to search for unmarked shard information from all shard information for each thread of the business node, where the shard information is the data information corresponding to the shard data obtained by sharding business data;
[0141] A locking unit 303, configured to lock an unmarked shard information and mark it;
[0142] An obtaining unit 304, configured to obtain corresponding shard data from the database according to the shard information;
[0143] A processing unit 305, configured to process the shard data and store the processed shard data into the database.
[0144] Further, the locking unit 303 is specifically configured to:
[0145] Lock the unmarked shard information by using a pessimistic lock;
[0146] Mark the shard information by using a thread identifier.
[0147] Further, the locking unit 303 is further configured to:
[0148] Determine whether the thread identifier corresponding to the shard information is the thread identifier of the thread;
[0149] If so, execute the step of obtaining corresponding shard data to be processed from the database according to the shard information;
[0150] If not, execute the step of searching for unmarked shard information from all shard information.
[0151] Further, the shard information includes a processing status of corresponding shard data; before the thread locks an unmarked shard information and marks it, the processing status of the shard information is unprocessed;
[0152] The locking unit 303 is further configured to:
[0153] After the thread locks an unmarked shard information and marks it, change the processing status of the shard information to be processed;
[0154] After the thread processes the shard data to be processed and stores the processed shard data into the database, change the processing status of the shard information to be processed to processed.
[0155] Further, the creating unit 301 is further configured to:
[0156] If all shard information has been marked, log off the thread.
[0157] Based on the same principle, the present invention further provides an electronic device, as Figure 4 shown, including:
[0158] It includes a processor 401, a memory 402, a transceiver 403, and a bus interface 404. The processor 401, the memory 402, and the transceiver 403 are connected through the bus interface 404.
[0159] The processor 401 is used to read the program in the memory 402 and execute the following method:
[0160] Create N threads, where N is an integer greater than 1;
[0161] For each thread of the service node, the thread searches for unmarked shard information from all shard information, and the shard information is the data information corresponding to the shard data obtained by sharding service data;
[0162] The thread locks an unmarked shard information and marks it;
[0163] The thread obtains the corresponding shard data from the database according to the shard information;
[0164] The thread processes the shard data and stores the processed shard data into the database.
[0165] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0166] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0167] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, causing a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process, such that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one flow Figure 1 one or more flows and / or blocks Figure 1 or steps for implementing the functions specified in a block or blocks.
[0168] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to cover the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0169] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A data processing method, characterized in that, applied to a business system, the business system includes a database and at least two business nodes, and the method includes: The business node creates N threads, where N is an integer greater than 1; For each thread of the business node, the thread searches for unmarked shard information from all shard information, and the shard information is the data information corresponding to the shard data obtained by sharding business data; The thread locks an unmarked shard information and marks it; The thread obtains the corresponding shard data from the database according to the shard information; The thread processes the shard data and stores the processed shard data into the database; Among them, the thread locks an unmarked shard information and marks it, including: The thread locks the unmarked shard information using a pessimistic lock; The thread marks the shard information using a thread identifier; After the thread marks the shard information using a thread identifier, it further includes: The thread determines whether the thread identifier corresponding to the shard information is the thread identifier of the thread; If so, the thread executes the step of obtaining the corresponding shard data to be processed from the database according to the shard information; If not, the thread executes the step of searching for unmarked shard information from all shard information.
2. The method according to claim 1, characterized in that, The shard information includes the processing status of the corresponding shard data; before the thread locks an unmarked shard information and marks it, the processing status of the shard information is unprocessed; After the thread locks an unmarked shard information and marks it, it further includes: The thread changes the processing status of the shard information to be processed to processing; After the thread processes the shard data to be processed and stores the processed shard data into the database, it further includes: The thread changes the processing status of the shard information to be processed to processed.
3. The method according to claim 1 or 2, characterized in that, After the thread searches for unmarked shard information from all shard information, it further includes: If all shard information has been marked, the business node cancels the thread.
4. The method according to claim 1, characterized in that, After the thread processes the shard data, it further includes: The business node sends a heartbeat message to the database at a set frequency; If the database does not receive the heartbeat message of the first business node within a set time, query the shard information corresponding to the first business node, and the first business node is any business node in the business system; If the process identifier of the shard information does not correspond to the first business node, the database determines the processing status of the shard information of the shard data; If the processing status of the shard information is processing failure, an alarm indication is sent; If the processing status of the shard information is processing, the processing status of the shard information is changed to unprocessed.
5. A data processing device, characterized in that, Applied to a business system, the business system includes a database and at least two business nodes, and the apparatus includes: A creation unit, configured to create N threads, where N is an integer greater than 1; A lookup unit, configured to, for each thread of the business nodes, look up unmarked shard information from all shard information, where the shard information is data information corresponding to sharded data obtained by sharding business data; A locking unit, configured to lock an unmarked shard information and mark it; An obtaining unit, configured to obtain corresponding sharded data from the database according to the shard information; A processing unit, configured to process the sharded data and store the processed sharded data into the database; The locking unit is specifically configured to: Lock the unmarked shard information by using a pessimistic lock; Mark the shard information by using a thread identifier; The processing unit is further configured to determine whether the thread identifier corresponding to the shard information is the thread identifier of the thread; if so, the thread executes the step of obtaining corresponding sharded data to be processed from the database according to the shard information; if not, the thread executes the step of looking up unmarked shard information from all shard information.
6. The apparatus according to claim 5, wherein, The processing unit is further configured to: If the thread fails to process the sharded data, determine the processing status of the shard information of the sharded data; If the processing status of the shard information is processing failed, send an alarm indication; If the processing status of the shard information is being processed, change the processing status of the shard information to unprocessed.
7. An electronic device, wherein, It includes: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-4.
8. A non-transitory computer-readable storage medium, wherein, The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the method according to any one of claims 1-4.
Citation Information
Patent Citations
Distributed text pre-detection method and device
CN111145028A