Data extraction method, system and device, medium and product

By using in-memory database for queue scheduling and detailed scheduling in the data extraction system, the performance and stability of the data extraction system under high business volume is solved, timely storage and stability of data is achieved, operation and maintenance costs are reduced, and network fluctuations and transaction volume changes are adapted.

CN120578468APending Publication Date: 2025-09-02AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510697345.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

Existing data extraction systems are difficult to ensure performance and stability under high business volume and limited maintenance costs, especially in unexpected situations such as network fluctuations. The classic task scheduling framework increases external dependence and operation and maintenance costs.

Method used

The in-memory database is used for queue scheduling, and the decimation job tasks are published through the dispatch node, and the work nodes that meet the response conditions are extracted and detailed scheduling. The key-value database Redis of the in-memory database is used to achieve atomic operations to avoid concurrency problems, and the decimation task is handled using four queues and five detailed structures in the in-memory database.

Benefits of technology

It improves the efficiency of data extraction, ensures timely storage and stability of data, reduces operation and maintenance costs, can cope with unexpected situations such as network fluctuations, and achieves stable operation and dynamic expansion and expansion of 7×24 hours.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120578468A_ABST
    Figure CN120578468A_ABST
Patent Text Reader

Abstract

The invention discloses a data extraction method, system and device, a medium and a product, and the method comprises the steps: issuing a data extraction operation task through a scheduling node, determining a data extraction time period, and carrying out the queue scheduling of the data extraction time period in a memory database; and performing data extraction and detail scheduling on the data extraction task in the data extraction time period through the working node meeting the response condition. According to the technical scheme, the performance and stability of data extraction are guaranteed, timely storage of service data is achieved, and unforeseen circumstances such as network fluctuation can be stably coped with.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a data extraction method, system, device, medium and product. Background Art

[0002] Third-party service platforms connect to banks through intermediaries. These connected banking systems are available to users 24 / 7, generating a significant volume of business. This business data is a valuable resource for banks and must be extracted and stored in a data integration facility for further processing and utilization.

[0003] However, extremely high business volume and limited maintenance costs place high demands on the performance and stability of the data extraction system. Summary of the Invention

[0004] The present invention provides a data extraction method, system, device, medium and product, which ensure the performance and stability of data extraction, realize the timely storage of business data, and can stably respond to unexpected situations such as network fluctuations.

[0005] In a first aspect, an embodiment of the present disclosure provides a data extraction method, which is applied to a data extraction system, wherein the data extraction system includes a scheduling node and multiple working nodes, the scheduling node being connected to each of the working nodes respectively, and the method includes:

[0006] Through the scheduling node, the number extraction task is issued, the number extraction time period is determined, and the number extraction time period is queued in the memory database;

[0007] Data extraction and detailed scheduling are performed on the extraction tasks in the extraction time period through the working nodes that meet the response conditions.

[0008] In a second aspect, an embodiment of the present disclosure provides a data extraction system, comprising: a scheduling node and a plurality of working nodes, wherein the scheduling node is connected to each of the working nodes respectively;

[0009] The scheduling node is used to issue a number extraction task, determine a number extraction time period, and perform queue scheduling for the number extraction time period in a memory database;

[0010] The working node that meets the response conditions is used to extract data and perform detailed scheduling on the extraction tasks in the extraction time period.

[0011] In a third aspect, an embodiment of the present disclosure provides an electronic device, including:

[0012] at least one processor; and

[0013] a memory communicatively connected to at least one processor; wherein,

[0014] The memory stores a computer program that can be executed by at least one processor, and the computer program is executed by at least one processor so that the at least one processor can execute a data extraction method provided by the above-mentioned first aspect embodiment.

[0015] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, which stores computer instructions. The computer instructions are used to enable a processor to implement a data extraction method provided in the embodiment of the first aspect when executed.

[0016] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements a data extraction method provided in the embodiment of the first aspect.

[0017] The data extraction method, system, device, medium and product of the embodiment of the present invention publish the extraction task through the scheduling node, determine the extraction time period, and queue schedule the extraction time period in the memory database; through the working node that meets the response conditions, the extraction task in the extraction time period is extracted and detailed scheduling is performed. The above technical solution completes data extraction in the memory database, improves the efficiency of data extraction, ensures the performance of data extraction, and realizes the timely storage of business data. The extraction working node must meet the response conditions to effectively ensure the stability of data extraction to cope with unexpected situations such as network fluctuations.

[0018] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0020] Figure 1 This is a flow chart of a data extraction method provided in Example 1 of the present invention;

[0021] Figure 2 This is a schematic diagram of a queue in a memory database provided by the first embodiment of the present invention;

[0022] Figure 3This is a detailed schematic diagram of a sampling time period provided by the first embodiment of the present invention;

[0023] Figure 4 This is a schematic diagram of a key identification data structure in a memory database provided by the first embodiment of the present invention;

[0024] Figure 5 This is a queue scheduling diagram for the current number extraction queue provided by the first embodiment of the present invention;

[0025] Figure 6 This is a queue scheduling diagram for a delayed extraction queue provided by the first embodiment of the present invention;

[0026] Figure 7 This is a queue scheduling diagram for a queue that fails to extract data, provided by the first embodiment of the present invention;

[0027] Figure 8 This is a detailed scheduling diagram for a sampling time period provided by the first embodiment of the present invention;

[0028] Figure 9 This is a structural diagram of a data extraction system provided by Embodiment 2 of the present invention;

[0029] Figure 10 This is a structural diagram of an electronic device provided in Example 3 of the present invention. DETAILED DESCRIPTION

[0030] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0031] It should be noted that the terms "first," "second," and "target" and the like in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this way are interchangeable where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatus.

[0032] Third-party payment institutions connect with banks through intermediaries, often referred to as express payment systems. These systems operate 24 / 7 and handle a massive transaction volume. This transaction data is a valuable intangible asset for banks and must be extracted into a data lake for further processing and utilization. However, the combination of high transaction volumes and limited operational and maintenance costs places high demands on the performance and stability of the extraction system. To ensure timely entry of all transaction data into the lake, the extraction system, like the express system, must operate stably 24 / 7 and must include a failover mechanism to mitigate unexpected situations such as network jitter. To accommodate future fluctuations in the express payment system's transaction volume, the extraction system must support dynamic scaling and efficient cluster scheduling to avoid bottlenecks in cluster size.

[0033] For periodic tasks such as data extraction (i.e., extracting data from the source data system to the target system, which in this case can be understood as the full data extraction of quick payment transaction details), using a classic task scheduling framework (Quartz) is the most intuitive idea. This type of framework has complete development documentation, is fully compatible with mainstream technology stacks, can be used to execute scheduled tasks, and also supports cluster deployment. In addition, this type of framework often supports defining and modifying the time expression of scheduled tasks, and supports persisting tasks in the database, so that scheduled tasks can still be executed smoothly after the database is restarted.

[0034] However, the classic task scheduling framework has two shortcomings that make it impractical for quick payment withdrawal systems. The first is database dependency, which requires a database to store task and scheduling information. This not only increases the system's external dependencies but also affects task scheduling performance, creating bottlenecks in the withdrawal system's performance and making it difficult to meet demand. The second is cluster complexity. Although the framework supports cluster mode and can distribute tasks across multiple server instances, setting up and maintaining a usable cluster environment is relatively complex. The database storage needs to be correctly configured, and the time and time synchronization configurations must be consistent across all nodes. This undoubtedly increases the operation and maintenance costs of the withdrawal system, which is contrary to the requirements of the withdrawal system.

[0035] It is understandable that before using the technical solutions disclosed in the embodiments of the present invention, the type, scope of use, usage scenarios, etc. of the personal information involved in the present invention should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant regulations.

[0036] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operation of the technical solution of the present invention based on the prompt message.

[0037] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0038] It is understandable that the above notification and user authorization obtaining process is merely illustrative and does not limit the implementation of the present invention. Other methods that meet relevant regulations may also be applied to the implementation of the present invention.

[0039] Example 1

[0040] Figure 1 This is a flow chart of a data extraction method provided in Example 1 of the present invention. This embodiment is applicable to scenarios where business data is extracted. The method can be executed by a data extraction system, which can be implemented in hardware and / or software. The data extraction system can be a cluster containing several containers / nodes. These containers compete for a scheduling node. Containers that fail to compete successfully serve as working nodes, ultimately forming a scheduling node and multiple working nodes. As a data extraction system, the scheduling node communicates with each working node separately.

[0041] like Figure 1 As shown, the method includes:

[0042] S101. Publish a number extraction task through a scheduling node, determine a number extraction time period, and perform queue scheduling for the number extraction time period in a memory database.

[0043] In this embodiment, the scheduling node can be understood as a node for overall scheduling, that is, a Scheduler node. The scheduling node is used to publish the extraction job tasks according to the configuration, and coordinate the status of the data extraction system. The extraction job tasks can be understood as tasks used to instruct the scheduling node to execute queue scheduling and instruct the working node to execute data extraction and detailed scheduling. The extraction time period can be understood as a collection of data extraction task details divided based on time periods. The in-memory database is a database management system that stores data completely or mainly in the computer memory. It greatly improves the processing speed and significantly reduces latency by directly operating the memory data. The in-memory database used in the embodiment of the present invention is the key-value database Redis. The built-in script environment in the in-memory database is used to ensure atomicity, so that the extraction cluster will not produce unexpected states due to concurrency.

[0044] The in-memory database includes at least four queue lists: Current, Delay, Success, and FAIL. Each list stores a single extraction time period. Each queue includes at least one extraction time period, and each extraction time period includes at least one extraction task. For each queue in the in-memory database, write operations occur on the right side of the list, and all read operations occur on the left side. Figure 2 This is a schematic diagram of a queue in a memory database provided by the first embodiment of the present invention. Figure 2As shown, {DE}CURRENTSCHEDULE is the current extraction queue, which stores newly released extraction time periods that have never been executed, such as 300_1714118100, 300_1714118700, and 300_1714118400; {DE}DELAY SCHEDULE is the delayed extraction queue, which stores blocked extraction time periods, such as 300_1714117200, 300_1714117500, and 300_1714117800. If the transaction details of a certain time period are not successfully entered into the lake for a long time, they will be moved to the delayed extraction queue to avoid affecting other normal extraction tasks of the current extraction queue; {DE}SUCCESS SCHEDULE is the successful extraction queue, which stores the extraction time periods in which data has been successfully extracted, such as 300_1714116300, 300_1714116900, and 300_1714115700. The current extraction progress can be known based on the successful extraction queue. {DE}FAILSCHEDULE is the failed extraction queue, which stores the extraction time periods in which data extraction failed due to various reasons, such as 300_1714116000, 300_1714115400, and 300_1714116600. These extraction time periods will be repeatedly retrieved until they are successful, ensuring that the AC flow is not lost during the entry into the lake.

[0045] The sampling time period is represented by a string in the form of 300_1714118100, where 300 is the length of the time period and 1714118100 is the timestamp of the end time of the time period. The unit of both is seconds.

[0046] Each draw time period includes multiple draw tasks, so you must also record the draw details for each draw time period. Each draw time period should include at least five detail sets: ALL, TODO, SUC, FAIL, and RUN. Figure 3 This is a detailed schematic diagram of a sampling time period provided by the first embodiment of the present invention, such as Figure 3As shown, taking the lottery time period 300_1714118100 as an example, taskStr is a character string describing the lottery task, {DE}_ALL is the total details of the lottery time period, which stores all lottery tasks in the lottery time period, such as task1Str, task2Str, task3Str, and task4Str; {DE}_TODO is the pending lottery details, which stores the lottery tasks to be executed in the lottery time period, such as task1Str; {DE}_SUC is the successful lottery details, which stores the successful lottery tasks in the lottery time period, such as task2Str; {DE}_FAIL is the failed lottery details, which stores the failed lottery tasks in the lottery time period, such as task3Str; {DE}_RUN is the current lottery details, which stores the lottery tasks being executed in the lottery time period, such as task4Str. In addition, the total amount of transaction details extracted in this time period is recorded in the data volume parameter VOLUME, and the publishing time of the extraction task in this time period is recorded in the publishing time parameter PUBLISH_TIME.

[0047] It can be understood that the total details include all the drawing tasks in the waiting drawing details, successful drawing details, failed drawing details and current drawing details. There are no repeated drawing tasks between the waiting drawing details, successful drawing details, failed drawing details and current drawing details. The same applies to the current drawing queue, delayed drawing queue, successful drawing queue and failed drawing queue. There are no repeated drawing time periods.

[0048] Figure 4 This is a schematic diagram of a key identification data structure in a memory database provided by the first embodiment of the present invention. Figure 4 As shown, in the data structure of the memory database Key, ENABLE is used as the extraction switch (ON means on, OFF means off), SUBMIT_UNTILL is used to record the release progress of the extraction task, ALL_TASK_SET is used to record all the extraction tasks required to be included in the newly released extraction time period, the scheduling node heartbeat parameter SCHEDULER_BEACON is used to record the heartbeat of the scheduling node in the data extraction system, and the extraction task heartbeat parameter TASK_BEACON_HASH is used to record the heartbeat of all the extraction tasks being executed. Among them, the field of TASK_BEACON_HASH is composed of the extraction time period and the extraction tasks of the extraction time period, and the value is composed of the timestamp of the last time the task execution heartbeat was received and the container name (Pod name) of the worker node that executes the corresponding extraction task task. Figure 4It can be seen that worker node 1 (Worker1) executes extraction task 1 (task1Str) in the extraction time period 300_1714118100, and worker node 2 (Worker2) executes extraction task 3 (task3Str) in the extraction time period 300_1714116600. The value of SCHEDULER_BEACON is composed of the timestamp of the last scheduler heartbeat time and the container name (pod name) of the scheduler.

[0049] Specifically, the scheduling node determines whether the current time meets the task release conditions. Specifically, the current time must be greater than the end time of the lottery time period corresponding to the last lottery task released, to avoid drawing unfinished transactions. If the task release conditions are met, the scheduling node releases the lottery task, dividing the 24-hour lottery task into several time periods at regular intervals. A lottery task detail is then created for each time period to record the execution of each lottery task. These time periods containing lottery task details are referred to as lottery time periods. All lottery tasks, ALL_TASK_SET, are copied to the total details (ALL) and pending lottery details (TODO) for that lottery time period, recording the release time (PUBLISH_TIME) in the details. The lottery task's release progress (SUBMIT_UNTILL) is updated, and the newly released lottery time period is placed in the current lottery queue (Current).

[0050] It is understandable that in addition to publishing the extraction job task and executing the queue scheduling, the scheduling node also needs to send a heartbeat. Each time a heartbeat is sent, the scheduling node will update the timestamp of the heartbeat SCHEDULER_BEACON corresponding to the scheduling node and the scheduling node identifier schedulerStr to the current value.

[0051] There is only one scheduling node in the data extraction system, which is competed by multiple working nodes. When there is a problem with the scheduling node sending heartbeats and the task of sending heartbeats is not executed within a certain period of time, a new scheduling node will be spontaneously competed from the working nodes.

[0052] The scheduling node executes the release of the extraction job task, queue scheduling and heartbeat sending in a round-robin manner.

[0053] S102. Data extraction and detailed scheduling are performed on the number extraction tasks in the number extraction time period through the working nodes that meet the response conditions.

[0054] In this embodiment, the response condition can be understood as a condition for determining whether a worker node can perform a data extraction task, such as a continuous heartbeat update of the worker node. A worker node can be understood as a node for performing data extraction, i.e., a worker node, which performs a specific extraction task.

[0055] After being connected to the data extraction cluster, each container becomes a worker node by default. Therefore, worker nodes make up the majority of the data extraction system. Each worker node must perform three tasks: competing for node scheduling, performing data extraction, and sending heartbeats. Worker nodes continuously perform these three tasks in a round-robin manner.

[0056] Wherein, during the competition scheduling node, each working node can periodically explore the scheduling node according to the heartbeat (SCHEDULER_BEACON) of scheduling node, if scheduling node loses response or current extraction cluster (data extraction system) also does not exist scheduling node, then there will be a working node competition to become scheduling node, and the competition mode present embodiment is reluctant to limit this. Each working node will update the description of this task together with self container name to TASK_BEACON_HASH when executing the extraction task, after this extraction task has been performed, can delete corresponding record among the TASK_BEACON_HASH, so that scheduling node judges whether the extraction task that working node performs loses response according to the update time of TASK_BEACON_HASH corresponding record and current time. If lose response, then do not meet response condition, if still have response, then determine that this working node meets response condition.Based on the working node that meets response condition, continue to carry out data extraction and detailed scheduling to the extraction task in the extraction time period based on the extraction operation task issued by scheduling node.

[0057] A data extraction method provided by an embodiment of the present invention includes obtaining an extraction time period through a scheduling node, storing the extraction time period in a memory database, and performing queue scheduling on the extraction time period stored in the memory database; and extracting data and performing detailed scheduling on the extraction tasks in the extraction time period through a working node that meets a response condition. The above technical solution completes data extraction in the memory database, improves the efficiency of data extraction, ensures the performance of data extraction, and realizes the timely storage of business data. The extraction working node must meet the response condition to effectively ensure the stability of data extraction to cope with unexpected situations such as network fluctuations.

[0058] The scheduling tasks of the scheduling node are relatively complex. The scheduling node not only needs to properly move a certain time period among the four queues of Current, Delay, Success, and Fail based on the extraction details of the extraction time period, but also needs to infer failure of the extraction tasks that have lost response based on TASK_BEACON_HASH. Specifically, the scheduling tasks can be divided into three categories according to the extraction queue in which the scheduled extraction time period is initially located (of the four queues, the Success queue does not need to be scheduled). Examples one to three respectively illustrate the scheduling logic for tasks initially located in the Current, Delay, and Fail queues, where SUC=ALL indicates whether the SUC and ALL in the details are equal, that is, whether all extraction tasks are successful.

[0059] As a first alternative embodiment of this embodiment, Figure 5 This is a queue scheduling diagram for the current number extraction queue provided by the first embodiment of the present invention, such as Figure 5 As shown in the figure, queue scheduling is performed for the extraction time period in the memory database, including:

[0060] a1. For the number drawing time period in the current number drawing queue, determine whether the number drawing details corresponding to the number drawing time period are empty.

[0061] In this embodiment, the current number extraction queue Current includes several number extraction time periods. For each number extraction time period, it is determined whether the to-be-extracted number details TODO included in the number extraction time period is empty.

[0062] b1. When the details of the numbers to be drawn are not empty, the time period for drawing numbers is retained in the current drawing number queue.

[0063] If the number to be drawn TODO is not empty, keep the number drawing time period in the current number drawing queue Current, and end the judgment logic.

[0064] c1. If the pending number extraction details are empty and the current number extraction details are not empty, determine whether the number extraction task in the current number extraction details has timed out. If not, retain the number extraction time period in the current number extraction queue. If timed out, schedule the number extraction time period from the current number extraction queue to the delayed number extraction queue.

[0065] In this embodiment, if the details of the number to be extracted TODO is empty, it is determined whether the current extraction details RUN is empty. If the details of the number to be extracted TODO is empty and the current extraction details RUN is not empty, it is determined whether the extraction task in the current extraction details RUN has timed out (for example, the time for executing data extraction is greater than the set data extraction time threshold). If it has timed out, the extraction time period is scheduled from the current extraction queue Current to the delayed extraction queue Delay, and the judgment logic ends; if it has not timed out, the extraction time period is retained in the current extraction queue Current, and the judgment logic ends.

[0066] d1. When the pending number extraction details are empty and the current number extraction details are empty, determine whether all the number extraction tasks in the total details have been completed. If so, schedule the number extraction time period from the current number extraction queue to the number extraction success queue. If not, schedule the number extraction time period from the current number extraction queue to the number extraction failure queue.

[0067] In this embodiment, if the details of the number to be drawn TODO is empty, it is determined whether the current drawing details RUN is empty. If the details of the number to be drawn TODO is empty and the current drawing details RUN is also empty, it is determined whether all the drawing tasks in the total details ALL have been completed (SUC=ALL). If SUC=ALL is satisfied, the drawing time period is scheduled from the current drawing queue Current to the drawing success queue Success, and the judgment logic ends; if SUC=ALL is not satisfied, the drawing time period is scheduled from the current drawing queue Current to the drawing failure queue FAIL, and the judgment logic ends.

[0068] As a second alternative embodiment of this embodiment, Figure 6 1 is a queue scheduling diagram for a delayed extraction queue provided by the first embodiment of the present invention, such as Figure 6 As shown in the figure, queue scheduling is performed for the extraction time period in the memory database, including:

[0069] a2. For the sampling time period in the delayed sampling queue, determine whether the current sampling details corresponding to the sampling time period are empty.

[0070] In this embodiment, the delayed sampling queue Delay includes several sampling time periods. For each sampling time period, it is determined whether the current sampling detail RUN included in the sampling time period is empty.

[0071] b2. When the current extraction detail is not empty, traverse the current extraction detail, schedule the extraction tasks that have lost the execution response in the current extraction detail from the current extraction detail to the extraction failure detail, and keep the extraction time period in the delayed extraction queue.

[0072] In this embodiment, for the extraction time period, if the current extraction detail RUN is not empty, the current extraction detail RUN is traversed, and the extraction task that has lost the execution response is determined based on TASK_BEACON_HASH. The extraction task that has lost the execution response in the current extraction detail RUN is scheduled to the extraction failure detail Fail, and the extraction time period is retained in the delayed extraction queue Delay, ending the judgment logic.

[0073] c2. When the current extraction details are empty, determine whether all extraction tasks in the total details have been completed. If so, schedule the extraction time period from the delayed extraction queue to the successful extraction queue. If not, schedule the extraction time period from the delayed extraction queue to the failed extraction queue.

[0074] In this embodiment, for the extraction time period, if the current extraction detail RUN is empty, it is determined whether all extraction tasks in the total detail ALL have been completed (SUC=ALL). If SUC=ALL is satisfied, the extraction time period is dispatched from the extraction queue Delay to the extraction success queue Success, and the judgment logic ends; if SUC=ALL is not satisfied, the extraction time period is dispatched from the extraction queue Delay to the extraction failure queue FAIL, and the judgment logic ends.

[0075] As a third alternative embodiment of this embodiment, Figure 7 This is a queue scheduling diagram for a queue with failed extraction, provided by the first embodiment of the present invention. Figure 7 As shown in the figure, queue scheduling is performed for the extraction time period in the memory database, including:

[0076] a3. For the number extraction time period in the number extraction failure queue, determine whether the current number extraction details corresponding to the number extraction time period are empty.

[0077] In this embodiment, the sampling failure queue FAIL includes several sampling time periods. For each sampling time period, it is determined whether the current sampling detail RUN included in the sampling time period is empty.

[0078] b3. When the current extraction detail is not empty, traverse the current extraction detail, schedule the extraction tasks that have lost the execution response in the current extraction detail from the current extraction detail to the extraction failure detail, and keep the extraction time period in the extraction failure queue.

[0079] In this embodiment, for the extraction time period, if the current extraction detail RUN is not empty, the current extraction detail RUN is traversed, and the extraction task that has lost the execution response is determined based on TASK_BEACON_HASH. The extraction task that has lost the execution response in the current extraction detail RUN is scheduled to the extraction failure detail Fail, and the extraction time period is retained in the extraction failure queue FAIL, ending the judgment logic.

[0080] c3. When the current number extraction details are empty, determine whether all number extraction tasks in the total details have been completed. If so, schedule the number extraction time period from the number extraction failure queue to the number extraction success queue. If not, keep the number extraction time period in the number extraction failure queue.

[0081] In this embodiment, for the extraction time period, if the current extraction detail RUN is empty, it is determined whether all extraction tasks in the total detail ALL have been completed (SUC=ALL). If SUC=ALL is satisfied, the extraction time period is scheduled from the extraction failure queue FAIL to the extraction success queue Success, and the judgment logic ends; if SUC=ALL is not satisfied, the extraction time period is retained in the extraction failure queue FAIL, and the judgment logic ends.

[0082] The logic for executing data extraction on the working node needs to be divided into two parts: normal execution of the extraction task that has not been executed and re-execution of the extraction task that has failed. As a fourth optional embodiment of this embodiment, Figure 8 This is a detailed scheduling diagram for a sampling time period provided by the first embodiment of the present invention, such as Figure 8 As shown, data extraction and detailed scheduling are performed for the number extraction tasks in the number extraction time period, including:

[0083] a4. Determine whether the first number extraction condition is currently met. The first number extraction condition is that the current number extraction queue is empty or the number to be extracted details of the number extraction time period corresponding to the head of the current number extraction queue is empty.

[0084] In this embodiment, the first extraction condition can be understood as a extraction condition used to determine the extraction task execution position. The first extraction condition is that the current extraction queue is empty or the details of the numbers to be extracted in the head extraction time period corresponding to the current extraction queue are empty. The head extraction time period can be understood as the first extraction time period in the current extraction queue.

[0085] Specifically, it is determined whether the current number extraction queue Current is empty or the number to be extracted details TODO of the head number extraction time period corresponding to the current number extraction queue Current is empty.

[0086] b4. If the first extraction condition is not met, the target extraction task is taken from the waiting extraction details of the extraction time period corresponding to the head of the current extraction queue, the target extraction task is scheduled from the waiting extraction details to the current extraction details, and the target extraction task is executed.

[0087] In this embodiment, the target extraction task can be understood as an extraction task for data extraction to be executed, which can be a head extraction task in an extraction time period, or other extraction tasks, and this embodiment does not set any limitation on this.

[0088] Specifically, if the current drawing queue Current is not empty and the details of the numbers to be drawn TODO of the head drawing time period corresponding to the current drawing queue Current is not empty, take the target drawing task from the details of the numbers to be drawn TODO of the head drawing time period corresponding to the current drawing queue, schedule the target drawing task from the details of the numbers to be drawn TODO to the current drawing details RUN, execute the target drawing task, and extract the corresponding business data into the lake.

[0089] c4. If the first extraction condition is met, determine whether the second extraction condition is currently met. The second extraction condition is that the extraction failure queue is empty or the extraction failure details of the extraction time period at the head of the corresponding queue of the extraction failure queue are empty.

[0090] In this embodiment, the second extraction condition can be understood as the extraction condition used to determine the extraction task execution position. The second extraction condition is that the extraction failure queue is empty or the extraction failure details of the extraction time period at the head of the extraction queue corresponding to the extraction failure queue are empty.

[0091] Specifically, if the current extraction queue Current is empty or the to-be-extracted number details TODO of the head extraction time period corresponding to the current extraction queue Current is empty, determine whether the current extraction failure queue FAIL is empty or the extraction failure details Fail of the head extraction time period corresponding to the head extraction queue FAIL is empty.

[0092] d4. If the second extraction condition is not met, the target extraction task is taken from the extraction failure details of the extraction time period at the head of the corresponding queue of the extraction failure queue, the target extraction task is scheduled from the extraction failure details to the current extraction details, and the target extraction task is executed.

[0093] In this embodiment, if the extraction failure queue FAIL is empty or the extraction failure detail Fail of the first extraction time period corresponding to the extraction failure queue FAIL is empty, the judgment logic ends; if the extraction failure queue FAIL is not empty and the extraction failure detail Fail of the first extraction time period corresponding to the extraction failure queue FAIL is not empty, the target extraction task is taken from the extraction failure detail Fail of the first extraction time period corresponding to the extraction failure queue FAIL, and the target extraction task is scheduled from the extraction failure detail Fail to the current extraction detail RUN, and the target extraction task is executed to extract the corresponding business data into the lake.

[0094] It can be seen that the working node will give priority to obtaining the extraction tasks that have not been executed, and then re-acquire the extraction tasks that have failed to execute.

[0095] e4. Determine whether the target extraction task has an execution error. If so, schedule the target extraction task from the current extraction detail to the extraction failure detail. If not, schedule the target extraction task from the current extraction detail to the extraction success detail and update the total extraction amount.

[0096] In this embodiment, when executing a lottery task, it is determined whether an error occurs during the task execution. If an error occurs, the target lottery task is scheduled from the current lottery detail RUN to the lottery failure detail Fail. If no error occurs, the target lottery task is scheduled from the current lottery detail RUN to the lottery success detail SUC, and the total amount of lottery task execution is updated and accumulated VOLUME.

[0097] The above technical solution uses an in-memory database to implement the scheduling of the sampling cluster, avoiding the cluster scale bottleneck caused by the use of traditional databases and ensuring the overall scheduling performance of the sampling cluster; the data extraction method adopted by the present invention is based on 4 sampling queues and 5 types of sampling details, which can handle common operation and maintenance situations such as sampling failure, loss of response of sampling applications, and expansion and contraction caused by changes in transaction volume, thereby achieving lower operation and maintenance costs; and achieving a balance between the performance and operation and maintenance costs of the sampling cluster.

[0098] Example 2

[0099] Figure 9 This is a structural diagram of a data extraction system provided by the second embodiment of the present invention. Figure 9 As shown, the system includes: a scheduling node 21 and multiple working nodes 22, wherein the scheduling node 21 is connected to each of the working nodes 22;

[0100] The scheduling node 21 is used to issue a number extraction task, determine a number extraction time period, and perform queue scheduling for the number extraction time period in a memory database;

[0101] The working node 22 that meets the response condition is used to extract data and perform detailed scheduling on the extraction tasks in the extraction time period.

[0102] The data extraction system adopted in this technical solution guarantees the performance and stability of data extraction, realizes the timely storage of business data, and can stably respond to unexpected situations such as network fluctuations.

[0103] Optionally, the memory database includes at least a current extraction queue, a delayed extraction queue, a successful extraction queue, and a failed extraction queue, each queue includes at least one extraction time period, and each extraction time period includes at least one extraction task;

[0104] For each lottery time period, each lottery time period includes at least total details, pending lottery details, successful lottery details, failed lottery details and current lottery details.

[0105] Optionally, the scheduling node 21 is specifically configured to:

[0106] For a number extraction time period in the current number extraction queue, determining whether the number to be extracted details corresponding to the number extraction time period is empty;

[0107] If the number to be drawn details is not empty, retain the number drawing time period in the current number drawing queue;

[0108] If the pending number extraction details are empty and the current number extraction details are not empty, determining whether the number extraction task in the current number extraction details has timed out, if not, retaining the number extraction time period in the current number extraction queue; if timed out, scheduling the number extraction time period from the current number extraction queue to the delayed number extraction queue;

[0109] When the to-be-drawn number details are empty and the current number-drawing number details are empty, determine whether all number-drawing tasks in the total details have been completed; if so, schedule the number-drawing time period from the current number-drawing queue to the number-drawing success queue; if not, schedule the number-drawing time period from the current number-drawing queue to the number-drawing failure queue.

[0110] Optionally, the scheduling node 21 is specifically configured to:

[0111] For a sampling time period in the delayed sampling queue, determining whether a current sampling detail corresponding to the sampling time period is empty;

[0112] If the current extraction details are not empty, traverse the current extraction details, schedule the extraction tasks that have lost execution responses in the current extraction details from the current extraction details to the extraction failure details, and keep the extraction time period in the delayed extraction queue;

[0113] When the current extraction details are empty, determine whether all extraction tasks in the total details have been completed. If so, schedule the extraction time period from the delayed extraction queue to the successful extraction queue. If not, schedule the extraction time period from the delayed extraction queue to the failed extraction queue.

[0114] Optionally, the scheduling node 21 is specifically configured to:

[0115] For the number extraction time period in the number extraction failure queue, determining whether the current number extraction details corresponding to the number extraction time period are empty;

[0116] If the current draw detail is not empty, traverse the current draw detail, schedule the draw task that has lost the execution response in the current draw detail from the current draw detail to the draw failure detail, and keep the draw time period in the draw failure queue;

[0117] When the current extraction details are empty, determine whether all extraction tasks in the total details have been completed. If so, schedule the extraction time period from the extraction failure queue to the extraction success queue. If not, keep the extraction time period in the extraction failure queue.

[0118] Optionally, the working node 22 is specifically configured to:

[0119] Determining whether a first number extraction condition is currently satisfied, wherein the first number extraction condition is that the current number extraction queue is empty or the number to be extracted details of the head number extraction time period corresponding to the current number extraction queue is empty;

[0120] If the first extraction condition is not met, taking a target extraction task from the waiting extraction details of the head extraction time period corresponding to the current extraction queue, scheduling the target extraction task from the waiting extraction details to the current extraction details, and executing the target extraction task;

[0121] If the first extraction condition is satisfied, determining whether a second extraction condition is currently satisfied, the second extraction condition being that the extraction failure queue is empty or the extraction failure details of the extraction time period at the head of the extraction failure queue are empty;

[0122] If the second extraction condition is not met, taking a target extraction task from the extraction failure details of the extraction time period at the head of the queue corresponding to the extraction failure queue, scheduling the target extraction task from the extraction failure details to the current extraction details, and executing the target extraction task;

[0123] Determine whether the target number extraction task has an execution error. If so, schedule the target number extraction task from the current number extraction detail to the number extraction failure detail. If not, schedule the target number extraction task from the current number extraction detail to the number extraction success detail, and update the total number of extractions.

[0124] The data extraction system provided by the embodiment of the present invention can execute the data extraction method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0125] Example 3

[0126] Figure 10: is a structural diagram of an electronic device provided in Example 3 of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0127] like Figure 10 As shown, the electronic device 30 includes at least one processor 31 and a memory, such as a read-only memory (ROM) 32, a random access memory (RAM) 33, etc., which is communicatively connected to the at least one processor 31. The memory stores a computer program that can be executed by the at least one processor. The processor 31 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 32 or the computer program loaded from the storage unit 38 into the random access memory (RAM) 33. Various programs and data required for the operation of the electronic device 30 can also be stored in the RAM 33. The processor 31, ROM 32, and RAM 33 are connected to each other via a bus 34. An input / output (I / O) interface 35 is also connected to the bus 34.

[0128] Multiple components in the electronic device 30 are connected to the I / O interface 35, including an input unit 36, such as a keyboard, a mouse, etc.; an output unit 37, such as various types of displays, speakers, etc.; a storage unit 38, such as a magnetic disk, an optical disk, etc.; and a communication unit 39, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 39 allows the electronic device 30 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0129] The processor 31 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the processor 31 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 31 executes the various methods and processes described above, such as the data extraction method.

[0130] In some embodiments, the data extraction method can be implemented as a computer program that is tangibly contained in a computer-readable storage medium, such as storage unit 38. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 30 via ROM 32 and / or communication unit 39. When the computer program is loaded into RAM 33 and executed by processor 31, one or more steps of the data extraction method described above can be performed. Alternatively, in other embodiments, processor 31 can be configured to perform the data extraction method in any other suitable manner (e.g., by means of firmware).

[0131] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0132] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0133] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0134] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0135] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0136] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0137] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0138] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A data extraction method, characterized in that: Applied to a data extraction system, the data extraction system includes a scheduling node and multiple working nodes, the scheduling node is connected to each of the working nodes respectively, and the method includes: Through the scheduling node, the number extraction task is issued, the number extraction time period is determined, and the number extraction time period is queued in the memory database; Data extraction and detailed scheduling are performed on the extraction tasks in the extraction time period through the working nodes that meet the response conditions.

2. The method according to claim 1, characterized in that The memory database includes at least a current extraction queue, a delayed extraction queue, a successful extraction queue, and a failed extraction queue, each queue includes at least one extraction time period, and each extraction time period includes at least one extraction task; For each lottery time period, each lottery time period includes at least total details, pending lottery details, successful lottery details, failed lottery details and current lottery details.

3. The method according to claim 2, characterized in that The performing queue scheduling on the extraction time period in the memory database includes: For a number extraction time period in the current number extraction queue, determining whether the number to be extracted details corresponding to the number extraction time period is empty; If the number to be drawn details is not empty, retain the number drawing time period in the current number drawing queue; If the pending number extraction details are empty and the current number extraction details are not empty, determining whether the number extraction task in the current number extraction details has timed out, if not, retaining the number extraction time period in the current number extraction queue; if timed out, scheduling the number extraction time period from the current number extraction queue to the delayed number extraction queue; When the to-be-drawn number details are empty and the current number-drawing number details are empty, determine whether all number-drawing tasks in the total details have been completed; if so, schedule the number-drawing time period from the current number-drawing queue to the number-drawing success queue; if not, schedule the number-drawing time period from the current number-drawing queue to the number-drawing failure queue.

4. The method according to claim 2, characterized in that The performing queue scheduling on the extraction time period in the memory database includes: For a sampling time period in the delayed sampling queue, determining whether a current sampling detail corresponding to the sampling time period is empty; If the current extraction details are not empty, traverse the current extraction details, schedule the extraction tasks that have lost execution responses in the current extraction details from the current extraction details to the extraction failure details, and keep the extraction time period in the delayed extraction queue; When the current extraction details are empty, determine whether all extraction tasks in the total details have been completed. If so, schedule the extraction time period from the delayed extraction queue to the successful extraction queue. If not, schedule the extraction time period from the delayed extraction queue to the failed extraction queue.

5. The method according to claim 2, characterized in that The performing queue scheduling on the extraction time period in the memory database includes: For the number extraction time period in the number extraction failure queue, determining whether the current number extraction details corresponding to the number extraction time period are empty; If the current draw detail is not empty, traverse the current draw detail, schedule the draw task that has lost the execution response in the current draw detail from the current draw detail to the draw failure detail, and keep the draw time period in the draw failure queue; When the current extraction details are empty, determine whether all extraction tasks in the total details have been completed. If so, schedule the extraction time period from the extraction failure queue to the extraction success queue. If not, keep the extraction time period in the extraction failure queue.

6. The method according to claim 2, characterized in that The data extraction and detailed scheduling of the number extraction tasks in the number extraction time period include: Determining whether a first number extraction condition is currently satisfied, wherein the first number extraction condition is that the current number extraction queue is empty or the number to be extracted details of the head number extraction time period corresponding to the current number extraction queue is empty; If the first extraction condition is not met, taking a target extraction task from the waiting extraction details of the head extraction time period corresponding to the current extraction queue, scheduling the target extraction task from the waiting extraction details to the current extraction details, and executing the target extraction task; If the first extraction condition is satisfied, determining whether a second extraction condition is currently satisfied, the second extraction condition being that the extraction failure queue is empty or the extraction failure details of the extraction time period at the head of the extraction failure queue are empty; If the second extraction condition is not met, taking a target extraction task from the extraction failure details of the extraction time period at the head of the queue corresponding to the extraction failure queue, scheduling the target extraction task from the extraction failure details to the current extraction details, and executing the target extraction task; Determine whether the target number extraction task has an execution error. If so, schedule the target number extraction task from the current number extraction detail to the number extraction failure detail. If not, schedule the target number extraction task from the current number extraction detail to the number extraction success detail, and update the total number of extractions.

7. A data extraction system, characterized in that: Executing the method according to any one of claims 1 to 6, the system comprises a scheduling node and a plurality of working nodes, the scheduling node being connected to each of the working nodes respectively; The scheduling node is used to issue a number extraction task, determine a number extraction time period, and perform queue scheduling for the number extraction time period in a memory database; The working node that meets the response conditions is used to extract data and perform detailed scheduling on the extraction tasks in the extraction time period.

8. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform a data extraction method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement a data extraction method according to any one of claims 1 to 6 when executed.

10. A computer program product, characterized in that The computer program product comprises a computer program, which implements a data extraction method according to any one of claims 1 to 6 when executed by a processor.