A data collection method, device, equipment and storage medium
By using daemon threads in the data acquisition system to obtain the main thread state and initialize the thread pool, and assign tasks based on node information, the data acquisition efficiency is improved and the processing capability is expanded, solving the problem of low data acquisition efficiency and the bottleneck of the main node in the existing technology.
Patent Information
- Application Number
- CN202111375705.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-19
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2041-11-19
AI Technical Summary
In the prior art, data acquisition efficiency is low, processing capability is not scalable, server failure recovery is not timely enough, and the main node is a centralized node, the main node server failure is unavailable, and the main node easily becomes a bottleneck in processing efficiency.
The main thread state information is obtained through the daemon thread. If the main thread state is in the open state, the thread pool of the worker thread is initialized. According to the node information, the target task is collected through the main thread, and the task is submitted to the target worker thread, and data collection is collected through the worker thread.
It realizes balancing load based on each node's own processing capabilities, and each node processes data acquisition tasks in parallel, improving data acquisition efficiency, solving the problem of unscalable processing capabilities and untimely server failure recovery, and avoiding the risk of master nodes becoming bottlenecks.
Smart Images

Figure CN114020472B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of computer technology, and in particular to a data acquisition method, device, equipment and storage medium. Background Art
[0002] With the rapid development of digitalization in the financial industry, the business system architecture of commercial banks will become more and more complex. Various business systems will generate a large amount of business data every day. Fully mining and utilizing business data assets can help commercial banks gain competitive advantages. How to efficiently, timely and stably collect business data is the basis and prerequisite for commercial banks to fully utilize data assets to create business value. There are generally two ways of conventional data collection: 1. Point-to-point single-node data collection method, the server uses HA method to achieve high availability; 2. A distributed method using a master-slave structure, the master node is responsible for task scheduling and allocation, and the slave node performs specific data processing tasks.
[0003] The point-to-point single data collection method mainly has the following problems: low data collection efficiency, non-scalable processing capacity, and untimely server failure recovery.
[0004] The main problems of the master-slave distributed approach are as follows: the master node is a centralized node, and if the master node server fails, the entire system becomes unavailable. The master node can easily become a bottleneck in processing efficiency. Summary of the invention
[0005] The embodiments of the present invention provide a data collection method, device, equipment and storage medium to achieve load balancing to each node according to the processing capacity of each node, and each node processes data collection tasks in parallel to improve data collection efficiency.
[0006] In a first aspect, an embodiment of the present invention provides a data collection method, including:
[0007] Get the main thread status information through the daemon thread;
[0008] If the main thread status information is in an open state, the thread pool of the working thread is initialized through the main thread;
[0009] According to the node information, the target task is obtained through the main thread, and the target task is submitted to the target working thread, and data collection is performed through the target working thread.
[0010] Further, the thread pool includes: a first number of worker threads;
[0011] Correspondingly, according to the node information, the target task is obtained through the main thread, and the target task is submitted to the target working thread, and data is collected through the target working thread, including:
[0012] Obtain node status information and a first number of working threads in a thread pool;
[0013] If the node status information is an online state, obtaining a second number of target working threads in a running state;
[0014] If the second number is less than the first number, the target task is obtained through the main thread, and the target task is submitted to the target working thread;
[0015] Data collection is performed through the target working thread.
[0016] Further, the target task is obtained through the main thread, and the target task is submitted to the target working thread, including:
[0017] Obtain the status information of the task lock in the database through the main thread;
[0018] If the status information of the task lock is unoccupied, updating the status information of the task lock to occupied;
[0019] Obtain the target task from the task list in the database through the main thread, or, if the task list in the database is empty, obtain the target task corresponding to the node in the offline state;
[0020] Update the status information of the task lock to an unoccupied state;
[0021] The received target task is submitted to the thread pool through the main thread, so that the thread pool submits the target task to the target working thread.
[0022] Furthermore, it also includes:
[0023] Get the check-in interval in the database;
[0024] Get the status information of the check-in processing lock in the database;
[0025] If the state information of the sign-in processing lock is an unoccupied state, the state information of the sign-in processing lock is updated to an occupied state, and the sign-in information is inserted into the sign-in table in the database; or, the sign-in information in the sign-in table in the database is updated;
[0026] Get the check-in information of other nodes;
[0027] Determine the status information of other nodes according to the check-in information of the other nodes;
[0028] The sign-in table is updated according to the status information of the other nodes, and the status information of the sign-in processing lock is updated to an unoccupied state.
[0029] Further, the check-in information of the other nodes includes: the last check-in time, check-in time and buffer time of the other nodes;
[0030] Correspondingly, determining the status information of other nodes according to the check-in information of the other nodes includes:
[0031] If the sum of the last check-in time, the check-in interval and the buffer time of the other nodes is less than the current system time, the status information of the other nodes is determined to be offline, wherein the buffer time is equal to N times the check-in time, wherein N is a positive integer.
[0032] Further, submitting the target task to a target work thread, and collecting data through the target work thread, includes:
[0033] Submitting the target task to the target working thread, wherein the target task carries task information;
[0034] Determine the stage information and target parameters of the target task according to the task information;
[0035] Data collection is performed according to the stage information and target parameters of the target task.
[0036] Furthermore, data collection is performed according to the stage information and target parameters of the target task, including:
[0037] If the stage information of the target task is the data extraction stage, data is extracted according to the target parameters corresponding to the data extraction stage, and the extracted data is inserted into a temporary table;
[0038] If the stage information of the target task is a data checking stage, performing data checking according to the target parameters corresponding to the data checking stage;
[0039] If the stage information of the target task is the data storage stage, data storage is performed according to the target parameters corresponding to the data storage stage.
[0040] In a second aspect, an embodiment of the present invention further provides a data acquisition device, the device comprising:
[0041] The acquisition module is used to obtain the main thread status information through the daemon thread;
[0042] An initialization module, used for initializing a thread pool of working threads through the main thread if the main thread status information is in an open state;
[0043] The collection module is used to obtain the target task through the main thread according to the node information, submit the target task to the target working thread, and collect data through the target working thread.
[0044] In a third aspect, an embodiment of the present invention further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, a data acquisition method as described in any one of the embodiments of the present invention is implemented.
[0045] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the data collection method as described in any one of the embodiments of the present invention.
[0046] The embodiment of the present invention obtains the main thread status information through the daemon thread; if the main thread status information is in the open state, the thread pool of the working thread is initialized through the main thread; the target task is obtained through the main thread according to the node information, and the target task is submitted to the target working thread, and data is collected through the target working thread, which not only solves the problems of low data collection efficiency, unscalable processing capacity, and insufficient server failure recovery, but also solves the problem that the main node is a centralized node, the entire system is unavailable due to the failure of the main node server, and the main node easily becomes a bottleneck of processing efficiency. The load can be balanced to each node according to the processing capacity of each node itself, and each node processes data collection tasks in parallel, thereby improving data collection efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0048] Figure 1 is a flow chart of a data collection method in an embodiment of the present invention;
[0049] Figure 1a is a schematic diagram of a data acquisition system in an embodiment of the present invention;
[0050] Figure 1b It is a flowchart of node task collection and processing scheduling in an embodiment of the present invention;
[0051] Figure 1c is a flowchart of a node status maintenance method in an embodiment of the present invention;
[0052] Figure 1d is a flow chart of another data collection method in an embodiment of the present invention;
[0053] Figure 2is a structural schematic diagram of a data acquisition device in an embodiment of the present invention;
[0054] Figure 3 is a schematic structural diagram of an electronic device in an embodiment of the present invention;
[0055] Figure 4 It is a schematic diagram of the structure of a computer-readable storage medium containing a computer program in an embodiment of the present invention. DETAILED DESCRIPTION
[0056] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only the parts related to the present invention, rather than all structures, are shown in the accompanying drawings. In addition, the embodiments of the present invention and the features in the embodiments may be combined with each other without conflict.
[0057] It should be mentioned before discussing exemplary embodiments in more detail that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe various operations (or steps) as sequential processes, many operations therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of various operations can be rearranged. The process can be terminated when its operation is completed, but can also have additional steps not included in the accompanying drawings. The process can correspond to methods, functions, procedures, subroutines, subprograms, etc. In addition, the embodiments in the present invention and the features in the embodiments can be combined with each other without conflict.
[0058] The term “including” and its variations used in the present invention are open inclusions, that is, “including but not limited to.” The term “based on” means “based at least in part on.” The term “one embodiment” means “at least one embodiment.”
[0059] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0060] Figure 1 A flowchart of a data acquisition method provided by an embodiment of the present invention is provided. This embodiment is applicable to data acquisition situations. The method can be executed by a data acquisition device in an embodiment of the present invention. The device can be implemented in software and / or hardware. Figure 1 As shown, the method specifically comprises the following steps:
[0061] S110, obtaining the main thread status information through the daemon thread.
[0062] The main thread status information may be in an enabled state or in an disabled state, which is not limited in the embodiment of the present invention.
[0063] Specifically, the current node periodically checks the main thread status information through the daemon thread.
[0064] S120: If the main thread status information is in an open state, the thread pool of the working thread is initialized through the main thread.
[0065] The method of initializing the thread pool of working threads through the main thread may be: constructing a thread pool of working threads corresponding to the maximum number of threads according to a preset maximum number of threads.
[0066] The thread pool of working threads includes at least two working threads.
[0067] Specifically, if the main thread status information is obtained through the daemon thread and is in an open state, a thread pool of working threads corresponding to the maximum number of threads is constructed according to a preset maximum number of threads.
[0068] S130, obtaining a target task through the main thread according to the node information, submitting the target task to a target working thread, and performing data collection through the target working thread.
[0069] The node information is current node information, and the node information may be obtained by obtaining the status of the current node in the sign-in table. The node information may be an online status or an offline status, which is not limited in the embodiment of the present invention.
[0070] Among them, the method of obtaining the target task through the main thread can be: if the node information is in an online state, the target task is determined according to the second data of the target working thread in the running state and the first number of working threads in the thread pool; the method of obtaining the target task through the main thread can also be: if the node information is in an online state, the first task set is determined according to the second data of the target working thread in the running state and the first number of working threads in the thread pool, and the target task with the most recent triggering time in the first task set is obtained; the method of obtaining the target task through the main thread can also be: if the node information is in an online state, the first task set is determined according to the second data of the target working thread in the running state and the first number of working threads in the thread pool, the first task set includes: a first subset, a second subset, a third subset and a fourth subset, the priorities of the first subset, the second subset, the third subset and the fourth subset are related to the attributes of the subsets, the first subset with the highest priority is selected, and the target task with the most recent triggering time in the first subset is obtained.
[0071] Specifically, the method of obtaining the target task through the main thread and submitting the target task to the target working thread can be: obtaining the status information of the task lock in the database through the main thread; if the status information of the task lock is unoccupied, updating the status information of the task lock to occupied; obtaining the target task in the task list in the database through the main thread, or if the task list in the database is empty, obtaining the target task corresponding to the node in the offline state; updating the status information of the task lock to unoccupied; submitting the obtained target task to the thread pool through the main thread, so that the thread pool submits the target task to the target working thread.
[0072] Among them, the target task is submitted to the target working thread, and the method of performing data collection through the target working thread can be: submitting the received target task to the thread pool through the main thread, and the thread pool submits the target task to the target working thread.
[0073] Among them, the method of collecting data through the target working thread can be: submitting the target task to the target working thread, wherein the target task carries task information; determining the stage information and target parameters of the target task according to the task information; and collecting data according to the stage information and target parameters of the target task. The method of collecting data through the target working thread can also be: submitting the target task to the target working thread, wherein the target task carries task information; determining the stage information and target parameters of the target task according to the task information; if the stage information of the target task is the data extraction stage, extracting data according to the target parameters corresponding to the data extraction stage, and inserting the extracted data into a temporary table; if the stage information of the target task is the data inspection stage, data inspection is performed according to the target parameters corresponding to the data inspection stage; if the stage information of the target task is the data warehousing stage, data warehousing is performed according to the target parameters corresponding to the data warehousing stage.
[0074] Specifically, Figure 1a As shown, an embodiment of the present invention provides a data collection system, which includes: a node cluster composed of multiple nodes, a database and a business system, and the node collects data from the corresponding business system according to the target task. The node cluster is connected to the commercial bank business system, and the data collection target is achieved through three mechanisms: node active state self-maintenance, distributed self-organizing node task collection and processing scheduling, and data collection processing task processing.
[0075] Optionally, if the main thread status information is in an unactivated state, the main thread is activated through a daemon thread.
[0076] Optionally, the thread pool includes: a first number of working threads;
[0077] Correspondingly, according to the node information, the target task is obtained through the main thread, and the target task is submitted to the target working thread, and data is collected through the target working thread, including:
[0078] Obtain node status information and a first number of working threads in a thread pool;
[0079] If the node status information is an online state, obtaining a second number of target working threads in a running state;
[0080] If the second number is less than the first number, the target task is obtained through the main thread, and the target task is submitted to the target working thread;
[0081] Data collection is performed through the target working thread.
[0082] The first number is a preset maximum number of threads. When the thread pool of working threads is initialized by the main thread, the thread pool of working threads is constructed according to the preset maximum number of threads.
[0083] Wherein, if the node status information is an online state, the method for obtaining the second number of target working threads in the running state may be: if the node status information is an online state, the second number of target working threads in the running state is obtained through the main thread. If the node status information is an online state, the method for obtaining the second number of target working threads in the running state may also be: if the node status information is an online state, the number of threads in the working thread list is obtained through the main thread, wherein the working thread list stores working threads in the working state.
[0084] Specifically, if the second number is less than the first number, the target task is obtained through the main thread and submitted to the target working thread. If the second number is greater than or equal to the first number, the status of the working thread continues to be obtained through the main thread. If the status of the working thread is idle (indicating that the task previously executed by the working thread has ended), the working thread in the idle state will be deleted from the working thread list.
[0085] Optionally, receiving the target task through the main thread and submitting the target task to the target working thread includes:
[0086] Obtain the status information of the task lock in the database through the main thread;
[0087] If the status information of the task lock is unoccupied, updating the status information of the task lock to occupied;
[0088] Obtain the target task from the task list in the database through the main thread, or, if the task list in the database is empty, obtain the target task corresponding to the node in the offline state;
[0089] Update the status information of the task lock to an unoccupied state;
[0090] The received target task is submitted to the thread pool through the main thread, so that the thread pool submits the target task to the target working thread.
[0091] The task list includes: unprocessed tasks and failed tasks.
[0092] Among them, the target task corresponding to the node in the offline state is the task received when the node is in the online state, the task has not been processed and completed, the node is offline due to node failure or other reasons, and the status of the task is a task that failed to be processed.
[0093] Specifically, after the target task in the task list in the database is obtained through the main thread, the status of the target task in the task list is changed to being processed. After obtaining the target task corresponding to the node in the offline state, the target task is added to the task list, and the status of the target task is recorded as being processed at the same time, so as to facilitate monitoring of the target task.
[0094] Wherein, if the state information of the task claim lock is unoccupied, updating the state information of the task claim lock to occupied state is to obtain the task claim lock. Updating the state information of the task claim lock to unoccupied state is to release the task claim lock. The above method can prevent different nodes from claiming the same task. Different nodes can claim tasks in series, and after claiming tasks, different nodes can process tasks in parallel.
[0095] In a specific example, Figure 1bAs shown, the daemon thread periodically checks the status of the main thread, and attempts to start the main thread if the main thread has not started. If the main thread is started, the thread pool is initialized, and the maximum number of threads is pre-set. Check whether the status of the current node in the sign-in table is online. If not, continue to check. If so, continue to execute the following steps: If the number of currently running working threads is less than the pre-set maximum number of threads, it means that the task can be executed. The main thread attempts to obtain the status information of the task lock. If other processing nodes are currently receiving tasks (the status information of the task lock is occupied), the current node waits for the lock to ensure the concurrent consistency of the node processing task collection. If the status information of the task lock is unoccupied, the status information of the task lock is updated to occupied, that is, the task lock is obtained, and the executable task set S is obtained from the task list of the day through the main thread. The executable task set contains 4 scenario subsets S={t 1 ,t 2 ,t 3 ,t 4}, where t 1 It is a normal pending task, that is, the task processing time has expired and the task status is a pending task set; t 2 The task that needs to failover is the task set whose processing status is in progress but the registered processing node status is offline. 3 The task is lost accidentally, that is, the task status is in processing and the task processing node is the current node, but the task is not in the current working thread; 4 To process failed tasks, that is, the task processing status is failed. The main thread takes a task from the executable task and submits it to the worker thread pool and records the worker thread. The task collection limit level is t 1 >t 2 >t 3 >t 4 Through the above steps, the processing capacity expansion of dynamic node addition and the failover of offline servers are realized. Update the task status (the processing node is the current node name, and the processing status is processing), and release the task processing lock. If the currently running worker threads are greater than or equal to the maximum number of worker threads, poll the worker thread list, delete the terminated worker threads, and update the number of running worker threads.
[0096] Specifically, each original data table corresponds to a collection task. After the collection batch is started, the executable collection tasks are loaded through the thread pool. The daemon thread is responsible for pulling up the main thread, and the main thread starts the worker thread to execute the specific collection task. Each node has only one active data collection main thread, but can have multiple worker threads.
[0097] Optionally, also include:
[0098] Get the check-in interval in the database;
[0099] Get the status information of the check-in processing lock in the database;
[0100] If the state information of the sign-in processing lock is an unoccupied state, the state information of the sign-in processing lock is updated to an occupied state, and the sign-in information is inserted into the sign-in table in the database; or, the sign-in information in the sign-in table in the database is updated;
[0101] Get the check-in information of other nodes;
[0102] Determine the status information of other nodes according to the check-in information of the other nodes;
[0103] The sign-in table is updated according to the status information of the other nodes, and the status information of the sign-in processing lock is updated to an unoccupied state.
[0104] The check-in interval is a pre-determined check-in interval, and the way to obtain the check-in interval in the database may be: obtaining the check-in interval in the database in the target server, that is, the database may be other servers except the node.
[0105] The sign-in information includes: node status, such as online status or offline status.
[0106] The sign-in table stores the node identifier, the node status corresponding to the node identifier, and the sign-in time.
[0107] Optionally, the check-in information of the other nodes includes: the last check-in time, check-in time and buffer time of the other nodes;
[0108] Correspondingly, determining the status information of other nodes according to the check-in information of the other nodes includes:
[0109] If the sum of the last check-in time, the check-in interval and the buffer time of the other nodes is less than the current system time, the status information of the other nodes is determined to be offline, wherein the buffer time is equal to N times the check-in time, wherein N is a positive integer.
[0110] The buffer time is generally double the buffer time.
[0111] In a specific example, Figure 1cAs shown, Step 1: The check-in interval is a system parameter, and each node checks in once every check-in interval. Step 2: Before checking in, each node needs to obtain the check-in processing lock first. If the lock is occupied by other nodes, it will wait. Step 3: When checking in, if there is no current node in the check-in table (in the case of a new node), insert the check-in record. If there is a check-in record, update the check-in time and node status. Step 4: After completing its own check-in, check the status of other nodes. If the following conditions are met: the last check-in time + check-in interval + buffer time < current system time, the current node is determined to be a node failure, and the node status is updated to offline. Among them, the buffer time is an integer multiple of the check-in interval, generally one times the check-in interval. Through the above steps, all nodes in the cluster jointly maintain the node status table, and each node can perceive all active nodes and inactive nodes in the cluster by checking the status table.
[0112] The self-maintenance of the active state of nodes is the premise for realizing task allocation and data collection. The technical solution provided by the embodiment of the present invention adopts the heartbeat mechanism of active node active check-in to realize the node state perception of decentralized nodes. Each node signs in in the database table at the agreed heartbeat interval. When each node signs in, it checks whether other nodes have overdue sign-in. If there is overdue sign-in, the node is judged to be faulty and its node state is updated.
[0113] It should be noted that batch servers are peer nodes to each other. If an individual node goes down, the tasks it is responsible for processing need to be transferred to normal nodes to ensure the business integrity of data collection. After the down node is restored, the unprocessed tasks can be reallocated to the node. Nodes can be dynamically added and removed. Newly added nodes can participate in data collection and processing at any time, and the tasks of the exited nodes can be taken over by the running nodes.
[0114] Optionally, submitting the target task to a target working thread, and performing data collection through the target working thread, includes:
[0115] Submitting the target task to the target working thread, wherein the target task carries task information;
[0116] Determine the stage information and target parameters of the target task according to the task information;
[0117] Data collection is performed according to the stage information and target parameters of the target task.
[0118] The stage information of the target task may be: data extraction stage, data checking stage, or data storage stage. It should be noted that the target parameters corresponding to the target task at different stages are different.
[0119] Optionally, data collection is performed according to the stage information and target parameters of the target task, including:
[0120] If the stage information of the target task is the data extraction stage, data is extracted according to the target parameters corresponding to the data extraction stage, and the extracted data is inserted into a temporary table;
[0121] If the stage information of the target task is a data checking stage, performing data checking according to the target parameters corresponding to the data checking stage;
[0122] If the stage information of the target task is the data storage stage, data storage is performed according to the target parameters corresponding to the data storage stage.
[0123] Specifically, if the stage information of the target task is the data extraction stage, data extraction is performed according to the target parameters corresponding to the data extraction stage, and the extracted data is inserted into a temporary table. The method can be: If the stage information of the target task is the data extraction stage, data extraction is performed according to the system identifier, table identifier and business data field, and the extracted data is inserted into a temporary table.
[0124] Specifically, if the stage information of the target task is the data checking stage, the method of performing data checking according to the target parameters corresponding to the data checking stage may be: if the stage information of the target task is the data checking stage, then a data check is performed on the data in the temporary table according to the target parameters corresponding to the data checking stage. For example, a data check may be performed on the data in the temporary table according to data standard parameters.
[0125] Specifically, if the stage information of the target task is the data warehousing stage, the method of performing data warehousing according to the target parameters corresponding to the data warehousing stage may be: performing warehousing processing on the data in the checked temporary table according to the storage location information.
[0126] In a specific example, Figure 1d As shown, after a node receives the target task through the main thread, it submits the target task to the working thread for specific data collection and processing. A data collection task is divided into three stages: data extraction, data inspection, and data storage. Step 1: The working thread checks the stage of the target task and runs the corresponding processing steps from the current stage. Step 2: When the task is in the data extraction stage, the working thread automatically generates data query statements based on the data table structure mapping relationship parameters to query data from the source system and insert it into the temporary table. Step 3: According to the data standards specified by the commercial bank, the data quality of the temporary table data is checked. Step 4: After the data quality check, the data is stored.
[0127] The embodiment of the present invention realizes an extensible peer-to-peer distributed data acquisition system. Compared with the traditional server deployment method, the embodiment of the present invention realizes a multi-active deployment method, with more timely fault switching and stronger parallel processing capabilities. Compared with the traditional master-slave distributed processing method, the embodiment of the present invention realizes a decentralized peer-to-peer deployment method, reducing the performance bottleneck risk of the master node under the master-slave distributed structure and the risk of failure of the master node. In addition, the embodiment of the present invention also realizes dynamic horizontal expansion capability, and the entire system can add new nodes to expand processing capabilities without shutting down.
[0128] The technical solution of this embodiment obtains the main thread status information through the daemon thread; if the main thread status information is in the open state, the thread pool of the working thread is initialized through the main thread; the target task is obtained through the main thread according to the node information, and the target task is submitted to the target working thread, and data is collected through the target working thread, which not only solves the problems of low data collection efficiency, unscalable processing capacity, and insufficient server failure recovery, but also solves the problem that the main node is a centralized node, the entire system is unavailable due to the failure of the main node server, and the main node easily becomes a bottleneck of processing efficiency. It can balance the load to each node according to the processing capacity of each node itself, and each node processes data collection tasks in parallel, thereby improving data collection efficiency.
[0129] Figure 2 This is a schematic diagram of the structure of a data acquisition device provided by an embodiment of the present invention. This embodiment is applicable to data acquisition. The device can be implemented in software and / or hardware. The device can be integrated in any device that provides data acquisition function, such as Figure 2 As shown, the data acquisition device specifically includes: an acquisition module 210 , an initialization module 220 and a collection module 230 .
[0130] Among them, the acquisition module is used to obtain the main thread status information through the daemon thread;
[0131] An initialization module, used for initializing a thread pool of working threads through the main thread if the main thread status information is in an open state;
[0132] The collection module is used to obtain the target task through the main thread according to the node information, submit the target task to the target working thread, and collect data through the target working thread.
[0133] The above-mentioned product can execute the method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0134] The technical solution of this embodiment obtains the main thread status information through the daemon thread; if the main thread status information is in the open state, the thread pool of the working thread is initialized through the main thread; the target task is obtained through the main thread according to the node information, and the target task is submitted to the target working thread, and data is collected through the target working thread, which not only solves the problems of low data collection efficiency, unscalable processing capacity, and insufficient server failure recovery, but also solves the problem that the main node is a centralized node, the entire system is unavailable due to the failure of the main node server, and the main node easily becomes a bottleneck of processing efficiency. It can balance the load to each node according to the processing capacity of each node itself, and each node processes data collection tasks in parallel, thereby improving data collection efficiency.
[0135] Figure 3 A schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. Figure 3 A block diagram of an electronic device 312 suitable for implementing embodiments of the present invention is shown. Figure 3 The electronic device 312 shown is only an example and should not limit the functions and scope of use of the embodiments of the present invention. The device 312 is a typical computing device for trajectory fitting functions.
[0136] like Figure 3 As shown, the electronic device 312 is in the form of a general purpose computing device. The components of the electronic device 312 may include, but are not limited to: one or more processors 316, a storage device 328, and a bus 318 connecting different system components (including the storage device 328 and the processor 316).
[0137] Bus 318 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor or a local bus using any of a variety of bus architectures. For example, these architectures include but are not limited to Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus and Peripheral Component Interconnect (PCI) bus.
[0138] The electronic device 312 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 312, including volatile and non-volatile media, removable and non-removable media.
[0139] The storage device 328 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 330 and / or cache memory 332. The electronic device 312 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 334 may be used to read and write non-removable, non-volatile magnetic media ( Figure 3 not shown, usually called a "hard drive"). Although Figure 3 Not shown in the figure, a disk drive for reading and writing a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing a removable non-volatile optical disk (e.g., a read-only optical disk (Compact Disc-Read Only Memory, CD-ROM), a digital video disk (Digital Video Disc-Read Only Memory, DVD-ROM) or other optical media) may be provided. In these cases, each drive may be connected to the bus 318 via one or more data medium interfaces. The storage device 328 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present invention.
[0140] A program 336 having a set (at least one) of program modules 326 may be stored, for example, in a storage device 328, such program modules 326 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment. The program modules 326 generally perform the functions and / or methods of the embodiments described herein.
[0141] The electronic device 312 may also communicate with one or more external devices 314 (e.g., keyboard, pointing device, camera, display 324, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 312, and / or communicate with any device that enables the electronic device 312 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication may be performed through an input / output (I / O) interface 322. In addition, the electronic device 312 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN) and / or a public network, such as the Internet) through a network adapter 320. As shown, the network adapter 320 communicates with other modules of the electronic device 312 through a bus 318. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 312, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, disk arrays (Redundant Arrays of Independent Disks, RAID) systems, tape drives, and data backup storage systems.
[0142] The processor 316 executes various functional applications and data processing by running the programs stored in the storage device 328, for example, implementing the data collection method provided in the above embodiment of the present invention:
[0143] Get the main thread status information through the daemon thread;
[0144] If the main thread status information is in an open state, the thread pool of the working thread is initialized through the main thread;
[0145] According to the node information, the target task is obtained through the main thread, and the target task is submitted to the target working thread, and data collection is performed through the target working thread.
[0146] Figure 4 Schematic diagram of the structure of a computer-readable storage medium containing a computer program in an embodiment of the present invention. The embodiment of the present invention provides a computer-readable storage medium 61 on which a computer program 610 is stored. When the program is executed by one or more processors, the data collection method provided in all the invention embodiments of the present application is implemented:
[0147] Get the main thread status information through the daemon thread;
[0148] If the main thread status information is in an open state, the thread pool of the working thread is initialized through the main thread;
[0149] According to the node information, the target task is obtained through the main thread, and the target task is submitted to the target working thread, and data collection is performed through the target working thread.
[0150] Any combination of one or more computer-readable media can be used. Computer-readable media can be computer-readable signal media or computer-readable storage media or any combination of the above two. Computer-readable storage media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or devices, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: electrical connections with one or more wires, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this document, computer-readable storage media can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device.
[0151] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, which carry computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0152] The program code embodied on the computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0153] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (Hyper Text Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0154] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0155] Computer program code for performing the operation of the present invention may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0156] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0157] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a unit does not, in some cases, limit the unit itself.
[0158] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0159] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0160] Note that the above are only preferred embodiments of the present invention and the technical principles used. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present invention, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A data collection method, characterized in that: The data collection method comprises: Get the main thread status information through the daemon thread; If the main thread status information is in an open state, the thread pool of the working thread is initialized through the main thread; According to the node information, the target task is obtained through the main thread, and the target task is submitted to the target working thread, and data is collected through the target working thread; Get the check-in interval in the database; Get the status information of the check-in processing lock in the database; If the state information of the sign-in processing lock is an unoccupied state, the state information of the sign-in processing lock is updated to an occupied state, and the sign-in information is inserted into the sign-in table in the database; or, the sign-in information in the sign-in table in the database is updated; Get the check-in information of other nodes; Determine the status information of other nodes according to the check-in information of the other nodes; The sign-in table is updated according to the status information of the other nodes, and the status information of the sign-in processing lock is updated to an unoccupied state.
2. The method according to claim 1, characterized in that The thread pool includes: a first number of working threads; Correspondingly, according to the node information, the target task is obtained through the main thread, and the target task is submitted to the target working thread, and data is collected through the target working thread, including: Obtain node status information and a first number of working threads in a thread pool; If the node status information is an online state, obtaining a second number of target working threads in a running state; If the second number is less than the first number, the target task is obtained through the main thread, and the target task is submitted to the target working thread; Data collection is performed through the target working thread.
3. The method according to claim 2, characterized in that Receiving the target task through the main thread and submitting the target task to the target working thread includes: Obtain the status information of the task lock in the database through the main thread; If the status information of the task lock is unoccupied, updating the status information of the task lock to occupied; Obtain the target task from the task list in the database through the main thread, or, if the task list in the database is empty, obtain the target task corresponding to the node in the offline state; Update the status information of the task lock to an unoccupied state; The received target task is submitted to the thread pool through the main thread, so that the thread pool submits the target task to the target working thread.
4. The method according to claim 1, characterized in that: The check-in information of the other nodes includes: the last check-in time, check-in time and buffer time of the other nodes; Correspondingly, determining the status information of other nodes according to the check-in information of the other nodes includes: If the sum of the last check-in time, the check-in interval and the buffer time of the other nodes is less than the current system time, the status information of the other nodes is determined to be offline, wherein the buffer time is equal to N times the check-in time, wherein N is a positive integer.
5. The method according to claim 1, characterized in that Submitting the target task to a target working thread, and collecting data through the target working thread, including: Submitting the target task to the target working thread, wherein the target task carries task information; Determine the stage information and target parameters of the target task according to the task information; Data collection is performed according to the stage information and target parameters of the target task.
6. The method according to claim 5, characterized in that Data collection is performed according to the stage information and target parameters of the target task, including: If the stage information of the target task is the data extraction stage, data is extracted according to the target parameters corresponding to the data extraction stage, and the extracted data is inserted into a temporary table; If the stage information of the target task is a data checking stage, performing data checking according to the target parameters corresponding to the data checking stage; If the stage information of the target task is the data storage stage, data storage is performed according to the target parameters corresponding to the data storage stage.
7. A data acquisition device, characterized in that: The data acquisition device comprises: The acquisition module is used to obtain the main thread status information through the daemon thread; An initialization module, used for initializing a thread pool of working threads through the main thread if the main thread status information is in an open state; A collection module, used to obtain a target task through the main thread according to the node information, submit the target task to a target working thread, and collect data through the target working thread; The device is also used to obtain the check-in interval in the database; obtain the status information of the check-in processing lock in the database; if the status information of the check-in processing lock is an unoccupied state, update the status information of the check-in processing lock to an occupied state, and insert the check-in information into the check-in table in the database; or, update the check-in information in the check-in table in the database; obtain the check-in information of other nodes; determine the status information of other nodes according to the check-in information of the other nodes; update the check-in table according to the status information of the other nodes, and update the status information of the check-in processing lock to an unoccupied state.
8. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the processors are enabled to implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium containing a computer program, wherein the computer program is stored thereon, characterized in that: When the program is executed by one or more processors, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Task scheduling method, device and equipment, and medium
CN111858012A