Data acquisition method and system

By splitting the acquisition tasks submitted by the user into multiple acquisition subtasks and using multi-thread execution, the automation and efficiency of data acquisition is achieved, and the problems of slow information acquisition speed and difficult to track error information in traditional methods are solved.

CN120045610APending Publication Date: 2025-05-27AISINO CORPORATION

Patent Information

Application Number
CN202411928221.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Traditional information collection methods are difficult to ensure timeliness and accuracy at the same time, and there are problems such as slow information acquisition speed and difficult to track error information.

Method used

A data acquisition method and system are proposed to obtain the acquisition tasks submitted by the user through the API interface, split them into multiple acquisition subtasks, and use multi-threads to execute the acquisition subtasks, monitor and analyze data in real time, and store them in the database.

Benefits of technology

The comprehensive automated data acquisition is realized, which greatly improves the efficiency of data acquisition. Users can quickly and accurately obtain target platform data, significantly enhance the efficiency of information acquisition, and visually display the status and performance of real-time monitoring of data acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045610A_ABST
    Figure CN120045610A_ABST
Patent Text Reader

Abstract

The invention discloses a data acquisition method and system, and the method comprises the steps: obtaining an acquisition task submitted by a user based on an API interface, and carrying out the task analysis of the acquisition task, so as to split the task into a plurality of acquisition subtasks; determining a corresponding task template based on the type of the acquisition sub-tasks, and distributing the plurality of acquisition sub-tasks to different acquisition message queues based on the determined task template; monitoring each acquisition message queue in real time by utilizing an acquisition program, acquiring an acquisition sub-task in each acquisition message queue by utilizing at least one thread, executing the acquisition sub-task by multiple threads to acquire original data, and writing the acquired original data into a summary queue; and consuming the original data by using a consumption program, analyzing the original data into structured format data according to a preset rule, and storing the structured format data into a database. According to the method, a single task is decomposed into a plurality of subtasks and concurrent execution is realized, so that the data acquisition efficiency is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data acquisition, and more particularly, to a data acquisition method and system. Background Art

[0002] As a comprehensive platform covering information queries in various regions across the country, the business platform has an increasing demand for users to frequently query and summarize information. Therefore, there is an urgent need for a general-purpose automation tool to quickly batch obtain information.

[0003] Traditional information acquisition methods often have difficulty in simultaneously ensuring timeliness and accuracy, and there are problems such as slow information acquisition speed and difficulty in tracking error information.

[0004] Therefore, a general data acquisition method is needed to achieve comprehensive automated data acquisition. Summary of the Invention

[0005] The present invention proposes a data acquisition method and system to solve the problem of how to efficiently implement data acquisition.

[0006] To solve the above problems, according to one aspect of the present invention, a data acquisition method is provided, and the method includes:

[0007] Obtaining a collection task submitted by a user based on an API interface, parsing the collection task to split the task into multiple collection subtasks, and establishing an association relationship between the collection subtasks and the corresponding collection tasks;

[0008] Determining a corresponding task template based on the type of the collection subtask, and distributing the multiple collection subtasks to different collection message queues based on the determined task template;

[0009] Using a collection program to monitor each collection message queue in real time, using at least one thread to obtain the collection subtasks in each collection message queue, executing the collection subtasks in multiple threads to collect raw data, and writing the collected raw data into a summary queue;

[0010] Using a consumption program to consume the raw data, and parsing the raw data into structured format data according to a preset rule and storing it in a database.

[0011] Preferably, the method further includes:

[0012] Obtaining the current timestamp, and performing duplicate removal processing on the collection subtasks in the collection message queue based on the timestamp.

[0013] Preferably, the method further includes:

[0014] Determine whether the collected message queue is empty. If it is not empty, sequentially obtain the task information of the collection subtasks in the collected message queue; and

[0015] Maintain the status of the collection task, and after executing the collection task, update the status of the collection task.

[0016] Preferably, the method further includes:

[0017] Perform real-time monitoring on the data collection task, record the execution information and task status of the collection subtasks, calculate the collection success rate according to the task status information, and display the execution information, task status and collection success rate in real time;

[0018] The execution information includes: the identifier of the data collection subtask, the name of the network device, the network device IP, the traffic port of the network device, the sampling period of the data collection subtask, the most recent collection time, and the status identifier of the collection process.

[0019] Preferably, the method further includes:

[0020] Based on the association relationship between the collection subtask and the corresponding collection task, associate the original data corresponding to the collection subtask with the collection task, and synchronize them to the storage bucket in real time for tracing the original data by querying the storage bucket.

[0021] According to another aspect of the present invention, a data collection system is provided, and the system includes:

[0022] A task parsing unit, configured to obtain the collection task submitted by the user based on the API interface, parse the collection task to split the task into multiple collection subtasks, and establish an association relationship between the collection subtasks and the corresponding collection tasks;

[0023] A task distribution unit, configured to determine the corresponding task template based on the type of the collection subtask, and distribute multiple collection subtasks to different collection message queues based on the determined task template;

[0024] A data collection unit, configured to use the collection program to monitor each collection message queue in real time, use at least one thread to obtain the collection subtasks in each collection message queue, execute the collection subtasks in multiple threads to collect the original data, and write the collected original data into the summary queue;

[0025] A data storage unit, configured to consume the original data using the consumption program, and parse the original data into structured format data according to the preset rules and store it in the database.

[0026] Preferably, the task distribution unit is further configured to:

[0027] Obtain the current timestamp, and perform deduplication processing on the collection subtasks in the collection message queue based on the timestamp.

[0028] Preferably, the system further includes: an update unit, configured to:

[0029] Determine whether the collection message queue is empty. If it is not empty, sequentially obtain the task information of the collection subtasks in the collection message queue; and

[0030] Maintain the status of the collection task, and update the status of the collection task after executing the collection task.

[0031] Preferably, the system further includes:

[0032] A monitoring unit, configured to perform real-time monitoring on the data collection task, record the execution information and task status of the collection subtasks, calculate the collection success rate according to the task status information, and perform real-time display on the execution information, task status, and collection success rate;

[0033] Wherein, the execution information includes: the identifier of the data collection subtask, the name of the network device, the network device IP, the traffic port of the network device, the sampling period of the data collection subtask, the most recent collection time, and the status identifier of the collection process.

[0034] Preferably, the system further includes:

[0035] A storage unit, configured to associate the original data corresponding to the collection subtask with the collection task based on the association relationship between the collection subtask and the corresponding collection task, and synchronize them to the storage bucket in real time for tracing the original data by querying the storage bucket.

[0036] The present invention provides a data collection method and system, including: obtaining a collection task submitted by a user based on an API interface, parsing the collection task to split the task into multiple collection subtasks, and establishing an association relationship between the collection subtasks and the corresponding collection task; determining a corresponding task template based on the type of the collection subtask, and distributing the multiple collection subtasks to different collection message queues based on the determined task template; using a collection program to monitor each collection message queue in real time, using at least one thread to obtain the collection subtasks in each collection message queue, executing the collection subtasks in multiple threads to collect raw data, and writing the collected raw data into a summary queue; using a consumption program to consume the raw data, and parsing the raw data into structured format data according to a preset rule and storing it in a database. By decomposing a single task into multiple subtasks and implementing concurrent execution, the present invention greatly improves the efficiency of data collection; users can quickly and accurately obtain data of a target platform, significantly enhancing the efficiency of information collection; the visual display of task collection records and interface success rate statistics enables users to monitor the status and performance of data collection in real time. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The exemplary embodiments of the present invention can be more fully understood by referring to the following drawings:

[0038] Figure 1 FIG. 100 is a flowchart of a data collection method according to an embodiment of the present invention;

[0039] Figure 2 FIG. 101 is a schematic diagram of a data collection process according to an embodiment of the present invention;

[0040] Figure 3 FIG. 300 is a schematic structural diagram of a data collection system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] Now, exemplary embodiments of the present invention will be described with reference to the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. These embodiments are provided to disclose the present invention in detail and completely, and to fully convey the scope of the present invention to those skilled in the art. The terms in the exemplary embodiments shown in the drawings are not intended to limit the present invention. In the drawings, the same unit / element is denoted by the same reference numeral.

[0042] Unless otherwise specified, the terms (including scientific and technical terms) used herein have the ordinary meaning understood by those skilled in the art. In addition, it can be understood that the terms defined in a commonly used dictionary should be understood as having a meaning consistent with the context of their related fields, and should not be understood as having an idealized or overly formal meaning.

[0043] Figure 1 This is a flowchart of a data collection method 100 according to an embodiment of the present invention. As Figure 1 shown, the data collection method provided by the embodiment of the present invention decomposes a single task into multiple subtasks and realizes concurrent execution, greatly improving the efficiency of data collection; users can quickly and accurately obtain data of the target platform, significantly enhancing the efficiency of information collection; the visual display of task collection records and interface success rate statistics enables users to monitor the status and performance of data collection in real time. The data collection method 100 provided by the embodiment of the present invention starts from step 101. At step 101, a collection task submitted by a user is obtained based on an API interface, and the collection task is parsed to split the task into multiple collection subtasks, and an association relationship between the collection subtasks and the corresponding collection tasks is established.

[0044] Preferably, the method further includes:

[0045] Obtaining the current timestamp and performing duplicate removal processing on the collection subtasks in the collection message queue based on the timestamp.

[0046] At step 102, a corresponding task template is determined based on the type of the collection subtask, and multiple collection subtasks are distributed to different collection message queues based on the determined task template.

[0047] At step 103, each collection message queue is monitored in real time by a collection program, at least one thread is used to obtain the collection subtasks in each collection message queue, the collection subtasks are executed by multiple threads to collect raw data, and the collected raw data is written into a summary queue.

[0048] At step 104, a consumption program is used to consume the raw data, and the raw data is parsed into structured format data according to a preset rule and then stored in a database.

[0049] Preferably, the method further includes:

[0050] Judging whether the collection message queue is empty. If it is not empty, the task information of the collection subtasks in the collection message queue is obtained in sequence; and

[0051] Maintaining the status of the collection task and updating the status of the collection task after the collection task is executed.

[0052] Preferably, the method further includes:

[0053] Monitor the data collection task in real time, record the execution information and task status of the collection subtask, calculate the collection success rate according to the task status information, and display the execution information, task status and collection success rate in real time;

[0054] Among them, the execution information includes: the identifier of the data collection subtask, the name of the network device, the network device IP, the traffic port of the network device, the sampling period of the data collection subtask, the most recent collection time, and the status identifier of the collection process.

[0055] Preferably, the method further includes:

[0056] Based on the association relationship between the collection subtask and the corresponding collection task, associate the original data corresponding to the collection subtask with the collection task, and synchronize it to the storage bucket in real time for tracing the original data by querying the storage bucket.

[0057] The present invention provides a complete set of distributed collection methods, including: task collection, task deduplication, multi-node distributed processing, original data aggregation, data parsing, original data storage, and data visualization. As shown in Figure 2 In the present invention, the task collection service is responsible for receiving the tasks submitted by users and refining and splitting the tasks, distributing the tasks to different processing queues through a preset template, and performing deduplication processing based on the current time. The collection program monitors each queue in real time, executes tasks to collect original data, and writes it into the aggregation queue. The aggregation queue supports multiple consumers to ensure that the original data is synchronously stored in the storage bucket in real time, and the data is parsed into a structured format according to the template and stored in the database. If there are data problems or questions during use, the storage bucket can be directly queried to achieve data traceability.

[0058] Specifically, in the present invention, the data collection process includes: the user submits a collection task based on the API interface, and the task collection service performs task refinement and splitting on the collection task to split the task into multiple collection subtasks, and distributes the collection subtasks to different message processing queues based on the configuration template; the collection program listens to different message processing queues through the adaptation template and executes tasks to collect original data; writes the collected original data into the aggregation queue; uses the consumption program to consume the original data, parses the original data into structured format data according to the preset rules and stores it in the database, and views the data collection situation on the visualization platform. In addition, the collected original data is synchronously stored in the storage bucket for easy traceability.

[0059] The method of the present invention greatly improves the efficiency of data collection through refined task splitting and a distributed multi-node processing mechanism. If there are any doubts about the collected data, it can be traced back to the corresponding original data immediately, eliminating the complex steps of checking one by one on the target platform. The task status and success rate during the collection process are recorded in detail and clearly displayed through a visualization platform, enabling users to intuitively grasp the success rate of the task and the collection progress.

[0060] Figure 3 FIG. 4 is a schematic structural diagram of a data collection system 300 according to an embodiment of the present invention. As Figure 3 shown, the data collection system 300 provided by the embodiment of the present invention includes: a task parsing unit 301, a task distribution unit 302, a data collection unit 303, and a data storage unit 304.

[0061] Preferably, the task parsing unit 301 is configured to obtain a collection task submitted by a user based on an API interface, parse the collection task, split the task into multiple collection subtasks, and establish an association relationship between the collection subtasks and the corresponding collection task.

[0062] Preferably, the task distribution unit 302 is configured to determine a corresponding task template based on the type of the collection subtask, and distribute the multiple collection subtasks to different collection message queues based on the determined task template.

[0063] Preferably, the task distribution unit 302 is further configured to:

[0064] Obtain the current timestamp, and perform duplicate removal processing on the collection subtasks in the collection message queue based on the timestamp.

[0065] Preferably, the data collection unit 303 is configured to use a collection program to monitor each collection message queue in real time, use at least one thread to obtain the collection subtasks in each collection message queue, execute the collection subtasks in multiple threads to collect the original data, and write the collected original data into a summary queue.

[0066] Preferably, the data storage unit 304 is configured to consume the original data using a consumption program, and parse the original data into structured format data according to a preset rule and store it in a database.

[0067] Preferably, the system further includes: an update unit, configured to:

[0068] Determine whether the collection message queue is empty. If it is not empty, sequentially obtain the task information of the collection subtasks in the collection message queue; and

[0069] Maintain the status of the collection task, and update the status of the collection task after the execution of the collection task.

[0070] Preferably, the system further includes:

[0071] A monitoring unit, configured to perform real-time monitoring on the data collection task, record the execution information and task status of the collection subtask, calculate the collection success rate according to the task status information, and perform real-time display on the execution information, task status, and collection success rate;

[0072] Wherein, the execution information includes: the identifier of the data collection subtask, the name of the network device, the network device IP, the traffic port of the network device, the sampling period of the data collection subtask, the most recent collection time, and the status identifier of the collection process.

[0073] Preferably, the system further includes:

[0074] A storage unit, configured to associate the original data corresponding to the collection subtask with the collection task based on the association relationship between the collection subtask and the corresponding collection task, and synchronize it to the storage bucket in real time for tracing the original data by querying the storage bucket.

[0075] The data collection system 300 of the embodiment of the present invention corresponds to the data collection method 100 of another embodiment of the present invention, and will not be elaborated here.

[0076] The present invention has been described by referring to a few embodiments. However, as is well known to those skilled in the art, as defined by the appended patent claims, other embodiments equivalent to those disclosed above of the present invention equally fall within the scope of the present invention.

[0077] Generally, all terms used in the claims are interpreted according to their ordinary meanings in the technical field, unless otherwise clearly defined therein. All references to "a / the [device, component, etc.]" are to be interpreted openly as at least one instance of the device, component, etc., unless otherwise clearly stated. The steps of any method disclosed herein need not be run in the exact order disclosed, unless clearly stated.

[0078] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0079] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0080] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0081] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operating steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the specific implementation manners of the present invention or make equivalent replacements. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. A data collection method, characterized in that: The method comprises: Acquire the collection task submitted by the user based on the API interface, perform task analysis on the collection task to split the task into multiple collection subtasks, and establish an association relationship between the collection subtasks and the corresponding collection task; Determine a corresponding task template based on the type of the acquisition subtask, and distribute the multiple acquisition subtasks to different acquisition message queues based on the determined task template; Using a collection program to monitor each collection message queue in real time, using at least one thread to obtain a collection subtask in each collection message queue, executing the collection subtask in multiple threads to collect raw data, and writing the collected raw data into a summary queue; The raw data is consumed by a consumption program, parsed into structured format data according to preset rules, and then stored in a database.

2. The method according to claim 1, characterized in that The method further comprises: A current timestamp is obtained, and based on the timestamp, duplicate collection subtasks in the collection message queue are deduplicated.

3. The method according to claim 1, characterized in that The method further comprises: Determine whether the collection message queue is empty, and if not, sequentially obtain task information of the collection subtasks in the collection message queue; and Maintain the status of the collection task, and update the status of the collection task after executing the collection task.

4. The method according to claim 1, characterized in that: The method further comprises: Monitor data collection tasks in real time, record the execution information and task status of collection subtasks, calculate the collection success rate based on the task status information, and display the execution information, task status and collection success rate in real time; The execution information includes: the identifier of the data collection subtask, the name of the network device, the IP address of the network device, the flow port of the network device, the sampling period of the data collection subtask, the most recent collection time and the status identifier of the collection process.

5. The method according to claim 1, characterized in that The method further comprises: Based on the association relationship between the collection subtask and the corresponding collection task, the original data corresponding to the collection subtask is associated with the collection task and synchronized to the storage bucket in real time, so as to trace the original data by querying the storage bucket.

6. A data acquisition system, characterized in that: The system comprises: A task parsing unit, used to obtain the collection task submitted by the user based on the API interface, perform task parsing on the collection task to split the task into multiple collection subtasks, and establish an association relationship between the collection subtasks and the corresponding collection task; A task distribution unit, used to determine a corresponding task template based on the type of the acquisition subtask, and distribute multiple acquisition subtasks to different acquisition message queues based on the determined task template; A data collection unit, used to monitor each collection message queue in real time using a collection program, obtain collection subtasks in each collection message queue using at least one thread, execute the collection subtasks in multiple threads to collect raw data, and write the collected raw data into a summary queue; The data storage unit is used to consume the original data using a consumption program, parse the original data into structured format data according to preset rules, and then store it in a database.

7. The system according to claim 6, characterized in that The task distribution unit is further used for: A current timestamp is obtained, and based on the timestamp, duplicate collection subtasks in the collection message queue are deduplicated.

8. The system according to claim 6, characterized in that The system further comprises an updating unit, configured to: Determine whether the collection message queue is empty, and if not, sequentially obtain task information of the collection subtasks in the collection message queue; and Maintain the status of the collection task, and update the status of the collection task after executing the collection task.

9. The system according to claim 6, characterized in that The system further comprises: The monitoring unit is used to monitor the data collection task in real time, record the execution information and task status of the collection subtask, calculate the collection success rate according to the task status information, and display the execution information, task status and collection success rate in real time; The execution information includes: the identifier of the data collection subtask, the name of the network device, the IP address of the network device, the flow port of the network device, the sampling period of the data collection subtask, the most recent collection time and the status identifier of the collection process.

10. The system according to claim 6, characterized in that The system further comprises: The storage unit is used to associate the original data corresponding to the collection subtask with the collection task based on the association relationship between the collection subtask and the corresponding collection task, and synchronize them to the storage bucket in real time, so as to trace the original data by querying the storage bucket.

Citation Information

Patent Citations

  • Server cluster performance sampling analysis method and device and electronic equipment

    CN110928732A

  • Data acquisition method and related equipment

    CN113760919A

  • Method and device for automatically generating traceability information in ETL, and electronic equipment

    CN116701500A

Cited By

  • Cross-border e-commerce settlement declaration processing method, electronic equipment and computer readable medium

    CN120430788A