A data acquisition method, device, apparatus and storage medium

CN115438056BActive Publication Date: 2026-09-29AGRICULTURAL BANK OF CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211062512.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-01
Publication Date
2026-09-29
Estimated Expiration
2042-09-01

AI Technical Summary

Benefits of technology

[0018]本发明实施例的技术方案,响应于目标数据获取请求,若待获取数据的数据获取方式为非实时获取方式,则根据目标数据获取请求,将目标数据获取请求对应的目标数据获取任务划分为至少两个子任务,并为每个子任务生成对应的备选数据获取请求,根据备选数据获取请求,并发控制至少两个工作线程基于对应的HashMap内锁,将各子任务对应的备选数据获取请求发送至对应的ES集群,获取至少两个备选数据,根据至少两个备选数据,确定目标获取数据。通过根据数据获取请求确定数据获取方式,进一步根据数据的获取方式,确定对应的数据获取策略,利用多线程并发,进行目标数据的获取,实现了从数据库中快速有效获取目标数据的方案,可以提高数据处理平台的工作效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115438056B_ABST
    Figure CN115438056B_ABST
Patent Text Reader

Abstract

The application discloses a data acquisition method, device and equipment and a storage medium. The method comprises the following steps: in response to a target data acquisition request, if the data acquisition mode of to-be-acquired data is a non-real-time acquisition mode, then according to the target data acquisition request, a target data acquisition task corresponding to the target data acquisition request is divided into at least two subtasks, and a corresponding candidate data acquisition request is generated for each subtask; according to the candidate data acquisition request, at least two working threads are controlled to send the candidate data acquisition request corresponding to each subtask to a corresponding ES cluster based on an inner lock in a corresponding HashMap, and at least two candidate data are acquired; and according to the at least two candidate data, target acquisition data is determined. According to the scheme provided by the application, a corresponding data acquisition strategy can be determined according to the data acquisition mode, and the target data can be quickly and effectively acquired by using multithreading concurrency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data, and in particular to a data acquisition method, apparatus, device, and storage medium. Background Technology

[0002] With the continuous development of big data technology, big data platforms are widely used to acquire data from various databases for subsequent analysis and log generation. For example, a log big data platform built on the ELK framework.

[0003] Because big data platforms need to acquire data with complex data types and large volumes, how to more effectively acquire data from databases and improve the efficiency of data processing platforms is an urgent problem to be solved. Summary of the Invention

[0004] This invention provides a data acquisition method, apparatus, device, and storage medium, which can determine the corresponding data acquisition strategy according to the data acquisition method, and at the same time utilize multi-threaded concurrency to achieve fast and effective acquisition of target data.

[0005] According to one aspect of the present invention, a data acquisition method is provided, comprising:

[0006] In response to a target data acquisition request, if the data acquisition method for the data to be acquired is a non-real-time acquisition method, then according to the target data acquisition request, the target data acquisition task corresponding to the target data acquisition request is divided into at least two sub-tasks, and a corresponding alternative data acquisition request is generated for each sub-task.

[0007] Based on the alternative data acquisition request, at least two worker threads are concurrently controlled to send the alternative data acquisition request corresponding to each subtask to the corresponding ES cluster based on the corresponding HashMap internal lock, and acquire at least two alternative data.

[0008] Based on the at least two alternative data, the target data is determined.

[0009] According to another aspect of the present invention, a data acquisition apparatus is provided, comprising:

[0010] The generation module is used to respond to the target data acquisition request. If the data acquisition method of the data to be acquired is a non-real-time acquisition method, then according to the target data acquisition request, the target data acquisition task corresponding to the target data acquisition request is divided into at least two sub-tasks, and a corresponding alternative data acquisition request is generated for each sub-task.

[0011] The first acquisition module is used to concurrently control at least two worker threads to send the candidate data acquisition requests corresponding to each subtask to the corresponding ES cluster based on the corresponding HashMap internal lock, and acquire at least two candidate data according to the candidate data acquisition request.

[0012] The first data determination module is used to determine the target data to be acquired based on the at least two candidate data.

[0013] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0014] At least one processor; and

[0015] A memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data acquisition method according to any embodiment of the present invention.

[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the data acquisition method described in any embodiment of the present invention.

[0018] The technical solution of this invention, in response to a target data acquisition request, if the data acquisition method is non-real-time, divides the target data acquisition task corresponding to the target data acquisition request into at least two sub-tasks, and generates a corresponding alternative data acquisition request for each sub-task. Based on the alternative data acquisition requests, at least two worker threads are concurrently controlled to send the alternative data acquisition requests corresponding to each sub-task to the corresponding Elasticsearch cluster based on the corresponding HashMap lock, acquiring at least two alternative data sets. Based on the at least two alternative data sets, the target data to be acquired is determined. By determining the data acquisition method based on the data acquisition request, and further determining the corresponding data acquisition strategy based on the data acquisition method, and utilizing multi-threaded concurrency to acquire the target data, a solution for quickly and effectively acquiring target data from the database is achieved, which can improve the working efficiency of the data processing platform.

[0019] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart of a data acquisition method provided in Embodiment 1 of the present invention;

[0022] Figure 2 This is a flowchart of a data acquisition method provided in Embodiment 2 of the present invention;

[0023] Figure 3 This is a flowchart of a data acquisition method provided in Embodiment 3 of the present invention;

[0024] Figure 4A This is a flowchart of a data acquisition method provided in Embodiment 4 of the present invention;

[0025] Figure 4B This is a flowchart of a data acquisition method provided in Embodiment 4 of the present invention;

[0026] Figure 5 This is a structural diagram of a data acquisition device provided in Embodiment 5 of the present invention;

[0027] Figure 6 This is a schematic diagram of the structure of the electronic device provided in Embodiment Six of the present invention. Detailed Implementation

[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "alternative," "target," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0030] Example 1

[0031] Figure 1 This is a flowchart of a data acquisition method provided in Embodiment 1 of the present invention. This embodiment is applicable to situations where a big data platform acquires data from a database, and is particularly applicable to situations where a big data platform acquires data using different acquisition strategies based on the data acquisition method. This method can be executed by a data acquisition device, which can be implemented in software and / or hardware and can be integrated into an electronic device with data acquisition capabilities, such as a log big data platform built on the ELK (Elasticsearch Logstash Kibana) framework. Here, the big data processing platform refers to a data stream processing component platform and a data storage platform. Figure 1 As shown, the method includes:

[0032] S101. In response to the target data acquisition request, if the data acquisition method of the data to be acquired is a non-real-time acquisition method, then according to the target data acquisition request, the target data acquisition task corresponding to the target data acquisition request is divided into at least two sub-tasks, and a corresponding alternative data acquisition request is generated for each sub-task.

[0033] In this context, "target data" refers to data obtained from a database, such as operational metric data acquired by a log big data platform. Specifically, the target data described in this embodiment is data that does not require real-time acquisition, i.e., batch metric data. Batch metric data primarily refers to data with large volumes and low real-time requirements. Metric data refers to secondary data obtained by the big data platform after statistical analysis of big data based on indexes. A target data acquisition request is a request to acquire the target data. Data acquisition methods include non-real-time acquisition and real-time acquisition. Alternative data acquisition requests refer to data acquisition requests corresponding to each subtask; that is, each subtask corresponds to one alternative data acquisition request.

[0034] Optionally, the big data processing platform can automatically generate batch data retrieval requests during preset time periods, such as early morning when the platform load is relatively low; or it can generate corresponding data retrieval requests upon detecting user input of data retrieval instructions.

[0035] Optionally, after detecting a data acquisition request, the big data platform can analyze the request and extract relevant information about the target data, such as the time period and data type. For example, it can acquire historical indicator data generated the previous day. Based on the relevant information about the target data, it can determine the data to be acquired and the preset data acquisition method.

[0036] Optionally, if the data to be acquired is acquired in a non-real-time manner, a target data acquisition task can be generated based on the relevant information of the target data in the target data acquisition request. Then, according to preset rules, such as the data type of the target data or preset index metrics, the target data acquisition task corresponding to the target data acquisition request can be divided into at least two subtasks. Alternatively, the target data acquisition task can be directly input into a pre-trained model, which outputs at least two subtasks. Furthermore, based on the target data for each subtask and the database cluster where the target data resides, a corresponding alternative data acquisition request is generated for each subtask.

[0037] For example, if the target data acquisition task is to acquire video data from the previous day, and the video data types include MP4 and WMV (Windows Media Video), then the target data acquisition task can be divided into two sub-tasks: acquiring MP4 video data from the previous day and acquiring WMV video data from the previous day.

[0038] S102. Based on the alternative data retrieval request, concurrently control at least two worker threads to send the alternative data retrieval request corresponding to each subtask to the corresponding ES cluster based on the corresponding HashMap internal lock, and retrieve at least two alternative data.

[0039] In this context, a worker thread (Thread) refers to a thread used to send data retrieval requests to the database cluster and obtain feedback data. A HashMap lock refers to an internal lock of type HashMap. Specifically, a HashMap type lock pool can be set up in the thread pool, i.e., setting a HashMap protection lock. A HashMap protection lock can include at least two HashMap internal locks. Each HashMap internal lock corresponds one-to-one with a database cluster. A database cluster refers to a cluster of commonly used databases; each database can include at least one database cluster. For Elasticsearch (ES) databases, they can be divided into at least two ES clusters according to preset partitioning rules, such as dividing the database into at least two clusters based on the time period of the records stored in each cluster.

[0040] Optionally, a thread pool can be set up, which contains the same number of core threads and maximum threads as the number of Elasticsearch clusters.

[0041] Optionally, based on the alternative data retrieval requests corresponding to each subtask, the ES cluster where each alternative data to be retrieved is located and the HashMap lock corresponding to the ES cluster can be determined. Further, based on the ES cluster where each alternative data to be retrieved is located, the same number of worker threads as the ES cluster can be created in the thread pool. Concurrent control can be exercised on each worker thread to send the alternative data retrieval requests corresponding to each subtask to the corresponding ES cluster based on the corresponding HashMap lock, so as to retrieve at least two alternative data.

[0042] Optionally, the identification information of the ES cluster and HashMap locks can be used to distinguish between different ES clusters and HashMap locks. For example, they can be distinguished based on the ES cluster and HashMap lock numbers. Accordingly, based on the alternative data retrieval requests, at least two worker threads are concurrently controlled to send the alternative data retrieval requests corresponding to each subtask to the corresponding ES cluster based on the corresponding HashMap locks, and retrieve at least two alternative data items. This includes: determining the ES cluster number and HashMap lock number corresponding to each subtask based on the alternative data retrieval requests; determining the number of worker threads to be called based on the number of subtasks; and concurrently controlling at least two worker threads to send the alternative data retrieval requests corresponding to each subtask to the corresponding ES cluster based on the corresponding HashMap locks, and retrieve the alternative data.

[0043] It should be noted that setting a HashMap type lock pool in the thread pool is the key to controlling the serial data acquisition of a single Elasticsearch cluster. Only by acquiring the lock corresponding to the Elasticsearch cluster can the worker thread send the access request to the Elasticsearch cluster and then obtain the query or aggregation result.

[0044] S103. Based on at least two alternative data, determine the target data to be acquired.

[0045] Among them, target data acquisition refers to the number of target data requested in the target data acquisition request.

[0046] Optionally, after determining at least two candidate data, the candidate data can be directly input into a pre-trained model to output the target acquisition data; alternatively, the candidate data can be processed according to preset rules to determine the target acquisition data. Specifically, determining the target acquisition data based on at least two candidate data includes: performing processing operations on at least two candidate data to generate the target acquisition data; wherein, the processing operations include merging and / or filtering operations.

[0047] Optionally, if there is abnormal data in at least two candidate data sets, the abnormal data can be removed, that is, a filtering operation can be performed on the candidate data, and the filtered candidate data can be merged to generate the target acquisition data; alternatively, at least two candidate data sets can be filtered according to preset rules, and the filtered candidate data can be used as the target acquisition data, that is, the target acquisition data can be generated; if there is no abnormal data in at least two candidate data sets, at least two candidate data sets can be merged directly to generate the target acquisition data.

[0048] The technical solution of this invention, in response to a target data acquisition request, if the data acquisition method is non-real-time, divides the target data acquisition task corresponding to the target data acquisition request into at least two sub-tasks, and generates a corresponding alternative data acquisition request for each sub-task. Based on the alternative data acquisition requests, at least two worker threads are concurrently controlled to send the alternative data acquisition requests corresponding to each sub-task to the corresponding Elasticsearch cluster based on the corresponding HashMap lock, acquiring at least two alternative data sets. Based on the at least two alternative data sets, the target data to be acquired is determined. By determining the data acquisition method based on the data acquisition request, and further determining the corresponding data acquisition strategy based on the data acquisition method, and utilizing multi-threaded concurrency to acquire the target data, a solution for quickly and effectively acquiring target data from the database is achieved, which can improve the working efficiency of the data processing platform.

[0049] Example 2

[0050] Figure 2This is a flowchart of a data acquisition method provided in Embodiment 2 of the present invention. Based on the above embodiments, this embodiment further explains in detail the following: "In response to a target data acquisition request, if the data acquisition method for the data to be acquired is a non-real-time acquisition method, then according to the target data acquisition request, the target data acquisition task corresponding to the target data acquisition request is divided into at least two sub-tasks, and a corresponding alternative data acquisition request is generated for each sub-task." Figure 2 As shown, the method includes:

[0051] S201. In response to the target data acquisition request, determine the data acquisition method for the data to be acquired.

[0052] The data acquisition methods include: non-real-time acquisition and real-time acquisition.

[0053] Optionally, the target data acquisition request can be analyzed to determine the type of data to be acquired. If the type of data to be acquired is historical data, that is, the recording time is within a historical time period, then the data acquisition method of the data to be acquired can be determined to be a non-real-time acquisition method. If the type of data to be acquired is data that needs to be acquired in real time, such as ledger page data, then the data acquisition method of the data to be acquired can be determined to be a real-time acquisition method.

[0054] S202. If the data to be acquired is acquired in a non-real-time manner, then determine the target data acquisition task and at least two indicators based on the target data acquisition request.

[0055] The metrics here refer to data index metrics. Metrics include at least one of the following: data access method, index log sending partition network segment, and the system to which the index belongs. Data access method refers to the way data is accessed based on different data transmission tools. Data access methods can include embedded framework access, external acquisition component access, and UDP (User Datagram Protocol) message transmission access, etc. Data transmission tools can include Filebeat (log file (collecting file data)), Flume (log collection system), Logback (open source logging component), Log4j (logging software), Log4net (a tool that helps programmers output log information to various targets (console, file, database, etc.)) and UDP, etc.

[0056] Optionally, a target data acquisition task can be generated based on the relevant information of the target data acquisition in the data acquisition request. Based on the target data acquisition task and preset rules, at least two indicators can be determined. Specifically, at least two indicators can be determined based on the different dimensions of the target data acquisition task.

[0057] For example, if the target data acquisition task is to obtain the student ID and name of students in class A, then the two indicators can be determined as student ID and name based on these two dimensions.

[0058] S203. Based on the number of indicators, divide the target data acquisition task into at least two sub-tasks, and generate corresponding alternative data acquisition requests for each sub-task.

[0059] The number of subtasks is the same as the number of metrics.

[0060] For example, if the target data acquisition task is to obtain the student ID and name of students in class A, and the two indicators are student ID and name, then the target data acquisition task can be divided into two sub-tasks based on these two indicators. Sub-task 1 is to obtain the student ID data of students in class A by student ID index, and sub-task 2 is to obtain the name data of students in class A by name index.

[0061] Optionally, after dividing the task into subtasks, alternative data retrieval requests can be generated for each subtask based on the content of each subtask and the database cluster where the target data is located.

[0062] S204. Based on the alternative data retrieval request, concurrently control at least two worker threads to send the alternative data retrieval request corresponding to each subtask to the corresponding ES cluster based on the corresponding HashMap internal lock, and retrieve at least two alternative data.

[0063] S205. Based on at least two alternative data, determine the target data to be acquired.

[0064] The technical solution of this invention, in response to a target data acquisition request, first determines the data acquisition method for the data to be acquired. When the data acquisition method is non-real-time acquisition, it determines the target data acquisition task and at least two indicators based on the target data acquisition request. Based on the number of indicators, the target data acquisition task is divided into at least two sub-tasks, and a corresponding alternative data acquisition request is generated for each sub-task. Finally, at least two worker threads are used to acquire alternative data from the corresponding Elasticsearch cluster and determine the target data to be acquired. This approach provides a feasible method for generating at least two alternative data acquisition requests when the data acquisition method is non-real-time acquisition. It allows for better division of the target data acquisition task, identification of accurate alternative data acquisition requests, and facilitates the subsequent generation of accurate target data, thereby improving the working efficiency of the big data processing platform.

[0065] Example 3

[0066] Figure 3This is a flowchart of a data acquisition method provided in Embodiment 3 of the present invention. Based on the above embodiments, this embodiment further explains in detail how the data acquisition process is performed when the data to be acquired is acquired in real time. Figure 3 As shown, the method includes:

[0067] S301. In response to the target data acquisition request, if the data acquisition method of the data to be acquired is real-time acquisition, then determine at least two data sources associated with the data to be acquired.

[0068] Here, the data source refers to the database cluster from which the data originates. Data sources can include Elasticsearch (ES) clusters and database (DB) clusters. A DB cluster is a cluster of databases within a database (DB) database.

[0069] Optionally, when the big data platform receives a request to load the ledger page from the browser, it can automatically generate a real-time target data acquisition request. Since the page data needs to be acquired from different data sources in real time, the data acquisition method for the data to be acquired is real-time acquisition.

[0070] Optionally, if the data to be acquired is acquired in real time, the relevant information of the target data in the target data acquisition request can be used to query the table that corresponds one-to-one with the database cluster storing the corresponding data to determine the minimum database cluster corresponding to the target data and, based on the type of database cluster, determine at least two data sources.

[0071] S302. For each data source, determine at least two worker threads.

[0072] Optionally, after determining at least two data sources associated with the data to be acquired, at least two worker threads can be created in the thread pool for each data source. That is, for each data source, at least two worker threads are used to acquire data from the database cluster of that data source.

[0073] S303. Control at least two worker threads to concurrently execute aggregation, query, or selection operations, and obtain at least two alternative data from each worker thread.

[0074] Aggregation refers to the operation of aggregating field values ​​using SQL statements, i.e., the aggs statement. Query refers to the operation of querying field values ​​using SQL statements, i.e., the query statement. Selection refers to the operation of filtering field values ​​using SQL statements, i.e., the select statement.

[0075] Optionally, control at least two worker threads to concurrently execute aggregation, query, or selection operations, and obtain at least two candidate data from each worker thread. This includes using a CountDownLatch counter to obtain at least two candidate data from each worker thread when it is detected that at least two worker threads have completed concurrent aggregation, query, or selection operations. Here, CountDownLatch is a counter used to count the number of threads out of N threads that have completed the task.

[0076] Optionally, after determining at least two worker threads for each data source, the number of worker threads that need to be executed concurrently can be determined. Based on the operations that each worker thread needs to perform, at least two worker threads are controlled to execute the corresponding operations concurrently. Although each worker thread starts executing operations at the same time, different operations require different execution times. Therefore, a CountDownLatch counter can be used to increment the CountDownLatch counter by 1 when each worker thread finishes execution, until the number recorded by the counter is the same as the number of worker threads that need to be executed concurrently. At this point, it is considered that at least two worker threads have completed the concurrent execution of aggregation, query, or selection operations. The feedback data obtained when each worker thread finishes execution is obtained, that is, at least two alternative data returned by each worker thread.

[0077] S304. Based on at least two alternative data, determine the target data to be acquired.

[0078] Optionally, processing operations can be performed on at least two candidate data sets to generate the target acquisition data, i.e., to determine the target acquisition data. These processing operations include merging and / or filtering operations. Specifically, the process of performing merging and / or filtering operations on at least two candidate data sets to generate the target acquisition data has been explained in detail in the above embodiments and will not be repeated here.

[0079] The technical solution of this invention, in response to a target data acquisition request, if the data acquisition method is real-time acquisition, determines at least two data sources associated with the data to be acquired. For each data source, at least two worker threads are determined, and these at least two worker threads are controlled to concurrently execute aggregation, query, or selection operations. At least two candidate data points are obtained from each worker thread, and finally, the target data to be acquired is determined based on these at least two candidate data points. This approach provides a feasible method for determining the target data to be acquired when the data acquisition method is real-time acquisition, by obtaining candidate data from different data sources. This method can better determine accurate candidate data, thereby generating accurate target data and improving the working efficiency of the big data processing platform.

[0080] Example 4

[0081] Figure 4A This is a flowchart of a data acquisition method provided in Embodiment 4 of the present invention. Figure 4B This is a flowchart of a data acquisition method provided in Embodiment 4 of the present invention. Based on the above embodiments, this embodiment provides preferred examples of how a big data platform can utilize multi-threaded concurrency to acquire data when the data acquisition method is real-time acquisition and non-real-time acquisition.

[0082] For example, such as Figure 4A As shown, for data that needs to be acquired in non-real-time, the methods for acquiring data include the following processes: Big data platforms can use the RESTful APIs provided by Java and Elasticsearch for data collection and processing.

[0083] Specifically, a main thread can be created first to generate the target data retrieval task. Then, for each index's data (i.e., based on the metric), the target data retrieval task is encapsulated into at least two Runnables (i.e., subtasks), such as Subtask 1, Subtask 2, and Subtask 3. An incremental task number can be set based on the Elasticsearch cluster to which the index data belongs. Specifically, an incremental and non-repeating value can be set in each task's Runnable for each index in each Elasticsearch cluster. Simultaneously, a fixed-size thread pool is started, where the core thread count and maximum thread count are both set to the number of Elasticsearch clusters. Before the thread pool starts running, the subtasks with counts of 1, 2, and 3 (i.e., Subtask 1, Subtask 2, and Subtask 3) are submitted to their corresponding worker threads. Then, before sending alternative data retrieval requests to the corresponding Elasticsearch cluster based on the index name (metric), the locks corresponding to each Elasticsearch cluster need to be acquired first.

[0084] Optionally, locking can be implemented by setting up a HashMap-type lock pool within the thread pool. This is crucial for controlling the serial data retrieval within a single Elasticsearch cluster. Only after acquiring the lock corresponding to the Elasticsearch cluster can a worker thread send an access request to the cluster to retrieve query or aggregation results. For example, the lock for Elasticsearch cluster 1 might be HashMap lock 1, the lock for Elasticsearch cluster 2 might be HashMap lock 2, and the lock for Elasticsearch cluster 3 might be HashMap lock 3. By processing the subtask data (potential data) retrieved by each worker thread, the target data can be determined.

[0085] For example, such as Figure 4BAs shown, for the scenario where the browser sends ledger page loading requests to various servers of the big data platform to obtain data in real time, the specific data acquisition method includes the following process: The big data platform can receive the ledger page loading requests (i.e., data acquisition requests for page data) sent by the browser. Based on the content of the request, it determines the number of associated data sources and calls the corresponding number of servers, such as server 1 and server 2, to obtain candidate data from different data sources. For the ES cluster data source, server 1 can simultaneously call worker thread 1 and worker thread 2 in the thread pool to perform query and aggregation operations respectively, that is, send query requests and aggregation requests to the ES cluster, and determine candidate data 1 and candidate data 2 based on the corresponding query and aggregation feedback. For the DB cluster data source, server 2 can simultaneously call worker thread 3 and worker thread 4 in the thread pool to perform selection operations respectively, that is, send corresponding selection requests to the DB cluster, and determine candidate data 3 and candidate data 4 based on the corresponding selection feedback. By processing the candidate data obtained by each worker thread, the target data to be obtained can be determined.

[0086] The technical solution of this invention determines the data acquisition method based on the data acquisition request, further determines the corresponding data acquisition strategy based on different data acquisition methods, and uses multi-threaded concurrency to acquire the target data. This achieves a solution for quickly and effectively acquiring target data from the database, improves the richness of data acquisition methods, and enhances the working efficiency of the data processing platform.

[0087] Example 5

[0088] Figure 5 This is a structural diagram of a data acquisition device provided in Embodiment 5 of the present invention. The data acquisition device provided in this embodiment of the present invention can execute a data acquisition method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.

[0089] like Figure 5 As shown, the device includes:

[0090] The generation module 501 is used to respond to the target data acquisition request. If the data acquisition method of the data to be acquired is a non-real-time acquisition method, then according to the target data acquisition request, the target data acquisition task corresponding to the target data acquisition request is divided into at least two sub-tasks, and a corresponding alternative data acquisition request is generated for each sub-task.

[0091] The first acquisition module 502 is used to concurrently control at least two worker threads to send the candidate data acquisition requests corresponding to each subtask to the corresponding ES cluster based on the corresponding HashMap internal lock, in order to acquire at least two candidate data according to the candidate data acquisition request.

[0092] The first data determination module 503 is used to determine the target acquisition data based on the at least two alternative data.

[0093] The technical solution of this invention, in response to a target data acquisition request, if the data acquisition method is non-real-time, divides the target data acquisition task corresponding to the target data acquisition request into at least two sub-tasks, and generates a corresponding alternative data acquisition request for each sub-task. Based on the alternative data acquisition requests, at least two worker threads are concurrently controlled to send the alternative data acquisition requests corresponding to each sub-task to the corresponding Elasticsearch cluster based on the corresponding HashMap lock, acquiring at least two alternative data sets. Based on the at least two alternative data sets, the target data to be acquired is determined. By determining the data acquisition method based on the data acquisition request, and further determining the corresponding data acquisition strategy based on the data acquisition method, and utilizing multi-threaded concurrency to acquire the target data, a solution for quickly and effectively acquiring target data from the database is achieved, which can improve the working efficiency of the data processing platform.

[0094] Furthermore, the generation module 501 is specifically used for:

[0095] In response to a target data acquisition request, the data acquisition method for the data to be acquired is determined; the data acquisition method includes: non-real-time acquisition method and real-time acquisition method;

[0096] If the data to be acquired is acquired in a non-real-time manner, then the target data acquisition task and at least two metrics are determined based on the target data acquisition request.

[0097] Based on the number of the indicators, the target data acquisition task is divided into at least two sub-tasks, and a corresponding alternative data acquisition request is generated for each sub-task.

[0098] Furthermore, the indicators include at least one of the following: data access method, index log sending partition network segment, and the system to which the index belongs.

[0099] Furthermore, the first acquisition module 502 is specifically used for:

[0100] Based on the alternative data acquisition request, determine the ES cluster number and HashMap lock number corresponding to each subtask;

[0101] Based on the number of subtasks, determine the number of worker threads that need to be invoked;

[0102] Based on the ES cluster number and the HashMap internal lock number, at least two worker threads are concurrently controlled to send the alternative data retrieval requests corresponding to each subtask to the corresponding ES cluster based on the corresponding HashMap internal lock, and retrieve the alternative data.

[0103] Furthermore, the aforementioned device also includes:

[0104] The data source determination module is used to respond to the target data acquisition request. If the data to be acquired is acquired in real time, it determines at least two data sources associated with the data to be acquired.

[0105] The thread determination module is used to determine at least two worker threads for each data source;

[0106] The second acquisition module is used to control at least two worker threads to concurrently execute aggregation, query or selection operations, and to acquire at least two alternative data from each worker thread.

[0107] The second data determination module is used to determine the target acquisition data based on the at least two candidate data.

[0108] Furthermore, the second acquisition module is specifically used for:

[0109] Using a CountDownLatch counter, when it is detected that at least two worker threads have completed executing aggregation, query, or selection operations concurrently, at least two alternative data points are obtained from each worker thread.

[0110] Furthermore, the first data determination module 503 is specifically used for:

[0111] Processing operations are performed on the at least two candidate data to generate target acquisition data; the processing operations include merging and / or filtering operations.

[0112] Example 6

[0113] Figure 6 This is a schematic diagram of the structure of the electronic device provided in Embodiment Six of the present invention. Figure 6 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0114] like Figure 6As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0115] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0116] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as data acquisition methods.

[0117] In some embodiments, the data acquisition method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the data acquisition method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the data acquisition method by any other suitable means (e.g., by means of firmware).

[0118] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0119] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0120] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0121] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0122] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0123] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0124] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0125] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A data acquisition method, applied to a log big data platform built using the ELK framework, characterized in that, include: In response to a target data acquisition request, determine the data acquisition method for the data to be acquired; The data acquisition methods include: non-real-time acquisition methods and real-time acquisition methods; If the data to be acquired is acquired in a non-real-time manner, then the target data acquisition task and at least two metrics are determined based on the target data acquisition request. Based on the number of the indicators, the target data acquisition task is divided into at least two sub-tasks, and a corresponding alternative data acquisition request is generated for each sub-task, wherein the number of sub-tasks is the same as the number of the indicators; Based on the alternative data acquisition request, at least two worker threads are concurrently controlled to send the alternative data acquisition request corresponding to each subtask to the corresponding ES cluster based on the corresponding HashMap internal lock, and acquire at least two alternative data. Based on the at least two candidate data, the target data to be acquired is determined; If the data to be acquired is acquired in real time, then at least two data sources associated with the data to be acquired are determined, including an ES cluster and a DB cluster. For each data source, identify at least two worker threads; Control at least two worker threads to concurrently execute aggregation, query, or selection operations, and obtain at least two alternative data from each worker thread; Based on the at least two alternative data, the target data is determined.

2. The method according to claim 1, characterized in that, in, The metrics include at least one of the following: data access method, index log sending partition network segment, and the system to which the index belongs.

3. The method according to claim 1, characterized in that, Based on the candidate data retrieval request, at least two worker threads are concurrently controlled to send the candidate data retrieval request corresponding to each subtask to the corresponding Elasticsearch cluster based on the corresponding HashMap internal lock, and retrieve at least two candidate data items, including: Based on the alternative data acquisition request, determine the ES cluster number and HashMap lock number corresponding to each subtask; Based on the number of subtasks, determine the number of worker threads that need to be invoked; Based on the ES cluster number and the HashMap internal lock number, at least two worker threads are concurrently controlled to send the alternative data retrieval requests corresponding to each subtask to the corresponding ES cluster based on the corresponding HashMap internal lock, and retrieve the alternative data.

4. The method according to claim 1, characterized in that, Control at least two worker threads to concurrently execute aggregation, query, or selection operations, and obtain at least two alternative data from each worker thread, including: Using a CountDownLatch counter, when it is detected that at least two worker threads have completed executing aggregation, query, or selection operations concurrently, at least two alternative data points are obtained from each worker thread.

5. The method according to claim 1, characterized in that, Based on the at least two candidate data, the target data to be acquired is determined, including: Processing operations are performed on the at least two candidate data to generate target acquisition data; the processing operations include merging and / or filtering operations.

6. A data acquisition device, integrated into a log big data platform built on the ELK framework, characterized in that, include: The generation module is used to: determine the data acquisition method for the data to be acquired in response to a target data acquisition request; The data acquisition methods include: non-real-time acquisition and real-time acquisition; if the data acquisition method for the data to be acquired is non-real-time acquisition, then based on the target data acquisition request, a target data acquisition task and at least two indicators are determined; based on the number of indicators, the target data acquisition task is divided into at least two sub-tasks, and a corresponding alternative data acquisition request is generated for each sub-task, wherein the number of sub-tasks is the same as the number of indicators; The first acquisition module is used to concurrently control at least two worker threads to send the candidate data acquisition requests corresponding to each subtask to the corresponding ES cluster based on the corresponding HashMap internal lock, and acquire at least two candidate data according to the candidate data acquisition request. The first data determination module is used to determine the target data to be acquired based on the at least two candidate data. The data source determination module is used to determine at least two data sources associated with the data to be acquired if the data acquisition method is real-time acquisition. The data sources include an ES cluster and a DB cluster. The thread determination module is used to determine at least two worker threads for each data source; The second acquisition module is used to control at least two worker threads to concurrently execute aggregation, query or selection operations, and to acquire at least two alternative data from each worker thread. The second data determination module is used to determine the target acquisition data based on the at least two candidate data.

7. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data acquisition method according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute the data acquisition method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Data insertion method and device

    CN105095261A

  • File retrieving system based on HDFS (Hadoop Distributed File System)

    CN106484877A

  • Data acquisition method, apparatus and system

    CN109101330A