Data processing method and device, electronic equipment and storage medium
Through the combination of task queues and thread pools, the problem of excessive use of thread resources in big data query tasks is solved, efficient query progress information management and timely acquisition of query results are achieved, and the capacity and user experience of query services are improved.
Patent Information
- Application Number
- CN202410030906.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-08
- Publication Date
- 2025-07-08
AI Technical Summary
In existing big data query tasks, each user request needs to create a corresponding thread to respond to the data processing task, resulting in excessive use of thread resources and affecting the efficiency of the query task.
Using the task queue and preset thread pool, after receiving the query task, it will be added to the task queue, and the corresponding thread processing task is determined from the preset thread pool, and query progress information is obtained through the asynchronous non-blocking interface, and stored in external storage space to reduce memory usage.
This greatly improves the capacity and efficiency of query services, reduces thread resource usage, avoids memory overflow, and improves the management efficiency and user experience of query progress information.
Smart Images

Figure CN120276814A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of big data technology, and in particular, to a data processing method, apparatus, electronic device, and storage medium. Background Art
[0002] A big data query task refers to a task of querying and analyzing massive data in a big data environment. With the rapid development of the Internet and the Internet of Things, the scale and complexity of big data have been continuously increasing, and traditional query methods can no longer meet the requirements for efficiently querying large-scale data. Therefore, the goal of big data query tasks is to provide fast and accurate query results by leveraging the capabilities of distributed computing and parallel processing.
[0003] Currently, mainstream big data query engines have all implemented the database access language "SQL" and also reused the database access interface "JDBC" in the programming interface. This solution has relatively well solved the compatibility problem when migrating from traditional databases to big data computing engines. However, in the processing of some "large-scale" query tasks (mainly reflected in a large number of tasks and a large query result set), this technical solution has obvious bottlenecks: a synchronous blocking programming model, where each user request requires creating a corresponding thread to respond to the data processing task, resulting in excessive occupation of thread resources by query tasks. Summary of the Invention
[0004] The present disclosure provides a data processing method, apparatus, electronic device, and storage medium. To overcome the problem that each user request requires creating a corresponding thread to respond to the data processing task, resulting in excessive occupation of thread resources by query tasks.
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a data processing method, including:
[0006] When receiving a first query task, adding the first query task to a task queue; wherein, the task queue is used to cache each query task received and updated.
[0007] Determining a first thread corresponding to the first query task from each to-be-called thread in a preset thread pool; wherein, the preset thread pool includes: at least two to-be-called threads.
[0008] Sending the first query task to a query engine through the first thread.
[0009] Obtaining query progress information of the first query task returned by the query engine.
[0010] In some embodiments, the method further includes:
[0011] In the case of receiving the first query task for the first time, if the query progress information of the first query task is successfully obtained, the query progress information of the first query task is sent to the application side.
[0012] In some embodiments, the method further includes:
[0013] In the case of obtaining the query progress information of the first query task from the query engine, an association relationship is established between the task identifier of the first query task and the first query task; wherein, the task identifier is: an identifier preset for uniquely identifying the query task;
[0014] Based on the association relationship, the query progress information of the first query task is updated to a preset storage space;
[0015] Wherein, the preset storage space is used to store the query progress information corresponding to each of the query tasks.
[0016] In some embodiments, the method further includes:
[0017] Receiving a first progress query request sent by the application side; wherein, the first progress query request carries the task identifier of the second query task;
[0018] Based on the task identifier of the second query task, the query progress information of the second query task is determined from the preset storage space;
[0019] The query progress information of the second query task is sent to the application side.
[0020] In some embodiments, the method further includes:
[0021] Sending a second progress query request to the query engine at a preset time interval; wherein, the second progress query request carries the task identifiers of at least one of the query tasks in the task queue;
[0022] Determining a second thread corresponding to the at least one query task from a preset thread pool;
[0023] Through each of the second threads, the query progress information of each of the query tasks obtained based on the task identifiers of the at least one query task is respectively obtained from the query engine;
[0024] The query progress information of each of the query tasks is updated to the preset storage space.
[0025] In some embodiments, the method further includes:
[0026] In the case of successfully obtaining the query progress information of the first query task, update the first query task based on the query progress information of the first query task to obtain a third query task;
[0027] Update the first query task in the task queue to the third query task.
[0028] In some embodiments, the method further includes:
[0029] Obtain the query progress information of the query tasks that have been successfully completed from the query engine;
[0030] Store the query results of the query tasks that have been successfully completed in a preset storage space, and send the query progress information of the query tasks that have been successfully completed to the application side.
[0031] In some embodiments, the storing the query results of the query tasks that have been successfully completed in a preset storage space includes:
[0032] Determine the query tasks that have been successfully completed based on the query progress information;
[0033] Determine a third thread corresponding to the query tasks that have been successfully completed from each of the to-be-called threads in the preset thread pool;
[0034] Use the third thread to obtain the query results of the query tasks that have been successfully completed from the query engine in a reactive stream processing manner, and store the query results of the query tasks that have been successfully completed in the preset storage space.
[0035] In some embodiments, the method further includes:
[0036] Receive a first result query request sent by the application side; wherein, the task identifier of the third query task is carried in the first result query request;
[0037] Based on the task identifier of the third query task, determine the query results of the third query task from the preset storage space.
[0038] According to the second aspect of the embodiments of the present disclosure, there is provided a data processing device, including:
[0039] A first receiving module, configured to add the first query task to a task queue when receiving the first query task; wherein, the task queue is used to cache each query task received and updated;
[0040] A first determination module, configured to determine a first thread corresponding to the first query task from each of the threads to be called in a preset thread pool; wherein, the preset thread pool includes: at least two of the threads to be called;
[0041] A first sending module, configured to send the first query task to a query engine through the first thread;
[0042] A first obtaining module, configured to obtain query progress information of the first query task returned by the query engine.
[0043] In some embodiments, the apparatus further includes:
[0044] A second sending module, configured to, in a case of first receiving the first query task, if successfully obtaining the query progress information of the first query task, send the query progress information of the first query task to the application side.
[0045] In some embodiments, the apparatus further includes:
[0046] An association module, configured to, in a case of obtaining the query progress information of the first query task from the query engine, establish an association relationship between a task identifier of the first query task and the first query task; wherein, the task identifier is: a preset identifier for uniquely identifying the query task;
[0047] A first update module, configured to update the query progress information of the first query task to a preset storage space based on the association relationship;
[0048] Wherein, the preset storage space is used to store query progress information corresponding to each of the query tasks.
[0049] In some embodiments, the apparatus further includes:
[0050] A second receiving module, configured to receive a first progress query request sent by the application side; wherein, the first progress query request carries a task identifier of a second query task;
[0051] A second determination module, configured to determine the query progress information of the second query task from the preset storage space based on the task identifier of the second query task;
[0052] A third sending module, configured to send the query progress information of the second query task to the application side.
[0053] In some embodiments, the apparatus further includes:
[0054] A fourth sending module, configured to send a second progress query request to the query engine at a preset time interval; wherein, the second progress query request carries the task identifiers of at least one of the query tasks in the task queue;
[0055] A third determining module, configured to determine a second thread corresponding to the at least one query task from a preset thread pool;
[0056] A second obtaining module, configured to respectively obtain, through each of the second threads, query progress information of each of the query tasks obtained based on the task identifiers of the at least one query task from the query engine;
[0057] A second updating module, configured to update the query progress information of each of the query tasks to a preset storage space.
[0058] In some embodiments, the device further includes:
[0059] A third updating module, configured to, in the case of successfully obtaining the query progress information of the first query task, update the first query task based on the query progress information of the first query task to obtain a third query task;
[0060] A fourth updating module, configured to update the first query task in the task queue to the third query task.
[0061] In some embodiments, the device further includes:
[0062] A third obtaining module, configured to obtain query progress information of successfully completed query tasks from the query engine;
[0063] A storage module, configured to store the query results of the successfully completed query tasks in a preset storage space, and send the query progress information of the successfully completed query tasks to the application side.
[0064] In some embodiments, the storage module is configured to:
[0065] Determine a successfully completed query task based on the query progress information;
[0066] Determine a third thread corresponding to the successfully completed query task from each of the to-be-called threads in the preset thread pool;
[0067] Use the third thread to obtain the query results of the successfully completed query tasks from the query engine in a reactive stream processing manner, and store the query results of the successfully completed query tasks in the preset storage space.
[0068] In some embodiments, the device further includes:
[0069] A third receiving module, configured to receive a first result query request sent by the application side; wherein, a task identifier of a third query task is carried in the first result query request;
[0070] A fourth determining module, configured to determine a query result of the third query task from the preset storage space based on the task identifier of the third query task.
[0071] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including:
[0072] A processor;
[0073] A memory configured to store processor-executable instructions;
[0074] Wherein, the processor is configured to: when executed, implement the steps in any one of the data processing methods in the first aspect above.
[0075] According to a fourth aspect of the embodiments of the present disclosure, there is provided a non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of a data processing device, enabling the device to execute the steps in any one of the data processing methods in the first aspect above.
[0076] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0077] When a first query task is received, adding the first query task to a task queue, determining a first thread corresponding to the first query task from each of the to-be-called threads in a preset thread pool, sending the first query task to a query engine through the first thread, and obtaining query progress information of the first query task returned by the query engine, wherein the task queue is used to cache each received and updated query task, and the preset thread pool includes: at least two of the to-be-called threads.
[0078] The technical solution of the present disclosure can add the query tasks received from the application side to the task queue of the query application server, and the query application server drives the processing of the entire query task, that is, each to-be-called thread in the preset thread pool processes all tasks in the task queue. This enables that after the query application server receives the first query task, it does not need to continuously occupy the application thread until the task is completed. Instead, it can add the first query task to the task queue and determine the first thread corresponding to the first query task from each to-be-called thread in the preset thread pool, and the first thread processes the first query task. Each to-be-called thread in the present disclosure can be used to process any query task in the task queue. Compared with the existing JDBC-based blocking task processing method, the 1:N (thread: task) task processing method in this solution can greatly improve the capacity of the query service. The service capacity is increased by more than a hundred times compared with the blocking task processing method, reducing the occupation of thread resources by large queries.
[0079] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.
[0081] Figure 1 is a flowchart of a data processing method shown according to an exemplary embodiment.
[0082] Figure 2 is a schematic diagram of the relationship between a task queue and a thread pool shown according to an exemplary embodiment.
[0083] Figure 3 is a processing flowchart of reactive stream processing shown according to an exemplary embodiment.
[0084] Figure 4 is a schematic diagram of a data processing scenario shown according to an exemplary embodiment.
[0085] Figure 5 is an interaction flowchart of a data processing method shown according to an exemplary embodiment.
[0086] Figure 6 is an interaction timing diagram of a data processing method shown according to an exemplary embodiment.
[0087] Figure 7 is a block diagram of a data processing device shown according to an exemplary embodiment.
[0088] Figure 8It is a hardware structure block diagram of an electronic device 800 shown according to an exemplary embodiment. Detailed implementation
[0089] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0090] Figure 1 It is a flowchart of a data processing method shown according to an exemplary embodiment. As Figure 1 shown, the method mainly includes the following steps:
[0091] In step 101, when a first query task is received, add the first query task to the task queue; wherein, the task queue is used to cache each query task received and updated.
[0092] In step 102, determine a first thread corresponding to the first query task from each to-be-called thread in a preset thread pool; wherein, the preset thread pool includes at least two to-be-called threads.
[0093] In step 103, send the first query task to the query engine through the first thread.
[0094] In step 104, obtain the query progress information of the first query task returned by the query engine.
[0095] It should be noted that this data processing method can be applied to an electronic device, and the electronic device may include a terminal device and an application server. The terminal device may include a mobile terminal and a fixed terminal. The mobile terminal may include a smart phone, a tablet computer, a wearable electronic device, etc., and the fixed terminal may include a desktop computer, an all-in-one computer, etc.
[0096] Among them, the query application server refers to a computer system used to provide query services in a computer network, which is responsible for processing query tasks and returning the processing results to the application side. The application server may be provided as a server, including but not limited to implementation manners such as traditional server architectures, cloud services, containerized architectures, Serverless architectures, distributed systems, and microservice architectures. These implementation manners have their respective advantages and applicable scenarios, and appropriate architectures and technologies can be selected according to requirements in terms of performance, maintainability, scalability, etc.
[0097] In an embodiment of the present disclosure, the first query task may be sent from the application side to the query application server.
[0098] It should be noted that the application side may include various electronic devices. The electronic devices may include terminal devices. Among them, the terminal devices may include mobile terminals and fixed terminals. For example, mobile phones, tablet computers, personal digital assistants, laptop computers, desktop computers, wearable devices, smart speakers, televisions, and vehicle-mounted terminals, etc. Among them, the wearable devices may include: portable devices directly worn on the user's body or integrated into the user's accessories. For example, it may include: wearable devices supported by the head (including headphones, glasses, helmets, headbands, etc.). It may also include: smart clothing, schoolbags, crutches, accessories, etc.
[0099] In an embodiment of the present disclosure, when the query application server receives the first query task sent by the application side, it may add the first query task to the task queue. Among them, the first query task may be the query task currently sent by the application side to the query application server, used to obtain the query result specified by the application side. The specified conditions include but are not limited to selected fields, query conditions, data sources, sorting fields, grouping fields, time ranges, filtering conditions, limiting the number of returned results, etc. The query results include but are not limited to formats such as JSON, CSV, and XML.
[0100] It should be noted that the task queue is used to cache each query task received and updated. Since the task queue is the only state of the entire query application server, to solve the problem of task state loss after the query application server restarts, the task queue can be externalized (such as Redis Queue).
[0101] Specifically, in the existing big data query task processing solution, the task state (including: query task and query progress information) is usually stored in the memory of the query application server. When the query application server restarts, the task state will be lost, and it is difficult to achieve lossless publishing of the client. By externalizing the task queue and not storing the query task in the memory of the application server, in this way, even if the query application server restarts, the task state will not be lost.
[0102] Figure 2 It is a schematic diagram showing the relationship between the task queue and the thread pool shown according to an exemplary embodiment, as Figure 2As shown, query tasks 1, 2, 3, etc. are stored in the task queue. Thread pool 1 can include at least one thread to be called. In some embodiments, the preset thread pool may include: the first thread pool (thread pool 1), and the threads to be called in the first thread pool are used to obtain query progress information; the second thread pool (thread pool 2), and the threads to be called in the second thread pool are used to obtain query results. In the embodiments of the present disclosure, when the application server adds the first query task to the task queue 201, it can determine the first thread corresponding to the first query task from each thread to be called in the preset thread pool 202. Among them, the thread to be called is used to obtain the first query task from the task queue for processing when the thread is idle. The thread states include busy, idle, faulty, etc. A thread can process at least one query task.
[0103] In some embodiments, each query task in the task queue can be sorted, and the threads to be called corresponding to each query task can be determined based on the sorting result. For example, it can be sorted in the order of "first in, first out", that is, the tasks added to the task queue first are processed first, and the tasks added to the task queue later are processed later.
[0104] In some other embodiments, the idle degrees of each thread to be called can also be sorted, and the threads corresponding to each query task can be determined from each thread to be called based on the sorting result. For example, when it is necessary to process the first query task, the idle degrees of each thread to be called in the preset thread pool can be sorted to determine the most idle thread, and then this thread is called to process the first query task.
[0105] In some other embodiments, the urgency levels of each query task in the task queue can also be sorted, and the threads to be called corresponding to each query task can be determined based on the sorting result. For example, the application side can assign a number to each query task according to the urgency level of the task. This number represents the urgency level of the task, and the number range is from 1 to 10. The larger the number, the more urgent the task and the faster it needs to be processed. At this time, a relatively idle thread can be called to process this task.
[0106] After determining the first thread corresponding to the first query task, the first query task can be processed by the first thread. For example, the first query task can be sent to the query engine by the first thread. After receiving the first query task, the query engine can obtain the query progress information of the first query task. It should be noted that the query engine is a software component or system used to execute data query operations. It can receive query tasks sent from the application side, interpret them, then execute corresponding query operations in the data storage, and finally return the results that meet the query conditions to the application side, including but not limited to various implementation methods such as MySQL Query Engine, Elasticsearch, Apache Hadoop, Apache Spark, etc.
[0107] Specifically, the query progress information includes a query status and a query log. Among them, the query status includes statuses such as waiting, compiling, executing, optimizing, scanning, aggregating, sorting, transmitting, waiting for a lock, completed, failed, and cancelled, and the query log includes information such as a query statement, a query start time and an end time, an execution plan, an execution status, error information, resource usage, execution time statistics, and resource usage.
[0108] The technical solution of the present disclosure can add the query tasks received from the application side to the task queue of the query application server, and the query application server drives the processing of the entire query task, that is, each pending thread in the preset thread pool processes all tasks in the task queue. This enables that after the query application server receives the first query task, it does not need to continuously occupy the application thread until the task is completed. Instead, the first query task can be added to the task queue, and a first thread corresponding to the first query task can be determined from each pending thread in the preset thread pool, and the first thread processes the first query task. Each pending thread in the present disclosure can be used to process any query task in the task queue. Compared with the existing JDBC-based blocking task processing method, the 1:N (thread: task) task processing method in this solution can greatly improve the capacity of the query service. The service capacity is increased by more than a hundred times compared with the blocking task processing method, reducing the occupation of thread resources by large queries.
[0109] In some embodiments, the method further includes:
[0110] In the case of first receiving the first query task, if the query progress information of the first query task is successfully obtained, the query progress information of the first query task is sent to the application side.
[0111] In the embodiments of the present disclosure, the first reception refers to the query application server receiving a new first query task sent by the application side. The new first query task refers to a query task generated by the application side according to actual usage requirements and has never been sent to the query application server before. For example, the query application server is currently processing a query task of "querying all users who have used the mobile phone for more than 5 hours", and at this time, it receives a query task of "querying all users who have used the mobile phone for more than 10 hours". At this time, "querying all users who have used the mobile phone for more than 10 hours" can be considered as a new first query task and is first received by the query application server.
[0112] In the embodiments of the present disclosure, the query application server receives a first query task sent by the application side and sends the first query task to the query engine. After receiving the first query task, the query engine immediately returns the query progress information of the first query task to the query application server through an asynchronous non-blocking interface. After successfully obtaining the query progress information of the first query task from the query engine, the query application server can send the query progress information of the first query task to the application side. It not only does not need to occupy the query application server thread for a long time until the first query task ends, which can reduce the occupation of the thread resources of the query application server and avoid memory overflow caused by excessive memory occupation, but also enables the application side to timely understand the query progress information of the first query task.
[0113] In some embodiments, the method further includes:
[0114] In the case of obtaining the query progress information of the first query task from the query engine, establishing an association relationship between the task identifier of the first query task and the first query task; wherein, the task identifier is: a preset identifier used to uniquely identify the query task;
[0115] Based on the association relationship, updating the query progress information of the first query task to a preset storage space.
[0116] In the embodiments of the present disclosure, before obtaining the query progress information of the first query task from the query engine, it is necessary to establish an association relationship between the task identifier of the first query task and the first query task. It should be noted that the task identifier is a preset identifier used to uniquely identify the query task, which can be calculated by a certain algorithm (such as a random number algorithm, a hash algorithm), and is usually a 16- to 32-bit string including numbers, letters, and characters, such as "550e8400e29b41d4a716446655440000". Each task identifier is globally unique, which can ensure that any task identifier generated at any time and place will not be repeated, and each task identifier can uniquely identify a query task for querying the query progress information of the corresponding query task.
[0117] Specifically, the association relationship refers to the correspondence between a task identifier and a query task, and each task identifier is associated with only one query task.
[0118] In an embodiment of the present disclosure, after obtaining the query progress information of the first query task from the query engine, the query progress information of the first query task can be stored in a preset storage space. It should be noted that the preset storage space is an external storage space different from the memory of the query application server, and mainly includes relational databases and non-relational databases. Among them, relational databases include, but are not limited to, implementation methods such as Oracle, DB2, PostgreSQL, Microsoft SQL Server, Microsoft Access, MySQL, and Inspur K-DB, and non-relational databases include, but are not limited to, implementation methods such as mongodb, cassandra, redis, hbase, and neo4j.
[0119] In an embodiment of the present disclosure, an association relationship between the task identifier of the first query task and the first query task is established, and the query progress information of the first query task is updated to the preset storage space based on the association relationship, which facilitates the external management of query tasks, can improve the query efficiency of query progress information, avoids caching a large amount of query progress information in the memory of the query application server, consumes less memory for the query application server, and solves the problem of memory overflow.
[0120] In some embodiments, the method further includes:
[0121] Receiving a first progress query request sent by the application side; wherein, the first progress query request carries the task identifier of the second query task;
[0122] Based on the task identifier of the second query task, determining the query progress information of the second query task from the preset storage space;
[0123] Sending the query progress information of the second query task to the application side.
[0124] In the embodiments of the present disclosure, when the query application server receives a first progress query request sent by the application side, it can parse the task identifier of the second query task from the first progress query request, and then determine the query progress information of the second query task from the preset storage space based on the task identifier of the second query task, and send the query progress information of the second query task to the application side. It should be noted that the first progress query request carries the task identifier of the second query task, and the query application server determines the corresponding query progress information from the preset storage space according to the task identifier of the second query task and returns it to the application side. Among them, the first progress query request refers to a request sent by the application side for obtaining the query progress information of the second query task being processed, and this request can be sent at any time during the processing of the second query task.
[0125] It should be noted that the second query task is a query task that already exists in the task queue, which can be a new query task or a historical query task, and no specific limitation is made here.
[0126] In the embodiments of the present disclosure, the query progress information is saved in the preset storage space instead of the memory of the query application server, which avoids the problem of the query progress information being lost due to the restart of the query application server. At the same time, obtaining the query progress information from the preset storage space will also reduce the query latency, thereby improving the user experience.
[0127] In some embodiments, the method further includes:
[0128] Sending a second progress query request to the query engine at a preset time interval; wherein, the second progress query request carries the task identifiers of at least one of the query tasks in the task queue;
[0129] Determining a second thread corresponding to the at least one query task from a preset thread pool;
[0130] Respectively obtaining, through each of the second threads, the query progress information of each of the query tasks obtained based on the task identifiers of the at least one query task from the query engine;
[0131] Updating the query progress information of each of the query tasks to the preset storage space.
[0132] In the embodiments of the present disclosure, the query application server will send a second progress query request to the query engine at a preset time interval. Among them, the preset time interval can be set according to actual applications and requirements. For example, the time interval can be set to 2 seconds or 5 seconds. As Figure 2As shown, the query application server calls the second thread in the preset thread pool 203 (thread pool 2) to process various tasks in the task queue at regular preset time intervals. It should be noted that the second progress query request carries the task identifiers of at least one of the query tasks in the task queue, where the second progress query request is used to obtain the query progress information of the query task corresponding to the task identifier it carries.
[0133] In the embodiments of the present disclosure, when the query application server sends the second progress query request to the query engine, it can determine the second thread corresponding to at least one query task from each of the to-be-called threads in the preset thread pool. After determining the second thread corresponding to the second progress query request, the second progress query request can be processed by the second thread. For example, the second thread can send the second progress query request to the query engine. After receiving the second progress query request, the query engine can obtain the query progress information of the corresponding query task and update it to the preset storage space.
[0134] In some embodiments, when there are multiple task identifiers carried by the second progress query request, there are correspondingly multiple query tasks, and thus multiple second threads can be determined. Of course, it is also possible to process multiple query tasks based on one second thread, as long as it can be realized that at least one query task is processed by the second thread and the query progress information of each query task can be obtained, and specific limitations are not made here.
[0135] In the embodiments of the present disclosure, the query application server actively polls and obtains the query progress information using the task queue, driving the state transition of the entire query task to the final state. It is not necessary to allocate a thread for each query task, and the state transition of multiple query tasks can be managed by a single thread, solving the problem of thread resource occupation by large queries, releasing unnecessary memory occupation, and avoiding memory overflow problems.
[0136] In some embodiments, the method further includes:
[0137] In the case of successfully obtaining the query progress information of the first query task, updating the first query task based on the query progress information of the first query task to obtain a third query task;
[0138] Updating the first query task in the task queue to the third query task.
[0139] In the embodiments of the present disclosure, after the query application server obtains the query progress information of the first query task, the query application server will update a series of new query tasks based on the first query task and the query progress information of the first query task, which are called third query tasks.
[0140] In the embodiments of the present disclosure, based on the first query task and the corresponding query progress information, the automatic update of the query task can be realized, the self-driving of the query task can be achieved, and at the same time, the application side can obtain the real-time query progress information of the query task, improving the user experience.
[0141] In some embodiments, the method further includes:
[0142] Obtaining the query progress information of the query tasks successfully completed from the query engine;
[0143] Storing the query results of the successfully completed query tasks in a preset storage space, and sending the query progress information of the successfully completed query tasks to the application side.
[0144] In the embodiments of the present disclosure, after the query application server obtains the query progress information of the successfully completed query tasks from the query engine, the query application server obtains the query results of the corresponding query tasks from the query engine and saves the query results in the preset storage space instead of the memory of the query application server, which can avoid memory overflow caused by excessive memory occupation. At the same time, obtaining the query results from the preset storage space can also reduce the query latency, thereby improving the user experience.
[0145] While storing the query results of the successfully completed query tasks in the preset storage space, sending the query progress information of the successfully completed query tasks to the application side can enable the application side to timely understand that the query tasks have been completed, and the user can perform corresponding operations based on this. For example, the query of the query results can be performed, etc.
[0146] In some embodiments, the storing the query results of the successfully completed query tasks in the preset storage space includes:
[0147] Determining the successfully completed query tasks based on the query progress information;
[0148] Determining a third thread corresponding to the successfully completed query task from each of the to-be-called threads in the preset thread pool;
[0149] Using the third thread to obtain the query results of the successfully completed query tasks from the query engine in a reactive stream processing manner, and storing the query results of the successfully completed query tasks in the preset storage space.
[0150] Figure 3 It is a processing flow chart of reactive stream processing shown according to an exemplary embodiment, as Figure 3As shown in the figure, in the embodiments of the present disclosure, while determining the successfully completed query tasks based on the query progress information, the query application server can determine a third thread corresponding to the successfully completed query tasks from each of the threads to be called in the preset thread pool. After determining the third thread corresponding to the successfully completed query tasks, the query result of the successfully completed query tasks can be processed through the third thread. For example, the query application server obtains the query result of the successfully completed query tasks from the query engine through the reactive stream processing method and stores the query result of the successfully completed query tasks in the preset storage space of the query application server. It should be noted that the preset thread pool can be implemented using an unbounded thread pool.
[0151] It should be noted that the third thread is used to obtain the query result of the successfully completed query tasks from the query engine and perform persistent storage, and at the same time, multiplexing technology can be used for parallel processing, thereby further improving the utilization rate of thread resources. Among them, persistent storage includes methods such as databases and files, such as implementation methods of Oracle, DB2, PostgreSQL, Microsoft SQL Server, Microsoft Access, MySQL, Inspur K-DB, mongodb, cassandra, redis, hbase, neo4j, CSV files, etc.
[0152] It should be noted that the reactive stream processing is used to implement the streaming reading and writing of query results. It can be considered to use the Json Stream (such as ndjson) method to store the query results, and this method is also optimal for retaining data types.
[0153] Specifically, to enable the query results to be readable by any node in the query application server cluster deployment environment, a distributed file storage solution (such as object storage, distributed file system) can be used to implement it.
[0154] In the embodiments of the present disclosure, when the query task is successfully completed, the third thread is used to obtain the query result from the query engine through the reactive stream processing method and save it to the preset storage space. Since the I / O time consumption is much greater than the CPU time consumption during the query result acquisition process, the advantages of reactive stream processing can be fully utilized, which can reduce the occupation of thread resources. At the same time, since the query results are written to disk in a timely manner, it can also effectively reduce the occupation of memory resources of the query application server and avoid memory overflow problems.
[0155] In some embodiments, the method further includes:
[0156] Receiving a first result query request sent by the application side; wherein, the task identifier of the third query task is carried in the first result query request;
[0157] Determine the query result of the third query task from the preset storage space based on the task identifier of the third query task.
[0158] In the embodiments of the present disclosure, after the query application server receives the first result query request sent by the application side, it determines the query result of the third query task from the preset storage space of the query application server based on the task identifier of the third query task carried in the first result query request. Among them, the first result query request is used for the application side to obtain the query result from the query application server in a reactive stream processing manner after the query task is successfully completed.
[0159] In the embodiments of the present disclosure, when the query task is successfully completed, the application side sends a first result query request to the query application server to obtain the query result. In order to decouple the result acquisition operation on the application side and the result acquisition operation on the query application server and to promptly store the query result to disk, the producer - consumer mode is adopted to implement the production and consumption of the query result data stream, reducing the occupation of the memory resources of the query application server and avoiding the memory overflow problem.
[0160] Figure 4 is a schematic diagram of a data - processing scenario shown according to an exemplary embodiment, as Figure 4 shown, the data - processing scenarios in which the application side 401 participates include:
[0161] Submit a query: The application side sends a first query task to the query application server through the front - end to query specified data. Here, the front - end can be a visualization interface of various electronic devices, such as a mobile phone browser page, a computer browser page, a smart - glasses browser page, etc.
[0162] Obtain the task progress in real time: The application side sends a first progress query request to the query application server through the front - end. The query application server actively polls the query progress information of the query task and returns the query progress information to the application side in real time.
[0163] Preview and download the result: The application side sends a first result query request to the query application server through the front - end to obtain the query result from the query application server.
[0164] Cancel the query: The application side sends a cancel query request to the query application server through the front - end, and the query application server ends the processing of the query task.
[0165] Figure 5 is an interaction flowchart of a data - processing method shown according to an exemplary embodiment, as Figure 5 shown, the method mainly includes the following steps:
[0166] In step 501, the query application server sends the first query task to the query engine through the first thread.
[0167] Here, the first thread is the thread determined by the query application server from each of the threads to be called in the preset thread pool for processing the first query task.
[0168] In step 502, the query application server simultaneously adds the first query task to the task queue.
[0169] Here, the query application server can first send the first query task to the query engine, or first add the first query task to the task queue, or perform the above two operations in parallel, and no specific limitation is made here.
[0170] In step 503, the query application server uses the second thread to obtain the first query task from the task queue.
[0171] Here, the second thread is the thread determined by the query application server from each of the threads to be called in the preset thread pool for processing the first query task, which can be the same as the first thread or different from the first thread.
[0172] In step 504, the query application server obtains the query progress information of the first query task from the query engine based on the task identifier of the first query task through the second thread, and updates the query progress information to the preset storage space.
[0173] In step 505, the query application server determines whether the first query task is successfully completed according to the obtained query progress information.
[0174] Here, if the first query task is not successfully completed, the first query task is updated based on the query progress information of the first query task to obtain a third query task, and the third query task is added to the task queue, and steps 502 to 505 are executed again.
[0175] In step 506, after the first query task is successfully completed, the query application server determines a third thread corresponding to the first query task from the preset thread pool, and uses the third thread to obtain the query result of the first query task from the query engine through the reactive stream processing method, and stores the query result in the preset storage space.
[0176] Figure 6 is an interaction timing diagram of a data processing method shown according to an exemplary embodiment, as Figure 6 shown, the method mainly includes the following steps:
[0177] In step 601, the application side submits a first query task to the front end;
[0178] In step 602, the front end sends the first query task submitted by the application side to the query application server;
[0179] In step 603, the query application server sends the first query task to the query engine through the first thread and adds the first query task to the task queue;
[0180] Here, the query application server can first send the first query task to the query engine, or first add the first query task to the task queue, or perform the above two operations in parallel, which is not specifically limited here.
[0181] In step 604, after receiving the first query task, the query engine immediately returns the query progress information of the first query task to the query application server through an asynchronous non-blocking interface;
[0182] In step 605, the query application server sends the returned query progress information to the front end;
[0183] In step 606, the front end returns the received query progress information to the application side;
[0184] In step 607, the query application server obtains the query progress information of the first query task from the query engine based on the task identifier of the first query task through the second thread;
[0185] In step 608, the query engine returns the query progress information to the query application server, and the query application server updates the query progress information to the preset storage space;
[0186] In step 609, before the first query task is successfully completed, the query application server continuously pushes the query progress information of the first query task to the front end;
[0187] In step 610, the front end returns the received query progress information to the application side;
[0188] In step 611, after the first query task is successfully completed, the query application server obtains the query result of the first query task from the query engine through the third thread using the reactive stream processing method;
[0189] In step 612, the query engine returns the query result of the first query task to the query application server using the reactive stream processing method, and the query application server stores the query result of the first query task in the preset storage space;
[0190] In step 613, the query application server sends the query progress information indicating that the first query task has been completed to the front end;
[0191] In step 614, the front end returns the query progress information indicating that the received first query task has been completed to the application side;
[0192] In step 615, the application side sends a first result query request to the front end;
[0193] In step 616, the front end sends the first result query request sent by the application side to the query application server;
[0194] In step 617, the query application server returns the query result of the first query task to the front end in a manner of reactive stream processing;
[0195] In step 618, the application side receives and views the query result returned by the query application server;
[0196] The technical solution of the present disclosure can add the query task received from the application side to the task queue of the query application server, and the query application server drives the processing of the entire query task, that is, all tasks in the task queue are processed by each pending thread in the preset thread pool. This enables the query application server not to continuously occupy the application thread until the task is completed after receiving the first query task, but can add the first query task to the task queue and determine the first thread corresponding to the first query task from each pending thread in the preset thread pool, and the first thread processes the first query task. Each pending thread in the present disclosure can be used to process any query task in the task queue. Compared with the existing JDBC-based blocking task processing method, the 1:N (thread: task) task processing method in this solution can greatly improve the capacity of the query service, and the service capacity is increased by more than a hundred times compared with the blocking task processing method, reducing the occupation of thread resources by large queries.
[0197] Figure 7 It is a block diagram of a data processing device shown according to an exemplary embodiment. As Figure 7 shown, the data processing device 700 mainly includes:
[0198] A first receiving module 701, configured to add the first query task to the task queue when receiving the first query task; wherein, the task queue is used to cache each received and updated query task;
[0199] A first determining module 702, configured to determine a first thread corresponding to the first query task from each pending thread in the preset thread pool; wherein, the preset thread pool includes: at least two of the pending threads;
[0200] A first sending module 703, configured to send the first query task to the query engine through the first thread;
[0201] The first acquisition module 704 is configured to acquire the query progress information of the first query task returned by the query engine.
[0202] In some embodiments, the apparatus 700 further includes:
[0203] The second sending module is configured to, when the first query task is received for the first time, if the query progress information of the first query task is successfully acquired, send the query progress information of the first query task to the application side.
[0204] In some embodiments, the apparatus 700 further includes:
[0205] The association module is configured to, when the query progress information of the first query task is acquired from the query engine, establish an association relationship between the task identifier of the first query task and the first query task; wherein, the task identifier is: an identifier preset for uniquely identifying the query task;
[0206] The first update module is configured to update the query progress information of the first query task to a preset storage space based on the association relationship;
[0207] wherein, the preset storage space is used to store the query progress information corresponding to each of the query tasks.
[0208] In some embodiments, the apparatus 700 further includes:
[0209] The second receiving module is configured to receive a first progress query request sent by the application side; wherein, the first progress query request carries the task identifier of the second query task;
[0210] The second determination module is configured to determine the query progress information of the second query task from the preset storage space based on the task identifier of the second query task;
[0211] The third sending module is configured to send the query progress information of the second query task to the application side.
[0212] In some embodiments, the apparatus 700 further includes:
[0213] The fourth sending module is configured to send a second progress query request to the query engine at a preset time interval; wherein, the second progress query request carries the task identifiers of at least one of the query tasks in the task queue;
[0214] The third determination module is configured to determine a second thread corresponding to the at least one query task from a preset thread pool;
[0215] A second acquisition module, configured to respectively acquire, via each of the second threads, query progress information of each of the query tasks obtained based on the task identifier of the at least one query task from the query engine;
[0216] A second update module, configured to update the query progress information of each query task to a preset storage space.
[0217] In some embodiments, the apparatus 700 further includes:
[0218] A third update module, configured to, in the case of successfully acquiring the query progress information of the first query task, update the first query task based on the query progress information of the first query task to obtain a third query task;
[0219] A fourth update module, configured to update the first query task in the task queue to the third query task.
[0220] In some embodiments, the apparatus 700 further includes:
[0221] A third acquisition module, configured to acquire query progress information of a query task that has been successfully completed from the query engine;
[0222] A storage module, configured to store the query result of the query task that has been successfully completed in a preset storage space, and send the query progress information of the query task that has been successfully completed to the application side.
[0223] In some embodiments, the storage module is configured to:
[0224] Determine a query task that has been successfully completed based on the query progress information;
[0225] Determine a third thread corresponding to the query task that has been successfully completed from each of the to-be-called threads in the preset thread pool;
[0226] Use the third thread to obtain the query result of the query task that has been successfully completed from the query engine in a reactive stream processing manner, and store the query result of the query task that has been successfully completed in the preset storage space.
[0227] In some embodiments, the apparatus 700 further includes:
[0228] A third receiving module, configured to receive a first result query request sent by the application side; wherein, the task identifier of the third query task is carried in the first result query request;
[0229] A fourth determination module, configured to determine the query result of the third query task from the preset storage space based on the task identifier of the third query task.
[0230] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0231] Figure 8 FIG. 7 is a block diagram of the hardware structure of an electronic device 800 shown according to an exemplary embodiment. For example, the device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0232] Refer to Figure 8 , the device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0233] The processing component 802 generally controls the overall operation of the device 800, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0234] The memory 804 is configured to store various types of data to support the operation of the device 800. Examples of these data include instructions for any application or method operating on the device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0235] The power component 806 provides power to various components of the device 800. The power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 800.
[0236] The multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0237] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0238] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.
[0239] The sensor component 814 includes one or more sensors for providing an assessment of the state of the device 800 in various aspects. For example, the sensor component 814 can detect the on / off state of the device 800, the relative positioning of components, such as the display and the keypad of the device 800. The sensor component 814 can also detect a change in the position of the device 800 or a component of the device 800, the presence or absence of user contact with the device 800, the orientation or acceleration / deceleration of the device 800, and the temperature change of the device 800. The sensor component 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 814 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0240] The communication component 816 is configured to facilitate communication between the device 800 and other devices in a wired or wireless manner. The device 800 may access a communication standard-based wireless network, such as WiFi, 4G, or 5G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0241] In an exemplary embodiment, the device 800 may be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0242] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the above instructions can be executed by a processor 820 of the device 800 to complete the above method. For example, the non-transitory computer-readable storage medium may be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0243] A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of a mobile terminal, enables the mobile terminal to execute a data processing method, and the method includes:
[0244] In the case of receiving a first query task, adding the first query task to a task queue; wherein, the task queue is used to cache each received and updated query task;
[0245] Determine a first thread corresponding to the first query task from each of the threads to be called in a preset thread pool; wherein, the preset thread pool includes: at least two of the threads to be called;
[0246] Send the first query task to a query engine through the first thread;
[0247] Obtain query progress information of the first query task returned by the query engine.
[0248] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0249] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A data processing method, characterized in that, Including: When receiving a first query task, adding the first query task to a task queue; wherein, the task queue is used to cache each received and updated query task; Determining a first thread corresponding to the first query task from each to-be-called thread in a preset thread pool; wherein, the preset thread pool includes at least two to-be-called threads; Sending the first query task to a query engine through the first thread; Obtaining query progress information of the first query task returned by the query engine.
2. The method according to claim 1, characterized in that, The method further includes: When first receiving the first query task, if successfully obtaining the query progress information of the first query task, sending the query progress information of the first query task to the application side.
3. The method according to claim 1, wherein The method further includes: When obtaining the query progress information of the first query task from the query engine, establishing an association relationship between a task identifier of the first query task and the first query task; wherein, the task identifier is a preset identifier for uniquely identifying the query task; Updating the query progress information of the first query task to a preset storage space based on the association relationship; Wherein, the preset storage space is used to store query progress information corresponding to each query task.
4. The method according to claim 3, characterized in that The method further includes: Receiving a first progress query request sent by the application side; wherein, the first progress query request carries a task identifier of a second query task; Determining the query progress information of the second query task from the preset storage space based on the task identifier of the second query task; Sending the query progress information of the second query task to the application side.
5. The method according to claim 1, wherein The method further includes: Sending a second progress query request to the query engine at a preset time interval; wherein, the second progress query request carries task identifiers of at least one query task in the task queue; Determining second threads corresponding to the at least one query task from the preset thread pool; Respectively obtaining, through each second thread, query progress information of each query task obtained based on the task identifiers of the at least one query task from the query engine; Updating the query progress information of each query task to the preset storage space.
6. The method according to claim 1, characterized in that, The method further includes: When successfully obtaining the query progress information of the first query task, updating the first query task based on the query progress information of the first query task to obtain a third query task; Updating the first query task in the task queue to the third query task.
7. The method according to claim 1, characterized in that The method further includes: Obtaining query progress information of a successfully completed query task from the query engine; Storing the query result of the successfully completed query task to the preset storage space, and sending the query progress information of the successfully completed query task to the application side.
8. The method according to claim 7, wherein The storing the query result of the successfully completed query task to the preset storage space includes: Determining the successfully completed query task based on the query progress information; Determine a third thread corresponding to the successfully completed query task from each of the to-be-invoked threads in the preset thread pool; Utilize the third thread to obtain the query result of the successfully completed query task from the query engine in a reactive stream processing manner, and store the query result of the successfully completed query task in the preset storage space.
9. The method according to claim 7, wherein The method further includes: Receiving a first result query request sent by the application side; wherein, the task identifier of the third query task is carried in the first result query request; Based on the task identifier of the third query task, determine the query result of the third query task from the preset storage space.
10. A data processing device, characterized in that, Including: A first receiving module, configured to add the first query task to the task queue when the first query task is received; wherein, the task queue is used to cache each received and updated query task; A first determining module, configured to determine a first thread corresponding to the first query task from each of the to-be-invoked threads in the preset thread pool; wherein, the preset thread pool includes at least two to-be-invoked threads; A first sending module, configured to send the first query task to the query engine through the first thread; A first obtaining module, configured to obtain the query progress information of the first query task returned by the query engine.
11. An electronic device, characterized in that, Including: A processor; A memory configured to store processor-executable instructions; Wherein, the processor is configured to: when executed, implement the steps in any one of the data processing methods in claims 1 to 9 above.
12. A non-transitory computer-readable storage medium, characterized in that When the instructions in the storage medium are executed by the processor of the data processing device, the device is enabled to execute the steps in any one of the data processing methods in claims 1 to 9 above.