Data processing method and device, electronic equipment and computer storage medium

By using loading threads and processing threads to process data loading and parallel processing respectively, and managing consumption data in sequence with data identification, the problem of excessive system resource occupation in the prior art is solved, and efficient and orderly data processing is achieved.

CN120179428APending Publication Date: 2025-06-20BEIJING VOYAGER TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311757993.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

While ensuring orderly processing of data, existing data concurrency processing methods lead to a large amount of system resources occupied, which in turn has an adverse impact.

Method used

By loading the pending data based on the loading thread and processing the data in parallel based on the processing thread, the consumption data is determined, and consumption and queue management are carried out according to the order of data identification, so as to avoid processing before starting after all data is loaded.

Benefits of technology

It reduces the large amount of system resources, improves data processing efficiency, and ensures orderly processing of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179428A_ABST
    Figure CN120179428A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data processing method and device, electronic equipment and a computer storage medium, and the method comprises the steps: loading at least one piece of to-be-processed data based on a loading thread, carrying out the parallel processing of the at least one piece of to-be-processed data based on at least one processing thread, and determining the corresponding consumption data, the to-be-processed data has a data identifier representing a loading sequence, and the consumption data has a data identifier consistent with the corresponding to-be-processed data. Therefore, data loading and data processing are completed through different threads, the situation that data processing can be started only when all data are loaded is avoided, and a large amount of occupied system resources can be reduced. And meanwhile, the to-be-processed data and the corresponding consumption data are identified through the data identifier, so that ordered processing of the data can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technologies, and particularly to a data processing method, apparatus, electronic device, and computer storage medium. Background Art

[0002] When dealing with a large amount of data, in order to process the data more quickly, multi-process or multi-thread methods are mostly used for parallel processing, and at the same time, it is hoped that the result after processing can still maintain the original data order. Taking the playback of an autonomous driving road test Bag as an example, it is often necessary to perform response processing on the original messages. However, since the data processing efficiency of a single process is low and cannot meet the requirements, it is necessary to use a multi-process method for data processing while ensuring the timeliness of the Bag information and still maintaining its original message order after the data processing is completed.

[0003] In existing data concurrent processing methods, in order to ensure ordered data processing, usually all the original data is first loaded and completed. After obtaining the complete list of original data, the data is then processed concurrently by multiple threads, and the order during loading is used as the order during consumption. However, this method will cause a large amount of system resources to be occupied, and further produce more adverse effects. Summary of the Invention

[0004] In view of this, an object of the embodiments of the present invention is to provide a data processing method to reduce the large occupation of system resources.

[0005] In a first aspect, an embodiment of the present invention aims to provide a data processing method, the method comprising:

[0006] Loading at least one data to be processed based on a loading thread, the data to be processed having a data identifier representing the loading order;

[0007] Performing parallel processing on at least one of the data to be processed based on at least one processing thread to determine corresponding consumed data, the consumed data having the same data identifier as the corresponding data to be processed.

[0008] Further, the method further comprises:

[0009] Consuming each of the consumed data according to the order of the data identifiers.

[0010] Further, the method further comprises:

[0011] Adding the consumed data to a queue, and the extraction order of each of the consumed data in the queue is determined according to the order of the data identifiers.

[0012] Further, the consuming each of the consumed data according to the order of the data identifiers includes:

[0013] Determine a target identifier, where the target identifier is used to represent the data identifier of the current data to be consumed;

[0014] Determine the corresponding consumed data from the queue according to the target identifier;

[0015] Consume the consumed data corresponding to the target identifier based on a consumption thread;

[0016] Update the target identifier.

[0017] Further, the method further includes:

[0018] In response to the consumed data being added to the queue, send a data ready notification.

[0019] Further, the determining the corresponding consumed data from the queue according to the target identifier includes:

[0020] In response to receiving the data ready notification, determine the corresponding consumed data from the queue according to the target identifier.

[0021] Further, the method further includes:

[0022] In response to all the consumed data being added to the queue, send a data completion notification.

[0023] Further, the consuming the consumed data according to the order of the data identifiers includes:

[0024] In response to receiving the data completion notification, consume the consumed data in the queue in sequence according to the order of the data identifiers based on a consumption thread.

[0025] Further, the loading of at least one data to be processed based on a loading thread includes:

[0026] Obtain the available status of system resources;

[0027] In response to there being available system resources, load at least one data to be processed based on a loading thread.

[0028] Further, the method further includes:

[0029] In response to the consumed data having been consumed, release the corresponding system resources.

[0030] Further, the method further includes:

[0031] Control the progress of the loading and the consumption through a resource controller.

[0032] Further, the queue is a minimum heap queue.

[0033] In a second aspect, an embodiment of the present invention aims to provide a data processing device, the device comprising:

[0034] A loading unit, configured to load at least one data to be processed based on a loading thread, the data to be processed having a data identifier characterizing the loading order;

[0035] A processing unit, configured to perform parallel processing on at least one of the data to be processed based on at least one processing thread to determine corresponding consumption data, the consumption data having a data identifier consistent with the corresponding data to be processed.

[0036] In a third aspect, an embodiment of the present invention aims to provide an electronic device, comprising a memory and a processor, the memory being configured to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method described in any one of the above.

[0037] In a fourth aspect, an embodiment of the present invention aims to provide a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the method steps described in any one of the above are implemented.

[0038] The technical solution of the embodiment of the present invention loads at least one data to be processed based on a loading thread, performs parallel processing on at least one of the data to be processed based on at least one processing thread to determine corresponding consumption data, so that data loading and data processing are completed by different threads, which can avoid starting data processing only after all data is loaded, and reduce a large amount of occupation of system resources. Moreover, since the data to be processed has a data identifier characterizing the loading order, and the consumption data has a data identifier consistent with the corresponding data to be processed, the data loading and consumption processes can be carried out in a certain order, thereby ensuring the orderly processing of data. Description of the Drawings

[0039] Through the following description of the embodiments of the present invention with reference to the drawings, the above and other objects, features and advantages of the present invention will become clearer. In the drawings:

[0040] Figure 1 is a flowchart of the data processing method according to the embodiment of the present invention;

[0041] Figure 2 is another flowchart of the data processing method according to the embodiment of the present invention;

[0042] Figure 3 is a flowchart of the data consumption method according to the embodiment of the present invention;

[0043] Figure 4 is a schematic diagram of the data processing method according to the embodiment of the present invention;

[0044] Figure 5 It is a schematic diagram of the data processing device according to an embodiment of the present invention;

[0045] Figure 6 It is another schematic diagram of the data processing device according to an embodiment of the present invention;

[0046] Figure 7 It is a schematic diagram of the electronic device according to an embodiment of the present invention. Detailed implementation manners

[0047] The following describes the present application based on embodiments, but the present application is not limited to these embodiments. In the following detailed description of the present application, some specific details are described in detail. Those skilled in the art can fully understand the present application without the description of these details. In order to avoid obscuring the essence of the present application, well-known methods, processes, procedures, components, and circuits are not described in detail.

[0048] In addition, those of ordinary skill in the art should understand that the drawings provided herein are all for illustrative purposes and are not necessarily drawn to scale.

[0049] Unless the context clearly requires otherwise, words such as "including" and "comprising" in the entire application document should be interpreted as having an inclusive meaning rather than an exclusive or exhaustive meaning; that is, it is the meaning of "including but not limited to".

[0050] In the description of the present application, it should be understood that terms such as "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.

[0051] For the solutions described in this specification and the embodiments, if they involve personal information processing, they will be processed on the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for performing a contract, etc.), and will only be processed within the specified or agreed scope. If the user refuses to process personal information other than the necessary information required for the basic functions, it will not affect the user's use of the basic functions.

[0052] During the existing data concurrent processing, since data loading and data processing are within one thread, before data processing starts, all data needs to be loaded, and the fully loaded data will occupy a large amount of system resources. At the same time, after data processing is completed, if a piece of data at the front has not been processed or the downstream fails to process data in a timely manner, a large amount of subsequent processed data will not be consumed in a timely manner due to waiting for the previous data, which will also occupy a large amount of system resources. In view of this, embodiments of the present invention aim to provide a data processing method to reduce the large occupation of system resources.

[0053] Figure 1 is a flowchart of the data processing method in the embodiments of the present invention. As Figure 1 shown, the data processing method in this embodiment includes the following steps.

[0054] In step S110, at least one data to be processed is loaded based on a loading thread.

[0055] In this embodiment, the data to be processed has a data identifier representing the loading order, so as to reflect the loading order of each data to be processed through the data identifier. At the same time, during the data loading process, data is loaded through a loading thread. Among them, data loading refers to the process of obtaining and importing the original data from a data source (such as a database, a file, an API, etc.) into a data analysis tool, platform or environment. Data loading is usually the starting point of data analysis, which provides the basis for data processing and consumption. The loaded data can be structured data, text data, image data, etc., and the specific loading method depends on the data source and the tools used.

[0056] Optionally, the data identifier in this embodiment can adopt a digital number or other identifier forms, arranged in ascending order, corresponding to the loading order from first to last. Taking the digital number as an example, the data identifier is set in the order of 1, 2,..., N (N is a natural number). For each data to be processed loaded, the data identifier is incremented by 1. Based on this, the data to be processed 1, the data to be processed 2,..., and the data to be processed N can be determined in sequence. Thus, by setting the above data identifier to represent the loading order, the data loading process becomes more orderly.

[0057] Optionally, since data loading requires a certain amount of system resources, when at least one data to be processed is loaded based on a loading thread, the available state of the system resources will be obtained first, and in response to the existence of available system resources, at least one data to be processed will be loaded based on the loading thread. Among them, the system resources can be resources such as CPU resources, storage resources, network resources, and node resources. Thus, through the above method, it can be ensured that data loading is carried out under the condition of available system resources, avoiding data loading interruption caused by insufficient system resources, ensuring the smooth progress of data loading, and thus being beneficial to improving the overall data processing efficiency.

[0058] In step S120, at least one piece of data to be processed is processed in parallel based on at least one processing thread, and corresponding consumption data is determined.

[0059] In this embodiment, different threads are used for the data processing thread and the loading thread. The data processing thread processes the loaded data to be processed to determine the consumption data corresponding to the data to be processed. Thus, in this embodiment, as long as there is data to be processed, the data can be processed by the processing thread without waiting for all the data to be processed to be loaded before starting the data processing, which can greatly reduce the occupation of system resources.

[0060] Optionally, the data processing in this embodiment includes operations such as data cleaning, data conversion, data integration, and data preparation on the loaded data to be processed, aiming to transform the original data into a more valuable and usable form for subsequent analysis, modeling, and application based on the consumption data determined by the data processing. Among them, data cleaning can be operations such as handling missing values and outliers, data conversion can be operations such as format conversion and feature engineering, and data integration can be operations such as merging multiple data sources and data integration.

[0061] Meanwhile, after the data to be processed is processed in this embodiment, the obtained consumption data has the same data identifier as the corresponding data to be processed. For example, when there are data to be processed 1, data to be processed 2, …, and data to be processed N, the corresponding consumption data are consumption data 1, consumption data 2, …, and consumption data N respectively. Thus, through the consistency of the data identifiers, the correspondence between the data to be processed and the consumption data can be guaranteed, facilitating data traceability and accuracy verification, and improving the accuracy of data processing.

[0062] Furthermore, in this embodiment, at least one piece of data to be processed obtained after loading is processed in parallel through at least one processing thread to determine the consumption data corresponding to each piece of data to be processed. Thus, by setting multiple processing threads and adopting a parallel processing method for data processing, the overall process of data processing can be accelerated and the data processing efficiency can be improved.

[0063] Optionally, the number of processing threads in this embodiment can be less than or equal to the number of data to be processed, and can be specifically set according to the actual usage scenario. Thus, by flexibly setting the number of processing threads, the balance between data processing efficiency and resource occupation can be coordinated while improving the data processing efficiency, and the overall performance of data processing can be optimized.

[0064] The technical solution of the embodiment of the present invention loads at least one piece of data to be processed based on a loading thread, and performs parallel processing on at least one piece of the data to be processed based on at least one processing thread to determine corresponding consumed data, so that data loading and data processing are completed by different threads, which can avoid starting data processing only after all data is loaded, and reduce the large occupation of system resources. Moreover, since the data to be processed has a data identifier representing the loading order, and the consumed data has the same data identifier as the corresponding data to be processed, the data loading and consumption processes can be carried out in a certain order, thereby ensuring the orderly processing of data.

[0065] Figure 2 is another flowchart of the data processing method of the embodiment of the present invention. As Figure 2 shown, the data processing method in this embodiment includes the following steps:

[0066] In step S210, at least one piece of data to be processed is loaded based on a loading thread.

[0067] In this embodiment, the method of loading at least one piece of data to be processed based on a loading thread is the same as the method introduced in the foregoing step S110, and will not be elaborated here.

[0068] In step S220, at least one piece of data to be processed is subjected to parallel processing based on at least one processing thread to determine corresponding consumed data.

[0069] In this embodiment, the method of determining the consumed data corresponding to the data to be processed is the same as the method introduced in the foregoing step S120, and will not be elaborated here.

[0070] In step S230, each consumed data is consumed according to the order of the data identifiers.

[0071] In this embodiment, the order of data identifiers represents the order of data consumption. When consuming data, each piece of consumed data is consumed in the order of data identifiers. For example, when there are consumed data 1, consumed data 2, …, and consumed data N that need to be consumed, the data consumption order is successively consumed data 1, consumed data 2, …, to consumed data N, which is consistent with the order at the time of data loading. The data loaded first is consumed first, and the data loaded later is consumed later. That is, the consumed data with a smaller data identifier is consumed first, and the consumed data with a larger data identifier can only be consumed after the consumed data corresponding to the smaller data identifier has been consumed. Thus, by consuming each piece of consumed data according to the order of data identifiers, the orderly completion of data consumption can be ensured. At the same time, since during the processes of data loading, data processing, and data consumption, the data to be processed and the consumed data corresponding to the same data are identified by the same data identifier, the consistency of the data source can be ensured, facilitating understanding the progress of the overall data processing process and being conducive to improving the overall data processing efficiency.

[0072] Optionally, data consumption in this embodiment refers to the process of extracting useful information, performing analysis, applying modeling, or making decisions from the processed data. Data consumption can include tasks such as statistical analysis, machine learning modeling, data visualization, report generation, prediction, and recommendation. Data consumption is the ultimate goal of data processing. By consuming the processed data, insights into business problems, model prediction results, business decision-making bases, etc. can be obtained. Thus, through data consumption operations, the data can be made to play a greater value and provide a reliable basis for subsequent analysis and decision-making.

[0073] Optionally, to further reduce the occupation of system resources in the overall data processing process, in this embodiment, the consumed data determined by parallel data processing is managed in the form of a queue. The order of taking out each piece of consumed data in the queue is determined according to the order of data identifiers.

[0074] To improve the management efficiency of data, in this embodiment, when adding consumption data to the queue, for each piece of data to be processed, after each data processing is completed, the corresponding consumption data and data identifier are added to the queue. For example, after the data to be processed 1 is processed, the corresponding consumption data 1 is added to the queue. The consumption data is the data after the corresponding data to be processed 1 is processed, and 1 is the data identifier. At the same time, to ensure the orderliness of data consumption, in this embodiment, when retrieving consumption data from the queue, the retrieval order of each consumption data in the queue is determined according to the order of the data identifiers. For example, assuming that there are consumption data 1, consumption data 2, …, and consumption data N in the queue, when data consumption is required to retrieve consumption data from the queue, the retrieval order of the consumption data is consumption data 1, consumption data 2, …, and consumption data N, that is, the consumption data with a smaller data identifier is retrieved first, and the consumption data with a larger data identifier can only be retrieved after the consumption data corresponding to the smaller data identifier is retrieved.

[0075] Furthermore, the queue in this embodiment can use a min-heap queue or other data structures to ensure the order of data. Among them, the min-heap queue is a special queue, which is implemented based on a min-heap. A min-heap is a complete binary tree, where the value of each node is less than or equal to the values of its children. In a min-heap queue, the element at the head of the queue is the minimum value among all elements.

[0076] Specifically, the implementation of a min-heap queue can be based on an array, where the position of each node is determined by its index in the array. When implementing a min-heap queue, it usually includes two main operations: insertion and deletion. The insertion operation is used to add a new element to the queue. When adding a new element, the new element is first added to the end of the array, and then it is adjusted to a min-heap. To maintain the properties of the heap, it is necessary to start from the new element and adjust the heap downward. The deletion operation is used to delete the minimum element from the queue. When deleting the minimum element, the first element of the array (i.e., the minimum element) is first deleted, and then the last element of the array is moved to the first position of the array, and then the heap is adjusted downward to restore the properties of the heap. In a min-heap queue, the time complexity of both the insertion and deletion operations is O(log n), where n is the number of elements in the queue. Due to the complete binary tree property of the min-heap, it is possible to quickly find the subscript of the left and right children of any node through the array and formula, improving the efficiency of the min-heap queue when processing a large amount of data.

[0077] Furthermore, the number of elements that can be tolerated in the minimum heap queue in this embodiment (i.e., the number of consumed data) can be set according to experience and actual usage scenarios. For example, when processing and consuming the data to be processed in batch loading, the number of consumed data in the minimum heap queue can be the same as the number of data included in each batch in the batch loading. Thus, by flexibly setting the number of consumed data that can be tolerated in the minimum heap queue, the setting and use of the minimum heap queue can be made more convenient, which is conducive to further improving the overall data processing efficiency.

[0078] Furthermore, in this embodiment, the consumed data is taken out from the minimum heap queue for data consumption. Specifically, when consuming each of the consumed data according to the order of the data identifiers, this embodiment is implemented through the following steps as shown in Figure 3 the following.

[0079] In step S310, a target identifier is determined.

[0080] In this embodiment, the target identifier is used to represent the data identifier of the currently to-be-consumed data.

[0081] In step S320, the corresponding consumed data is determined from the queue according to the target identifier.

[0082] In this embodiment, after determining the data identifier of the currently to-be-consumed data, a query is made in the queue according to the target identifier to determine whether there is consumed data corresponding to the target identifier in the queue. When there is consumed data corresponding to the target identifier in the queue, the consumed data corresponding to the target identifier is locked to facilitate subsequent data consumption. Specifically, assume that the current target identifier is 1, then the corresponding to-be-consumed data is consumed data 1. At this time, check whether there is consumed data 1 in the queue, and when there is consumed data 1 with the data identifier 1 in the queue, determine consumed data 1 as the consumed data corresponding to the target identifier 1.

[0083] Optionally, to facilitate determining the consumed data of the target identifier from the queue, in this embodiment, a data ready notification is sent in response to the addition of consumed data to the queue; and, in response to receiving the data ready notification, the corresponding consumed data is determined from the queue according to the target identifier. That is, a data ready notification is sent each time a consumed data is added to the queue, and receiving the data ready notification is used as the trigger condition for determining the consumed data corresponding to the target identifier from the queue.

[0084] In an alternative implementation, the data ready notification may include the data identifier of the consumption data currently enqueued. When determining the corresponding consumption data from the queue based on the target identifier, the data identifier in the data ready notification can be first matched with the target identifier. When the data identifier in the data ready notification is consistent with the data identifier corresponding to the target identifier, it indicates that the consumption data corresponding to the target identifier already exists in the queue. At this time, the consumption data corresponding to the target identifier can be determined from the queue. Alternatively, when the data identifier in the data ready notification is inconsistent with the target identifier, stop determining the corresponding consumption data from the queue based on the target identifier and continue to wait for the next data ready notification. Thus, by adding the data identifier in the data ready notification, when determining the corresponding consumption data based on the data identifier, only the data identifier needs to be matched to determine the corresponding consumption data, which can reduce the workload of data search and is beneficial to further improving the data processing efficiency.

[0085] In another alternative implementation, the data ready notification does not include the data identifier of the consumption data enqueued. When determining the corresponding consumption data from the queue based on the target identifier, directly search for the consumption data corresponding to the target identifier in the queue. If the consumption data corresponding to the target identifier exists, take out the consumption data from the queue for subsequent data consumption; if the consumption data corresponding to the target identifier does not exist, it indicates that the consumption data corresponding to the target identifier has not been generated yet. At this time, stop the search and continue to wait until the above operation is repeated after receiving the next data ready notification.

[0086] In step S330, consume the consumption data corresponding to the target identifier based on the consumption thread.

[0087] In this embodiment, after determining the consumption data corresponding to the current target identifier from the queue, consume the consumption data corresponding to the target identifier based on the consumption thread.

[0088] Optionally, to ensure that all the consumption data in the queue is consumed in a timely manner, in this embodiment, a data completion notification is sent in response to all the consumption data being added to the queue. At this time, it indicates that all the consumption data corresponding to the data identifiers (including the consumed and unconsumed consumption data) have been added to the current queue. The consumption data corresponding to the target identifier can be directly determined from the queue based on the target identifier, or in response to receiving the data completion notification, consume the consumption data in the queue (i.e., all the unconsumed consumption data in the queue) based on the consumption thread in the order of the data identifiers. Thus, in this embodiment, through the data completion notification, the addition situation of the consumption data in the queue can be understood in a timely manner, which is beneficial to further improving the data consumption efficiency and the overall data processing efficiency.

[0089] Further, in this embodiment, to facilitate determining whether all consumption data has been added to the queue, when adding consumption data to the queue, the identifier of the consumption data entering the queue can be recorded, so as to determine that the consumption data corresponding to each piece of data to be processed in all the data to be processed has been added to the queue according to the recorded data identifier. Alternatively, it is also possible to determine that all the data to be processed has been processed by recording the data identifier of the data to be processed in the processing thread, and after all the data to be processed on each processing line has been processed and the corresponding consumption data is generated, determine that all the consumption data has been added to the queue. Thus, by determining that all the consumption data has been added to the queue in different ways, the timing of sending the data completion notification can be set flexibly, making the operation of sending the data completion notification more convenient, and at the same time improving the processing efficiency of the consumption data in the queue.

[0090] Further, since data consumption will occupy a certain amount of system resources, to further reduce the occupation of system resources, in this embodiment, the corresponding system resources will be released in response to the consumption of the consumption data. That is, after each consumption data is consumed, the system resources occupied by the consumption data will be released. Further, in this embodiment, the release of system resources can be controlled by a resource controller. Thus, in this embodiment, by promptly releasing the system resources occupied by data consumption, unnecessary occupation of system resources can be reduced, and at the same time, system resources can be coordinated quickly, providing convenience for subsequent data consumption.

[0091] Further, to avoid the processing gap between the two sub-processes of data loading and data consumption from having an adverse impact on the overall data processing efficiency, in this embodiment, the resource controller is also used to control the progress of data loading and consumption, so as to balance the processing progress of data loading and data consumption and reduce the problem of long-term occupation of system resources caused by too long a gap time.

[0092] In step S340, update the target identifier.

[0093] In this embodiment, after the consumption data corresponding to the current target identifier enters the consumption thread for consumption or is consumed, the target identifier will be updated to facilitate subsequent data consumption.

[0094] Optionally, for the update of the target identifier, when the data identifier is set as the sequentially arranged digital number as described above, in this embodiment, the target identifier can be updated by incrementing the number of the target identifier by 1. For example, after the consumption data 1 with the target identifier of 1 is consumed, the target identifier is updated to 2 to consume the consumption data 2 in the next data consumption until all the data to be processed is consumed.

[0095] Optionally, in this embodiment, after receiving the data completion notification, the target identifier can be updated once for each consumed consumption data; alternatively, the unconsumed data in the queue can be directly consumed in sequence until all the consumption data is consumed, and the target identifier of the last consumed consumption data is incremented by 1 for update, or the overall data processing process is ended. Thus, by providing a flexible target identifier update method, the data processing requirements in different scenarios can be met; at the same time, by reducing the update frequency of the target identifier in the latter method, the overall data processing efficiency can be further improved.

[0096] The technical solution of the embodiment of the present invention loads at least one data to be processed based on a loading thread, and performs parallel processing on at least one of the data to be processed based on at least one processing thread to determine the corresponding consumption data, so that data loading and data processing are completed by different threads, which can avoid starting data processing only after all the data is loaded, reduce the large occupation of system resources, and improve the processing efficiency of the overall data processing process. Moreover, since the data to be processed has a data identifier representing the loading order, and the consumption data has the same data identifier as the corresponding data to be processed, the data loading and consumption processes can be carried out in a certain order, thereby ensuring the orderly processing of data.

[0097] Figure 4 is a schematic diagram of the data processing method of the embodiment of the present invention. As Figure 4 shown, in this embodiment, a loading thread is used to load the data to be processed. When loading the data, a data identifier is set for each data to be processed, and the data identifier can represent the loading order. For example, when setting the data identifier, starting from 1, it is sequentially 2, 3, 4, …, and so on, to obtain the data to be processed 1, the data to be processed 2, the data to be processed 3, …. Optionally, before starting the data loading, first attempt to obtain system resources from the resource controller. If there are available system resources, then start loading the data to be processed; if there are no available system resources, then wait until the system resources are released and then perform data loading. Among them, the system resources can be resources of types such as storage resources, CPU resources, network resources, and node resources.

[0098] After the data to be processed is loaded, in this embodiment, the data to be processed is processed by a processing thread different from the loading thread. The number of processing threads can be multiple, and the multiple processing threads are used to perform parallel processing on different data to be processed to determine the consumption data corresponding to each data to be processed.

[0099] Meanwhile, for each piece of data to be processed that is completed, the consumed data obtained after processing, together with the data identifier, is added to the minimum heap queue. The data identifier of the consumed data is the same as that of the corresponding data to be processed. For example, the data identifier of the consumed data corresponding to the data to be processed 3 is also 3. Optionally, after the consumed data is added to the minimum heap queue, a data ready notification is sent through a notification sender. Further, after all the data to be processed is completed or all the consumed data is added to the minimum heap queue (i.e., the consumed data corresponding to all the data to be processed has been added to the queue), a data completion notification is sent through the notification sender.

[0100] Moreover, when consuming data through a consumption thread, the consumed data is consumed in the order of the data identifiers. Specifically, first, a target identifier is set to represent the data identifier of the currently pending consumed data. After the notification receiver receives the data ready notification, it triggers the data consumption operation and checks whether there is consumed data corresponding to the target identifier in the minimum heap queue. When there is consumed data corresponding to the target identifier in the minimum heap queue, the consumed data corresponding to the target identifier is consumed. Meanwhile, the resource controller releases the resources and updates the target identifier, and continues to search for the consumed data corresponding to the current target identifier in the queue. Or, when there is no consumed data corresponding to the target identifier in the minimum heap queue, no operation is performed, and the data ready notification is waited for continuously. Or, after the notification receiver receives the data completion notification, it directly triggers the consumption thread to consume all the consumed data in the minimum heap queue sequentially.

[0101] For example, at the start of the first data consumption operation, the initial value of the target identifier (representing the data identifier of the currently pending consumed data) for data consumption is set to 1. The notification receiver is responsible for waiting for and receiving the data ready notification. After receiving the data ready notification, it searches the minimum heap queue to determine whether there is consumed data with a data identifier of 1, that is, consumed data 1, in the minimum heap queue at this time. If there is consumed data 1 in the queue, the consumed data 1 is consumed based on the consumption thread, and the target identifier is updated to 2. At the same time, 1 system resource is released through the resource controller, and then the consumed data 2 is searched for in the minimum heap queue, and so on. Conversely, if there is no consumed data 1 in the queue, the data ready notification is waited for continuously. Or, after the notification receiver receives the data completion notification, it stops searching for the consumed data and updating the target identifier, and directly takes out each consumed data from the minimum heap queue in the order of the data identifiers for consumption.

[0102] Accordingly, in this embodiment, the data is loaded, processed, and consumed by the loading thread, the processing thread, and the consuming thread respectively, without waiting for all the data to be loaded before starting data processing, which can avoid a large amount of system resource occupation. At the same time, by using the same data identifier to identify the data to be processed and the corresponding consumed data, the data loading and data consumption can be ensured to proceed in an orderly manner, thereby improving the overall data processing efficiency. Moreover, by using multiple processing threads to perform parallel processing on the data to be processed after loading to determine the corresponding consumed data, and parallel processing data loading and data consumption, the data processing efficiency can be further improved. Furthermore, by playing the role of a resource controller in data loading and data consumption to control the progress of data loading and data consumption, the coordination of data loading and data consumption can be ensured, making the overall data processing process more orderly and efficient, and thus improving the overall data processing efficiency.

[0103] Figure 5 is a schematic diagram of the data processing device according to an embodiment of the present invention. As Figure 5 shown, the data processing device in this embodiment includes: a loading unit 1, a processing unit 2, and a consuming unit 3. Among them, the loading unit 1 is configured to load at least one data to be processed based on a loading thread, and the data to be processed has a data identifier representing the loading order. The processing unit 2 is configured to perform parallel processing on at least one data to be processed based on at least one processing thread to determine the corresponding consumed data, and the consumed data has the same data identifier as the corresponding data to be processed. The consuming unit 3 is configured to consume each consumed data according to the order of the data identifiers.

[0104] Optionally, the loading unit 1 in this embodiment is further configured to obtain the available state of the system resources; in response to the existence of available system resources, load at least one data to be processed based on a loading thread.

[0105] Optionally, as Figure 6 shown, this embodiment further includes a queue unit 4, and the queue unit 4 is configured to add the consumed data to the queue. Among them, the extraction order of each consumed data in the queue is determined according to the order of the data identifiers. Further, the queue in this embodiment is a minimum heap queue.

[0106] Further, the consuming unit 3 in this embodiment is further configured to determine a target identifier, determine the corresponding consumed data from the queue according to the target identifier; consume the consumed data corresponding to the target identifier based on a consuming thread; and update the target identifier. Wherein, the target identifier is used to represent the data identifier of the currently to-be-consumed data.

[0107] Optionally, this embodiment further includes a notification unit 5, which is configured to send a data ready notification in response to the consumption data being added to the queue. Further, the consumption unit 3 is further configured to determine the corresponding consumption data from the queue according to the target identifier in response to receiving the data ready notification.

[0108] Optionally, the notification unit 5 in this embodiment is further configured to send a data completion notification in response to all the consumption data being added to the queue. Further, the consumption unit 3 is further configured to consume the consumption data in the queue in sequence according to the order of the data identifiers based on the consumption thread in response to receiving the data completion notification.

[0109] Optionally, this embodiment further includes a control unit 6, which is configured to release the corresponding system resources in response to the consumption data being consumed; and control the progress of loading and consumption.

[0110] The technical solution of the embodiment of the present invention loads at least one data to be processed by a loading unit based on a loading thread, and parallelly processes the at least one data to be processed by a processing unit based on at least one processing thread to determine the corresponding consumption data, so that data loading and data processing are completed by different threads, which can avoid starting data processing only after all data is loaded, and reduce the large occupation of system resources. Moreover, since the data to be processed has a data identifier representing the loading order, and the consumption data has the same data identifier as the corresponding data to be processed, the data loading and consumption processes can be carried out in a certain order, thereby ensuring the orderly processing of data.

[0111] Figure 7 is a schematic diagram of the electronic device according to the embodiment of the present invention. As Figure 7 shown, Figure 7 The electronic device shown is a general address query device, which includes a general computer hardware structure, and at least includes a processor 61 and a memory 62. The processor 61 and the memory 62 are connected through a bus 63. The memory 62 is adapted to store instructions or programs executable by the processor 61. The processor 61 can be an independent microprocessor or a set of one or more microprocessors. Thus, the processor 61 executes the instructions stored in the memory 62 to execute the method flow of the embodiment of the present invention as described above to implement the processing of data and the control of other devices. The bus 63 connects the above-mentioned multiple components together, and at the same time connects the above-mentioned components to a display controller 64, a display device, and an input / output (I / O) device 65. The input / output (I / O) device 65 can be a mouse, a keyboard, a modem, a network interface, a touch input device, a somatosensory input device, a printer, and other devices well known in the art. Typically, the input / output (I / O) device 65 is connected to the system through an input / output (I / O) controller 66.

[0112] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, apparatuses (devices) or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can be implemented as a computer program product on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0113] The present application is described with reference to the flowcharts of methods, apparatuses (devices) and computer program products according to the embodiments of the present application. It should be understood that each process in the flowchart can be implemented by computer program instructions.

[0114] These computer program instructions can be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that implements the process Figure 1 specified functions in one or more of these processes.

[0115] These computer program instructions can also be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the Figure 1 specified functions in one or more of these processes.

[0116] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program, and the computer-readable program is used for a computer to execute the above-mentioned partial or all method embodiments.

[0117] That is, those skilled in the art can understand that all or part of the steps in implementing the above-mentioned embodiment methods can be completed by specifying relevant hardware through a program, and the program is stored in a storage medium, including several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. And the aforementioned storage media include: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks or optical disks and other various media that can store program codes.

[0118] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A data processing method, characterized in that, The method includes: Loading at least one data to be processed based on a loading thread, where the data to be processed has a data identifier characterizing the loading order; Performing parallel processing on at least one of the data to be processed based on at least one processing thread to determine corresponding consumption data, where the consumption data has the same data identifier as the corresponding data to be processed.

2. The method according to claim 1, characterized in that, The method further includes: Consuming each of the consumption data according to the order of the data identifiers.

3. The method according to claim 1, characterized in that, The method further includes: Adding the consumption data to a queue, and the extraction order of each of the consumption data in the queue is determined according to the order of the data identifiers.

4. The method according to claim 3, characterized in that, The consuming each of the consumption data according to the order of the data identifiers includes: Determining a target identifier, where the target identifier is used to characterize the data identifier of the currently to-be-consumed data; Determining corresponding consumption data from the queue according to the target identifier; Consuming the consumption data corresponding to the target identifier based on a consumption thread; Updating the target identifier.

5. The method according to claim 4, characterized in that, The method further includes: Sending a data ready notification in response to the consumption data being added to the queue.

6. The method according to claim 5, characterized in that, The determining corresponding consumption data from the queue according to the target identifier includes: In response to receiving the data ready notification, determining corresponding consumption data from the queue according to the target identifier.

7. The method according to claim 3, characterized in that, The method further includes: Sending a data completion notification in response to all of the consumption data being added to the queue.

8. The method according to claim 7, characterized in that, The consuming each of the consumption data according to the order of the data identifiers includes: In response to receiving the data completion notification, consuming the consumption data in the queue in sequence according to the order of the data identifiers based on a consumption thread.

9. The method according to claim 1, characterized in that, The loading at least one data to be processed based on a loading thread includes: Obtaining the available state of system resources; In response to there being available system resources, loading at least one data to be processed based on a loading thread.

10. The method according to claim 2, characterized in that, The method further includes: Releasing corresponding system resources in response to the consumption data having been consumed.

11. The method according to claim 2, characterized in that, The method further includes: Controlling the progress of the loading and the consumption through a resource controller.

12. The method according to claim 3, characterized in that, The queue is a minimum heap queue.

13. A data processing device, characterized in that,The apparatus includes: A loading unit, configured to load at least one data to be processed based on a loading thread, where the data to be processed has a data identifier characterizing the loading order; A processing unit, configured to perform parallel processing on at least one of the data to be processed based on at least one processing thread to determine corresponding consumption data, where the consumption data has the same data identifier as the corresponding data to be processed.

14. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store one or more computer program instructions, where the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1-12.

15. A computer-readable storage medium, characterized in that, A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps according to any one of claims 1-12 are implemented.