A method and device for processing a data set
A dual-threaded data processing method for face recognition turnstiles synchronizes data download and processing, addressing inefficiencies in manual USB switching by enabling simultaneous operations.
Patent Information
- Application Number
- CN202210226206.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-07
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-03-07
AI Technical Summary
In the prior art, due to the limited storage space of the gate and insufficient USB power supply capacity, it is necessary to manually switch the USB disk to process large data sets, resulting in low testing efficiency.
The dual-threading method is adopted to download and push slice data to the message queue through child threads. The main thread processes the data in a timely manner to realize the synchronization of data download and processing.
It shortens the waiting time of data processing services, realizes efficient processing of data sets, and improves testing efficiency.
Smart Images

Figure CN114595082B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data processing, and particularly to a method and device for processing a data set. Background Art
[0002] A turnstile is a channel blocking device, which is widely used at the entrance channels of various toll and access control occasions. For example, a face recognition turnstile completes identity recognition based on the comparison between the collected face information and the read certificate information. In order to ensure that the face recognition turnstile can accurately and quickly complete the identity recognition process, tests such as recognition performance need to be carried out before it is put into production and use.
[0003] In the related art, a test data set is input into the SDK (Software Development Kit) of the turnstile to achieve automated testing. Considering that the amount of data in the test data set is large, and the turnstile is a small terminal device with a small storage space and limited power supply capacity of its USB port, in the related art, the test data set is scattered and stored in multiple USB flash drives, and each group of test data is processed by switching the USB flash drives.
[0004] Although the method adopted in the related art solves the problem of data storage, however, since this method requires manual switching of the USB flash drive, the data processing is not continuous, resulting in low test efficiency. Summary of the Invention
[0005] In order to solve the above technical problems, the present application provides a method and device for processing a data set, which shortens the waiting time of the data processing service and realizes the efficient processing of the data set.
[0006] The embodiments of the present application disclose the following technical solutions:
[0007] On the one hand, the embodiments of the present application provide a method for processing a data set, and the method includes:
[0008] Start a child thread;
[0009] If the remaining storage capacity of the data storage space is greater than or equal to the size of a sliced data, determine the to-be-downloaded sliced data from N sliced data; the N sliced data are obtained by slicing a target data set;
[0010] Download the to-be-downloaded sliced data to the data storage space through the child thread according to the download address in the download list, and push the data information of the to-be-downloaded sliced data that has been downloaded to the message queue; the download list further includes N pieces of data information, and the N pieces of data information are used to identify the N sliced data;
[0011] When it is detected that there is data information stored in the message queue, the main thread is used to obtain target data information from the message queue;
[0012] Through the main thread, according to the target data information, the target slice data corresponding to the target data information is read from the data storage space, and the target slice data is processed.
[0013] On the other hand, an embodiment of the present application provides a data set processing device, and the device includes a start unit, a determination unit, a download unit, an acquisition unit, and a data processing unit:
[0014] The start unit is used to start a sub-thread;
[0015] The determination unit is used to, if the storage margin of the data storage space is greater than or equal to the size of one slice data, determine the slice data to be downloaded from N slice data; the N slice data is obtained by slicing a target data set;
[0016] The download unit is used to, through the sub-thread, download the slice data to be downloaded to the data storage space according to the download addresses in the download list, and push the data information of the slice data to be downloaded that has been downloaded to the message queue; the download list further includes N pieces of data information, and the N pieces of data information are used to identify the N slice data;
[0017] The acquisition unit is used to, when it is detected that there is data information stored in the message queue, obtain target data information from the message queue through the main thread;
[0018] The data processing unit is used to, through the main thread, read the target slice data corresponding to the target data information from the data storage space according to the target data information, and process the target slice data.
[0019] As can be seen from the above technical solution, when processing a target data set, the download list of the target data set can be obtained first. The download list includes data information and a download address. The data information is used to identify the sliced data obtained after slicing the target data set. Then, a sub-thread is started. When there is a storage margin in the data storage space that can store at least one sliced data, the sub-thread downloads the sliced data to be downloaded to the data storage space according to the download address, and pushes the data information of the sliced data to be downloaded that has been downloaded to the message queue. Further, when it is detected that there is data information stored in the message queue, the main thread obtains the target data information from the message queue, reads the target sliced data according to the target data information, and processes the target sliced data. Specifically, the sliced data to be downloaded is determined from the N sliced data obtained by slicing the target data set according to the storage margin of the data storage space. The sub-thread downloads the sliced data to be downloaded and pushes the data information of the sliced data to be downloaded that has been downloaded to the message queue. When it is detected that there is data information stored in the message queue, the main thread can start processing the target sliced data. It can be seen that by adopting the dual-thread method, the main thread responsible for data processing only needs to wait for the download time of one sliced data to start the data processing service. Thus, the data download and data processing are synchronized, shortening the waiting time of the data processing service and realizing the efficient processing of the data set. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0021] Figure 1 is a flowchart of a method for processing a data set provided by an embodiment of the present application;
[0022] Figure 2 is a schematic diagram of a strategy for a method for processing a data set provided by an embodiment of the present application;
[0023] Figure 3 is a logic diagram of a method for processing a data set provided by an embodiment of the present application;
[0024] Figure 4 is a device structure diagram of a device for processing a data set provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solutions in the embodiments of this application in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.
[0026] As a channel blocking device, the turnstile is widely used in various charging, access control and other occasions. For example, the face recognition turnstile completes identity recognition based on the comparison of the collected face information and the read certificate information. In scenarios such as railway stations, the face recognition turnstile is used to automatically verify the identity of inbound passengers. Whether the identity verification process can be accurately and quickly completed depends to a large extent on the performance of the face recognition turnstile. Therefore, before being put into use, performance testing is required.
[0027] In the related art, the dataset for testing is input into the SDK of the turnstile to achieve automated testing. However, the turnstile is a small terminal device with a small storage space and limited power supply capacity of its own USB port. Thus, for a dataset with a large amount of data, it needs to be dispersed and stored in multiple USB flash drives, and then the USB flash drives are switched to process each group of test data to complete the entire test process and obtain a complete test result. Since this method requires manual switching of USB flash drives, the data processing is not continuous, resulting in low test efficiency.
[0028] To solve the above technical problems, this application provides a dataset processing method and device, which realizes efficient processing of the dataset.
[0029] Specifically, it is described through the following embodiments:
[0030] Figure 1 It is a method flow chart of a dataset processing method provided by an embodiment of this application. The method includes S101 - S105:
[0031] It should be noted that before starting the processing of the target dataset, the method may further include: obtaining the download list of the target dataset.
[0032] Wherein, the download list includes N pieces of data information and download addresses, and the N pieces of data information are used to identify N sliced data obtained by slicing the target dataset.
[0033] To save local storage space, in a possible implementation, the target dataset can be saved on the server, and when it is necessary to use the target dataset for testing the performance of the turnstile, it can be downloaded online.
[0034] When the target data set is stored on the server, in order to quickly complete the download of the data to a certain data storage space and then start data processing by using the downloaded data for synchronization, that is, to synchronize data download and data processing by dynamically controlling data transmission, so as to more efficiently complete the foregoing related tests. Therefore, in one possible implementation, for a target data set with a large amount of data, slicing processing is performed to obtain N slice data, and the N slice data of the target data set are stored on the server.
[0035] In one possible implementation, the slice data obtained by slicing the target data set can be compressed and packaged and then stored on the server.
[0036] When starting to process the target data set, first obtain the download list of the target data set. The download list
[0037] includes N pieces of data information and download addresses, where the N pieces of data information are used to identify the
[0038] N slice data of the target data set.
[0039] In one possible implementation, the data information can be the unique encoding serial number of each slice data. Specifically, in the process of slicing the target data set to obtain N slice data, a unique encoding serial number is determined for each slice data, such as an auto-increment serial number, 1, 2,..., N. In addition, the data information can also be a field for identifying information such as the data content of the slice data. This application does not make any limitation on this.
[0040] S101: Start a sub-thread.
[0041] S102: If the storage margin of the data storage space is greater than or equal to the size of one slice data, determine the slice data to be downloaded from the N slice data according to the storage margin.
[0042] After obtaining the download list of the target data set, start a sub-thread, which is responsible for downloading relevant data. Specifically, when the data storage space has a storage margin capable of storing at least one slice data, determine the slice data to be downloaded from the N slice data according to the current storage margin of the data storage space.
[0043] In one possible implementation, when the data storage space has a storage margin capable of storing at least one slice data, the determined slice data to be downloaded can be one, which is convenient for quickly completing the download task of the slice data to be downloaded. It can be understood that, in combination with the specific network environment where the download task is located, the number of slice data to be downloaded can also be determined to be multiple based on the storage margin of the data storage space. This application does not make any limitation on this.
[0044] S103: According to the download addresses in the download list, the sub-thread downloads the slice data to be downloaded into the data storage space, and pushes the data information of the slice data to be downloaded that has been successfully downloaded into the message queue.
[0045] After determining the slice data to be downloaded, the sub-thread downloads the slice data to be downloaded into the data storage space according to the download addresses in the download list, and pushes the data information of the slice data that has been successfully downloaded in the slice data to be downloaded into the message queue.
[0046] In a possible implementation, when the sub-thread downloads the slice data to be downloaded, it can dynamically detect the data download progress. When it detects that there is slice data to be downloaded that has been successfully downloaded, it pushes the data information of the successfully downloaded slice data into the message queue.
[0047] In a possible implementation, the above download address can be a URL address, and the sub-thread directly downloads the slice data stored on the server through the URL address. When the slice data is stored on the server in the form of a compressed package, correspondingly, the downloaded slice data packet according to the URL address needs to be decompressed, and then it is saved in the data storage space and its data information is pushed into the message queue.
[0048] In a possible implementation, for N slice data of a target data set, a unified download address is determined. Specifically, the download address can be configured according to the storage path of the target data set. Then, after determining the slice data to be downloaded, obtain the data information of the slice data to be downloaded. The sub-thread searches for and downloads the slice data to be downloaded from the N slice data based on the data information according to the unified download address and the data information of the slice data to be downloaded.
[0049] In order to more directly complete the download of the slice data, in a possible implementation, for N slice data of a target data set, correspondingly N different download addresses are determined. Specifically, a unique download address can be configured for each slice data according to the save path of each slice data under the root directory of the target data set. Then, after determining the slice data to be downloaded, obtain the download address of the slice data to be downloaded. The sub-thread directly downloads the slice data according to the download address of each slice data to be downloaded, saving steps such as searching and matching based on data information during the data download process.
[0050] The data storage space is used to store the slice data downloaded by the sub-thread, and its form can be selected according to the actual situation. In a possible implementation, the data storage space is a target directory or the message queue.
[0051] Considering the security of data storage, a target directory can be configured locally to store the sliced data downloaded by the sub-thread. Correspondingly, the main thread obtains the target sliced data from the target directory. It can be understood that a message queue can also be used as a data storage space to store the sliced data downloaded by the sub-thread, thereby saving the overhead of local storage space. Correspondingly, the main thread obtains the target sliced data from the message queue. This application does not make any restrictions on this.
[0052] Since the operation of a thread consumes computer resources, in a possible implementation, after the sub-thread downloads the sliced data to be downloaded to the data storage space according to the download addresses in the download list and pushes the data information of the sliced data to be downloaded that has been downloaded to the message queue, it further includes:
[0053] Deleting the data information of the sliced data to be downloaded that has been downloaded from the download list;
[0054] When it is determined that the download list does not contain data information, changing the status of the download flag from the download status to the download completed status and ending the sub-thread.
[0055] Thus, after S103, the data information of the sliced data to be downloaded that has been downloaded is deleted from the download list. It can be seen that by dynamically updating the download list, the download list is used to identify the undownloaded part of the N sliced data for the target data set. Thus, when it is determined that the download list does not contain any data information, it can be determined that all the N sliced data for the target data set have been downloaded. At this time, the download task of the sliced data responsible by the sub-thread has been completed, and the sub-thread can be ended to avoid unnecessary waste of computer resources.
[0056] In a possible implementation, before starting the sub-thread, it further includes:
[0057] Setting the status of the download flag to the download status.
[0058] Before starting the sub-thread, the download flag is initialized and its status is set to the download status to indicate that there is a download task for the target data set.
[0059] When it is determined that the download list does not contain any data information, that is, the download task for the target data has been all completed, changing the status of the download flag from the download status to the download completed status.
[0060] S104: When it is detected that the message queue stores data information, the main thread obtains the target data information from the message queue.
[0061] For the processing of the target data set, the main thread is responsible for the data processing service. Since the main thread and the sub-threads run independently, in a possible implementation, the message queue can be dynamically detected to obtain the data download progress of the sub-threads in real time.
[0062] Specifically, when it is detected that there is data information stored in the message queue, it indicates that the sub-thread has completed the download of some slice data. Correspondingly, some slice data has been stored in the data storage space. Then, the main thread starts the data processing service and obtains the target data information from the message queue.
[0063] Considering the characteristics of the message queue, for the main thread to obtain the target data information from the message queue, in a possible implementation, the data information located at the head of the queue can be directly obtained as the target data information starting from the head of the message queue. In another possible implementation, starting from the head of the message queue, a certain number of data information sorted from the head of the queue can be sequentially obtained as the target data information. For example, the first 3 data information sorted from the head of the queue can be obtained as the target data information. This can be set according to the actual situation (such as the data processing capacity of the current device, the data storage space, etc.), and the present application does not make any limitation on this.
[0064] S105: Through the main thread, according to the target data information, read the target slice data corresponding to the target data information from the data storage space, and process the target slice data.
[0065] The main thread reads the target slice data corresponding to the target data information from the data storage space according to the obtained target data information, and processes the target slice data.
[0066] The message queue is a shared variable between the main thread and the sub-threads. The main thread obtains the target data information from the message queue in response to the detection result that there is data information stored in the message queue, so as to read and process the target slice data. Correspondingly, the main thread can also respond to the detection result that there is no data information stored in the message queue. Therefore, in a possible implementation, after the main thread processes the target slice data, it further includes:
[0067] When it is determined that there is no data information in the message queue, obtain the status of the download flag;
[0068] If the status of the download flag is the download completed status, end the main thread.
[0069] When the detection result shows that the message queue does not contain data information, obtain the status of the download flag, where the status of the download flag is used to identify whether there is a download task for the target data set, specifically including: the download status is used to indicate that there is a download task for the target data set, and the download completed status is used to indicate that there is no download task for the target data set.
[0070] Therefore, when the status of the obtained download flag is the download completed status, it indicates that all download tasks related to the target data set have been completed, and the message queue also does not contain any data information, which means that all data processing tasks related to the target data set have been completed. At this time, the main thread is ended to save computer resources.
[0071] In order to avoid unexpected situations such as interruption of the processing process of the target slice data due to reasons such as data reading interruption, in a possible implementation, S105 includes the following steps:
[0072] S1051: Through the main thread, according to the target data information, read the target slice data corresponding to the target data information from the data storage space;
[0073] S1052: Save the read target slice data to the data processing middleware and process the target slice data;
[0074] S1053: After the main thread finishes processing the target slice data, delete the target slice data in the data processing middleware.
[0075] Specifically, the main thread saves the read target slice data to the data processing middleware and processes the target slice data. The data processing middleware is a secure and efficient data exchange platform. Using it as the storage container for the target slice data can make data reading and data processing run relatively independently and effectively connect these two tasks.
[0076] To save the storage resources of the data processing middleware, after the main thread finishes processing the target slice data, delete the target slice data saved in the data processing middleware to release the storage space.
[0077] During the entire processing of the target data set, the main thread and the sub-thread run synchronously and are related, and the process involves data flow control for sliced data. Specifically, different data processing processes are presented based on the magnitude relationship between the data processing speed of the main thread and the data download speed of the sub-thread. For example, when there is a difference between the data processing speed of the main thread and the data download speed of the sub-thread, problems such as excessive waste of storage space due to storing a large amount of downloaded data and waiting time for data processing operations due to insufficient downloaded data for data processing operations may occur.
[0078] Therefore, based on the magnitude relationship between the data processing speed of the main thread and the data download speed of the sub-thread, the embodiments of the present application provide a schematic diagram of a strategy for a data set processing method as shown in Figure 2 Specifically:
[0079] As shown in Figure 2 (a) of, if the data processing speed of the main thread is the same as the data download speed of the sub-thread, it indicates that for the data flow processing of sliced data, one sliced data is processed while one sliced data is downloaded. It can be seen that the utilization of computer resources in this case is the best, so there is no need to adjust the download strategy of the sub-thread.
[0080] Furthermore, in order to avoid related problems caused by the difference between the data processing speed of the main thread and the data download speed of the sub-thread, such as data download failure due to insufficient storage space, in a possible implementation, the method further includes:
[0081] S11: Obtain the data processing speed of the main thread and the data download speed of the sub-thread;
[0082] S12: If the data processing speed and the data download speed are different, adjust the download strategy of the sub-thread.
[0083] Specifically, by obtaining the data processing speed of the main thread and the data download speed of the sub-thread, when the data processing speed and the data download speed are different, the download strategy of the sub-thread can be adjusted correspondingly according to their magnitude relationship.
[0084] Among them, the data processing speed can be used to identify the time required for the main thread to process one sliced data, and the data download speed can be used to identify the time required for the sub-thread to download one sliced data. It can be understood that the corresponding data processing speed and data download speed can also be determined based on the size of the sliced data. For example, the data processing speed is 3.2M / s and the data download speed is 3M / s. The present application does not make any limitations on this.
[0085] As shown in Figure 2As shown in (b) of [the figure], when the data processing speed is less than the data download speed, a queue length upper limit needs to be set for the message queue. Otherwise, if the data is downloaded too fast, the data storage space will be quickly filled up, leading to problems such as data storage overflow and data download failure due to insufficient remaining space in the data storage. To avoid the above problems, in a possible implementation, S12 includes:
[0086] If the data processing speed is less than the data download speed, obtain the remaining queue length of the message queue;
[0087] If the remaining queue length is less than the size of one data message, adjust the sub-thread to enter the waiting download state.
[0088] Specifically, when the data processing speed is less than the data download speed, obtain the remaining queue length of the message queue, and determine that when the remaining queue length of the message queue is not sufficient to store one data message, the sub-thread needs to be adjusted to enter the waiting download state, wait for the message queue to be idle, and resume the sub-thread to download data.
[0089] Since one data message corresponds to one slice of data, in a possible way, the queue length upper limit of the message queue can be set according to the size of the data storage space. For example, if the size of the data storage space can store M slice data, then correspondingly, the queue length upper limit of the message queue can be set to M, which means that the message queue can store M data messages.
[0090] Such as Figure 2 As shown in (c) of [the figure], when the data processing speed is greater than the data download speed, the data processing speed being too fast may cause the main thread to enter the process of waiting for the sub-thread to download data, and due to the existence of the data processing service waiting duration, the data processing efficiency for the target data set will be reduced. Therefore, in a possible implementation, S12 includes:
[0091] If the data processing speed is greater than the data download speed, increase the data download speed of the sub-thread.
[0092] Specifically, when the data processing speed is greater than the data download speed, increase the data download speed of the sub-thread to make it transform into Figure 2 the state shown in (a) of [the figure] or Figure 2 the state shown in (b) of [the figure] to avoid the data processing service responsible by the main thread from entering the state of waiting for queue data. Regarding how to increase the data download speed of the sub-thread, an appropriate method can be selected according to the actual situation, such as:
[0093] In a possible implementation, the data download speed of the sub-thread for downloading sliced data can be increased by obtaining sliced data with a smaller data volume through smaller data slicing of the target data set and by increasing the compression ratio of the sliced data, etc.
[0094] In a possible implementation, the data download volume can be increased by starting multiple sub-threads for multi-threaded concurrent downloading and then merging, so as to provide sufficient data to be processed for the main thread.
[0095] It can be understood that in addition to the above ways to increase the data download speed of the sub-thread, other ways can also be used, such as using a network environment with a higher download speed, etc. The present application does not make any limitation in this regard.
[0096] Figure 3 This is a logic diagram of a data set processing method provided by an embodiment of the present application. Specifically, at the beginning of data set processing, the download flag can be initialized first, the status of the download flag is set to the download status, and the download list of the target data set is passed in, and the sub-thread is started.
[0097] During the running process of the sub-thread, first judge the message queue. If the message queue is full, the sub-thread can be adjusted to enter the sleep state and data download is not performed temporarily until the message queue can store at least one data message. Then the sub-thread starts to download sliced data, decompress the sliced data and push the data message of the sliced data to the message queue. Further, judge whether the download list is empty. When it is judged that the download list is empty, that is, when the download list does not contain data messages, the status of the download flag is set to the download completed status, indicating that the download task for the target data set has been completed. At this time, the sub-thread ends.
[0098] During the running process of the main thread, first judge the message queue. When it is detected and judged that the message queue is not empty, that is, when there are data messages stored in the message queue, the main thread obtains the target data message from the message queue, reads the target sliced data corresponding to the target data message, and processes the target sliced data. After the main thread finishes processing the target sliced data, the target sliced data is deleted; when it is detected and judged that the message queue is empty, that is, when there are no data messages in the message queue, further obtain the status of the download flag. When it is judged that the status of the download flag is not the download completed status, it indicates that there are still unfinished download tasks for the target data set. At this time, the main thread can be adjusted to enter the sleep state, wait for queue data, and the download strategy of the sub-thread can be adjusted; when it is judged that the status of the download flag is the download completed status, it indicates that both the download task and the processing task for the target data set have been completed. At this time, the main thread ends.
[0099] It can be seen that when processing the target data set, the download list of the target data set can be obtained first. The download list includes data information and download addresses. The data information is used to identify the slice data obtained after slicing the target data set. Then, a sub-thread is started. When there is a storage margin in the data storage space that can store at least one slice data, the sub-thread downloads the slice data to be downloaded to the data storage space according to the download address, and pushes the data information of the slice data to be downloaded that has been downloaded to the message queue. Further, when it is detected that there is data information stored in the message queue, the main thread obtains the target data information from the message queue, reads the target slice data according to the target data information, and processes the target slice data. Specifically, the slice data to be downloaded is determined from the N slice data obtained by slicing the target data set according to the storage margin of the data storage space. The sub-thread downloads the slice data to be downloaded and pushes the data information of the slice data to be downloaded that has been downloaded to the message queue. When it is detected that there is data information stored in the message queue, the main thread can start processing the target slice data. It can be seen that by adopting the dual-thread method, the main thread responsible for data processing only needs to wait for the download time of one slice data to start the data processing service. Thus, data download and data processing are synchronized, the waiting time of the data processing service is shortened, and efficient processing of the data set is achieved.
[0100] Figure 4 FIG. 4 is a structural diagram of a data set processing apparatus provided by an embodiment of the present application. The apparatus includes a start unit 401, a determination unit 402, a download unit 403, an acquisition unit 404, and a data processing unit 405:
[0101] The start unit 401 is configured to start a sub-thread;
[0102] The determination unit 402 is configured to, if the storage margin of the data storage space is greater than or equal to the size of one slice data, determine the slice data to be downloaded from the N slice data; the N slice data are obtained by slicing the target data set;
[0103] The download unit 403 is configured to download the slice data to be downloaded to the data storage space according to the download address in the download list through the sub-thread, and push the data information of the slice data to be downloaded that has been downloaded to the message queue; the download list further includes N pieces of data information, and the N pieces of data information are used to identify the N slice data;
[0104] The acquisition unit 404 is configured to, when it is detected that there is data information stored in the message queue, obtain the target data information from the message queue through the main thread;
[0105] The data processing unit 405 is configured to read, by the main thread, target slice data corresponding to the target data information from the data storage space, and process the target slice data.
[0106] In a possible implementation, the data processing unit is further configured to:
[0107] Read, by the main thread, target slice data corresponding to the target data information from the data storage space;
[0108] Save the read target slice data to the data processing middleware, and process the target slice data;
[0109] After the main thread finishes processing the target slice data, delete the target slice data in the data processing middleware.
[0110] In a possible implementation, after the download unit downloads the to-be-downloaded slice data to the data storage space according to the download addresses in the download list by the sub-thread and pushes the data information of the to-be-downloaded slice data that has been downloaded to the message queue, the download unit is further configured to:
[0111] Delete the data information of the to-be-downloaded slice data that has been downloaded from the download list;
[0112] When it is determined that the download list does not contain data information, change the status of the download flag from the download status to the download completed status, and end the sub-thread.
[0113] In a possible implementation, before starting the sub-thread, the download unit is further configured to:
[0114] Set the status of the download flag to the download status.
[0115] In a possible implementation, after the main thread finishes processing the target slice data, the data processing unit is further configured to:
[0116] When it is determined that the message queue does not contain data information, obtain the status of the download flag;
[0117] If the status of the download flag is the download completed status, end the main thread.
[0118] In a possible implementation, the download unit is further configured to:
[0119] Obtain the data processing speed of the main thread and the data download speed of the sub-thread;
[0120] If the data processing speed is different from the data download speed, adjust the download policy of the sub-thread.
[0121] In a possible implementation, the download unit is further configured to:
[0122] If the data processing speed is less than the data download speed, obtain the remaining queue length of the message queue;
[0123] If the remaining queue length is less than the size of one piece of data information, adjust the sub-thread to enter the waiting download state.
[0124] In a possible implementation, the download unit is further configured to:
[0125] If the data processing speed is greater than the data download speed, increase the data download speed of the sub-thread.
[0126] In a possible implementation, the data storage space is the target directory or the message queue.
[0127] Thus, when processing the target data set, the download list of the target data set can be obtained first. The download list includes data information and download addresses, and the data information is used to identify the sliced data obtained by slicing the target data set. Then, start the sub-thread. When the data storage space has a storage margin capable of storing at least one sliced data, download the sliced data to be downloaded to the data storage space according to the download address through the sub-thread, and push the data information of the downloaded sliced data to be downloaded to the message queue. Further, when it is detected that there is data information stored in the message queue, obtain the target data information from the message queue through the main thread, read the target sliced data according to the target data information, and process the target sliced data. Specifically, determine the sliced data to be downloaded from the N sliced data obtained by slicing the target data set according to the storage margin of the data storage space, download the sliced data to be downloaded through the sub-thread, and push the data information of the downloaded sliced data to be downloaded to the message queue. When it is detected that there is data information stored in the message queue, the main thread can start processing the target sliced data. It can be seen that by adopting the dual-thread method, the main thread responsible for data processing only needs to wait for the download time of one sliced data to start the data processing service, thereby realizing the synchronous progress of data download and data processing, shortening the waiting time of the data processing service, and realizing the efficient processing of the data set.
[0128] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial descriptions of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0129] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of another identical element in the process, method, article or device comprising the element.
[0130] The above has introduced in detail a method and device for processing a data set provided by an embodiment of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method of the present application. At the same time, for those of ordinary skill in the art, there will be changes in the specific implementation manner and application scope according to the method of the present application.
[0131] In summary, the content of this specification should not be construed as a limitation to the present application. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered within the protection scope of the present application. Moreover, based on the implementation manners provided in the above aspects of the present application, further combinations can be made to provide more implementation manners.
Claims
1. A method for processing a data set, characterized in that, The method includes: Start a child thread; If the remaining storage capacity of the data storage space is greater than or equal to the size of a slice of data, determine the slice of data to be downloaded from N slices of data; the N slices of data are obtained by slicing a target data set; Download the slice of data to be downloaded to the data storage space according to the download address in the download list through the child thread, and push the data information of the slice of data to be downloaded that has been downloaded to the message queue; the download list further includes N pieces of data information for identifying the N slices of data; When it is detected that there is data information stored in the message queue, obtain the target data information from the message queue through the main thread; Read the target slice of data corresponding to the target data information from the data storage space through the main thread according to the target data information, and process the target slice of data.
2. The method according to claim 1, wherein The step of reading the target slice of data corresponding to the target data information from the data storage space through the main thread according to the target data information, and processing the target slice of data includes: Read the target slice of data corresponding to the target data information from the data storage space through the main thread according to the target data information; Save the read target slice of data to the data processing middleware, and process the target slice of data; After the main thread finishes processing the target slice of data, delete the target slice of data in the data processing middleware.
3. The method according to claim 1, characterized in that, After downloading the slice of data to be downloaded to the data storage space according to the download address in the download list through the child thread, and pushing the data information of the slice of data to be downloaded that has been downloaded to the message queue, it further includes: Delete the data information of the slice of data to be downloaded that has been downloaded from the download list; When it is determined that the download list does not contain data information, switch the status of the download flag from the download status to the download completed status, and end the child thread.
4. The method according to claim 3, wherein Before starting the child thread, it further includes: Set the status of the download flag to the download status.
5. The method according to claim 3, characterized in that, After the main thread finishes processing the target slice of data, it further includes: When it is determined that the message queue does not contain data information, obtain the status of the download flag; If the status of the download flag is the download completed status, end the main thread.
6. The method according to claim 1, wherein It further includes: Obtain the data processing speed of the main thread and the data download speed of the child thread; If the data processing speed and the data download speed are different, adjust the download strategy of the child thread.
7. The method according to claim 6, wherein The step of adjusting the child thread if the data processing speed and the data download speed are different includes: If the data processing speed is less than the data download speed, obtain the remaining queue length of the message queue; If the remaining queue length is less than the size of one piece of data information, adjust the child thread to enter the waiting download state.
8. The method according to claim 6, characterized in that, The step of adjusting the child thread if the data processing speed and the data download speed are different includes: If the data processing speed is greater than the data download speed, increase the data download speed of the child thread.
9. The method according to any one of claims 1-8, characterized in that, The data storage space is the target directory or the message queue.
10. A data set processing device, characterized in that, The device includes a start unit, a determination unit, a download unit, an acquisition unit, and a data processing unit: The start unit is used to start a child thread; The determination unit is used to determine the slice data to be downloaded from the N slice data according to the storage margin if the storage margin of the data storage space is greater than or equal to the size of one slice data; the N slice data are obtained by slicing a target data set; The download unit is used to download the slice data to be downloaded to the data storage space through the child thread according to the download address in the download list, and push the data information of the slice data to be downloaded that has been downloaded to the message queue; the download list further includes N pieces of data information, and the N pieces of data information are used to identify the N slice data; The acquisition unit is used to obtain target data information from the message queue through the main thread when it is detected that there is data information stored in the message queue; The data processing unit is used to read the target slice data corresponding to the target data information from the data storage space through the main thread according to the target data information, and process the target slice data.
Citation Information
Patent Citations
Data transmission method and device
CN105187533A
Multi-thread linked list processing method and device and computer readable storage medium
CN110598054A