A data scheduling method, device, apparatus and readable storage medium
By comparing the processing time of the current batch of data with that of the previous batch of data, the data volume scheduling ratio is dynamically adjusted, which solves the high load problem caused by unstable system load and achieves efficient data desensitization processing.
Patent Information
- Application Number
- CN202210956574.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-10
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2042-08-10
AI Technical Summary
Existing data scheduling methods are prone to high load when system load changes, blocking the processing of high-priority packets and resulting in low data desensitization efficiency.
By obtaining the first unit duration of the current batch of data and comparing it with the second unit duration of the pre-stored previous batch of data, the data volume scheduling ratio of the next batch of data is dynamically adjusted, a batch data planning and scheduling record is generated, and the data scheduling process is optimized.
This effectively avoids the high load on the data desensitization system, improves the desensitization efficiency of massive image data, and ensures the efficient operation of the system.
Smart Images

Figure CN115310129B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of data scheduling, and more particularly, to a data scheduling method, device, equipment and readable storage medium. BACKGROUND
[0002] With the development of information technology, people's demand for data is increasing, such as people need to understand the world by obtaining images and texts. These data inevitably involve sensitive information, such as the need to capture environmental images in unmanned driving and capture sensitive information such as pedestrians and license plates, and such as the personal resumes obtained in the recruitment system involve a lot of personal information, which need to be desensitized to remove the private information in the sensitive data.
[0003] Because the amount of data to be desensitized is huge, in the process of data desensitization, the data to be desensitized needs to be dispatched to the desensitization system in batches for desensitization, and the efficiency of the data scheduling process can be improved by optimizing the efficiency of the data scheduling process.
[0004] The commonly used data desensitization scheduling method is to set a division threshold to divide a large number of images into multiple packages, and then hand over each package to the system for desensitization execution. Because the load of the system is constantly changing, it is easy to cause the system to have a high load for a long time to process a single package, which blocks the processing of high-priority packages, resulting in low efficiency.
[0005] By dynamically adjusting the data volume of the data to be dispatched, the phenomenon of high load of the system is avoided, and the mass image data is desensitized efficiently. SUMMARY
[0006] In view of the above problems, the present application is proposed to provide a data scheduling method, device, equipment and readable storage medium, to avoid the phenomenon of high load of the system, and to desensitize the mass image data efficiently.
[0007] In order to achieve the above purpose, the specific scheme is as follows:
[0008] A data scheduling method comprises:
[0009] Upon receiving a data to-be-scheduled signal, a first unit time length for desensitizing a current batch of data is obtained;
[0010] The first unit time length is compared with a second unit time length for desensitizing a previous batch of data of the current batch of data, and a data volume scheduling ratio of a next batch of data of the current batch of data in a previously established batch data planning scheduling record is determined.
[0011] According to the data amount scheduling ratio corresponding to the next batch of data, the next batch of data is adjusted to obtain to-be-scheduled batch data, and the to-be-scheduled batch data is scheduled.
[0012] Optionally, the establishing process of the batch data planning scheduling record comprises:
[0013] According to the directory address to which each to-be-assigned data belongs, the to-be-assigned data is divided to obtain a plurality of batches of data, and the data amount of each batch of data is determined;
[0014] The data scheduling sequence of each batch of data in the plurality of batches of data is determined.
[0015] According to the data amount of each batch of data in the plurality of batches of data and the data scheduling sequence of each batch of data, a batch data planning scheduling record is established.
[0016] Optionally, the first unit time is compared with a second unit time at which the previous batch of data of the current batch of data is desensitized in advance to determine the data amount scheduling ratio of the next batch of data of the current batch of data in the batch data planning scheduling record established in advance, comprising:
[0017] If the first unit time is greater than the second unit time at which the previous batch of data of the current batch of data is desensitized in advance, the data amount scheduling ratio of the next batch of data of the current batch of data in the existing batch data planning scheduling record is determined as a first ratio.
[0018] If the first unit time is not greater than the second unit time at which the previous batch of data of the current batch of data is desensitized in advance, the data amount scheduling ratio of the next batch of data of the current batch of data in the existing batch data planning scheduling record is determined as a second ratio.
[0019] Optionally, according to the data amount scheduling ratio corresponding to the next batch of data, the next batch of data is adjusted to obtain to-be-scheduled batch data, comprising:
[0020] When the data amount scheduling ratio corresponding to the next batch of data is the first ratio, a first part of data in the next batch of data is selected as to-be-scheduled batch data, and the data amount of the first part of data is the product of the data amount of the next batch of data and the first ratio.
[0021] Optionally, after the first part of data in the next batch of data is selected as to-be-scheduled batch data, the method further comprises:
[0022] data in the next batch of data other than the first part of data is determined as load deferred execution data, the load deferred execution data being data after the first part of data in scheduling order.
[0023] Optionally, the next batch of data is adjusted according to a data amount scheduling ratio corresponding to the next batch of data, to obtain to-be-scheduled batch data, including:
[0024] When the data amount scheduling ratio corresponding to the next batch of data is a second ratio, a second part of data in the next batch of data is selected as to-be-scheduled batch data, and a data amount of the second part of data is a result of multiplication of a data amount of the next batch of data and the second ratio.
[0025] Optionally, the to-be-scheduled batch data is scheduled, including:
[0026] The to-be-scheduled batch data is added to an existing task queue, so that a data desensitization processor for data desensitization obtains the to-be-scheduled batch data from the task queue.
[0027] Optionally, a first unit processing duration of current batch data for desensitization processing is obtained, including:
[0028] A data amount of current batch data for desensitization processing and a total processing time are obtained.
[0029] A ratio of the total processing time to the data amount is taken as a first unit duration of the current batch data for desensitization processing.
[0030] Optionally, the ratio of the total processing time to the data amount is taken as the first unit duration of the current batch data for desensitization processing, including:
[0031] The ratio of the total processing time to the data amount is determined.
[0032] A thousandth of the ratio is rounded off, an estimated value before a thousandth of the ratio is reserved, and a duration corresponding to the estimated value is taken as the first unit duration of the current batch data for desensitization processing.
[0033] An apparatus for data scheduling, including:
[0034] A unit duration obtaining unit, configured to, when a data to-be-scheduled signal is received, obtain a first unit duration of current batch data for desensitization processing.
[0035] The scheduling ratio determination unit is used to compare the first unit duration with the second unit duration of the previous batch of data of the current batch data after de-identification processing, which is pre-stored, to determine the data volume scheduling ratio of the next batch of data of the current batch data in the pre-established batch data planning and scheduling record.
[0036] The data to be scheduled determination unit is used to adjust the next batch of data according to the data volume scheduling ratio corresponding to the next batch of data to obtain the batch of data to be scheduled.
[0037] The data scheduling unit is used to schedule the batch of data to be scheduled.
[0038] Optionally, the device may also include:
[0039] The first scheduling record establishment unit is used to divide each data to be allocated into several batches of data according to the directory address of each data to be allocated locally, and to determine the data volume of each batch of data.
[0040] The second scheduling record establishment unit is used to determine the data scheduling order of each batch of data in the plurality of batches of data;
[0041] The third scheduling record establishment unit is used to establish a batch data planning and scheduling record based on the data volume of each batch of data and the data scheduling order of each batch of data.
[0042] Optionally, the scheduling ratio determination unit includes:
[0043] The first ratio determination unit is used to determine the data volume scheduling ratio of the next batch of data in the existing batch data planning and scheduling record as the first ratio if the first unit duration is greater than the second unit duration of the previous batch of data of the current batch data for desensitization processing, which is stored in advance.
[0044] The second ratio determination unit is used to determine the data volume scheduling ratio of the next batch of data in the existing batch data planning and scheduling record as the second ratio if the first unit duration is not greater than the second unit duration of the previous batch of data of the current batch data that has undergone desensitization processing.
[0045] Optionally, the data to be scheduled determination unit includes:
[0046] The first proportion multiplication unit is configured to select a first part of data in the next batch of data as the batch of data to be scheduled when the data volume scheduling proportion corresponding to the next batch of data is a first proportion, and a data volume of the first part of data is a result of multiplication of a data volume of the next batch of data and the first proportion.
[0047] Optionally, the apparatus further comprises:
[0048] The load data determination unit is configured to determine data other than the first part of data in the next batch of data as load deferred execution data, and the load deferred execution data is data after the first part of data in a scheduling sequence.
[0049] Optionally, the data to be scheduled determination unit comprises:
[0050] The second proportion multiplication unit is configured to select a second part of data in the next batch of data as the batch of data to be scheduled when the data volume scheduling proportion corresponding to the next batch of data is a second proportion, and a data volume of the second part of data is a result of multiplication of a data volume of the next batch of data and the second proportion.
[0051] Optionally, the data to be scheduled determination unit comprises:
[0052] The queue data addition unit is configured to add the batch of data to be scheduled to an existing task queue, so that a data desensitization processor for data desensitization obtains the batch of data to be scheduled from the task queue.
[0053] Optionally, the unit time length acquisition unit comprises:
[0054] The processing information acquisition unit is configured to acquire a data volume of a current batch of data for desensitization processing and a total processing time.
[0055] The unit time length calculation unit is configured to take a ratio of the total processing time to the data volume as a first unit time length for the current batch of data for desensitization processing.
[0056] Optionally, the unit time length calculation unit comprises:
[0057] The time length result determination unit is configured to determine a ratio of the total processing time to the data volume.
[0058] The time length result precision unit is configured to round off a thousandth of the ratio, retain an estimated value before the thousandth of the ratio, and take a time length corresponding to the estimated value as the first unit time length for the current batch of data for desensitization processing.
[0059] An apparatus for data scheduling comprises a memory and a processor.
[0060] The memory is configured to store a program.
[0061] The processor is configured to execute the program to implement each step of the method for data scheduling.
[0062] A readable storage medium has a computer program stored thereon, and the computer program, when executed by a processor, implements each step of the method for data scheduling.
[0063] According to the technical solution, when a data scheduling signal is received, a first unit time for desensitizing the current batch data is obtained, the first unit time is compared with a second unit time for desensitizing the previous batch data of the current batch data, a data amount scheduling ratio of the next batch data of the current batch data in a batch data planning scheduling record established in advance is determined, the next batch data is adjusted according to the data amount scheduling ratio corresponding to the next batch data, the batch data to be scheduled is obtained, and the batch data to be scheduled is scheduled. As can be seen, the result of the current batch data processing is compared with the result of the nearest previous batch data processing, the data amount of the next batch data is controlled, the load of the data desensitizing system is prevented from being too high, and finally the batch data to be scheduled is scheduled to the data desensitizing processor, so that the massive image data can be desensitized efficiently. BRIEF DESCRIPTION OF DRAWINGS
[0064] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The accompanying drawings are included to provide a description of the preferred embodiments, and are not intended to limit the scope of the present application. Moreover, like reference numerals designate like parts throughout the several views in the drawings. In the drawings:
[0065] Figure 1 A flowchart of a data scheduling method provided by an embodiment of the present application is shown.
[0066] Figure 2 An architecture diagram of a data scheduling system provided by an embodiment of the present application is shown.
[0067] Figure 3 A device structure diagram of a data scheduling device provided by an embodiment of the present application is shown.
[0068] Figure 4 A structure diagram of a data scheduling device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION
[0069] With reference to the drawings, the technical solutions in the embodiments of the present application will be clearly and completely described below. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.
[0070] The technical solutions in the present application can be implemented based on a terminal with data processing capability, which can be a data scheduler. The terminal can be a computer, a server, a cloud, etc.
[0071] Next, the data scheduling method of the present application will be described in detail with reference to the drawings. Figure 1 The data scheduling method of the present application can include the following steps:
[0072] In step S110, when a data scheduling signal is received, a first unit time length for desensitizing the current batch data is obtained.
[0073] Specifically, the data scheduling signal can be received after the desensitization of the current batch data is completed. The first unit time length for desensitizing the current batch data can represent the average time length for desensitizing each data in the current batch data, or the total time length for desensitizing the current batch data.
[0074] The terminal can obtain the first unit time length for desensitizing the current batch data in the module for supervising data desensitization.
[0075] In step S120, the first unit time length is compared with a second unit time length for desensitizing the previous batch data of the current batch data, and a data amount scheduling ratio of the next batch data of the current batch data in a previously established batch data planning scheduling record is determined.
[0076] Specifically, the batch data planning scheduling record can be established before scheduling each batch data, and each batch data can be scheduled according to the order of each batch data in the batch data planning scheduling record.
[0077] The previous batch data can represent the batch data before the current batch data in the scheduling order, and the next batch data can represent the batch data after the current batch data in the scheduling order.
[0078] It can be understood that the change of the second unit time length to the first unit time length can represent the change of the load of the system, when the first unit time length is not greater than the second unit time length, it can represent that the system load is not heavy, and when the first unit time length is greater than the second unit time length, it can represent that the system load is heavy, and there may be a risk of high load. Based on this, the data amount scheduling ratio of the next batch data of the current batch data in the pre-established batch data planning scheduling record can be determined.
[0079] In step S130, the next batch data is adjusted according to the data amount scheduling ratio corresponding to the next batch data, to obtain the to-be-scheduled batch data, and the to-be-scheduled batch data is scheduled.
[0080] Specifically, the next batch data can be pre-stored locally and temporarily by the terminal, or can be obtained by querying the database according to the information of the next batch data of the current batch data in the pre-established batch data planning scheduling record.
[0081] It can be understood that, due to the different results of the first unit time length and the second unit time length, the next batch data needs to be adjusted according to the data amount scheduling ratio corresponding to the next batch data, and the adjustment is corresponding to the result. For example, when the first unit time length is not greater than the second unit time length, it represents that the load of the system is normal or not high, and then the next batch data can be adjusted by a small amount or not reduced, and when the first unit time length is greater than the second unit time length, it represents that the load of the system has a high risk, and then the next batch data can be adjusted by a large amount.
[0082] For example Figure 2 , Figure 2 A system architecture of a data scheduling system is shown, which can include a data scheduler, a task queue, a data desensitization processor, and a static queue. When the data desensitization processor finishes processing the current batch data, it can send the information of the first unit time length of the desensitization processing of the current batch data and the data to-be-scheduled signal to the static queue. The data scheduler can receive the data to-be-scheduled signal from the static queue and obtain the first unit time length, compare the first unit time length with the pre-stored second unit time length, determine the data amount scheduling ratio of the next batch data and adjust the next batch data, and finally adjust the next batch data as the to-be-scheduled batch data, and schedule it to the data desensitization processor through the task queue.
[0083] The method for data scheduling provided in the embodiment can obtain a first unit time duration for desensitization processing of current batch data when a data-to-be-scheduled signal is received, compare the first unit time duration with a second unit time duration for desensitization processing of previous batch data of the current batch data stored in advance, determine a data amount scheduling ratio of next batch data of the current batch data in a batch data planning scheduling record established in advance, adjust the next batch data according to the data amount scheduling ratio corresponding to the next batch data, obtain batch data to be scheduled, and schedule the batch data to be scheduled. As can be seen, the data amount of next batch data is controlled according to the result of current batch data processing and the result of the most adjacent batch data processing, so as to avoid high load of a data desensitization system, and finally the batch data to be scheduled is scheduled to a data desensitization processor, so that massive image data can be desensitized efficiently.
[0084] In some embodiments of the present application, the establishment process of the batch data planning scheduling record mentioned in the above embodiments is introduced, which can include:
[0085] S1, according to the directory address to which each data to be allocated belongs, the data to be allocated is divided to obtain a plurality of batch data, and the data amount of each batch data is determined.
[0086] Specifically, all data to be allocated can be temporarily stored in a local storage, each data to be allocated can have a directory address belonging to it, and the data to be allocated with the same directory address can be classified and batched to obtain a plurality of batch data, and the number of data to be allocated in each batch data can be determined.
[0087] The directory address can be a directory address selected or fixed by the data to be allocated when uploading.
[0088] S2, the data scheduling order of each batch data in the plurality of batch data is determined.
[0089] Specifically, the data scheduling order can be the time order of the generation time of each batch data, or the upload time order of each batch data when uploading.
[0090] S3, according to the data amount of each batch data in the plurality of batch data and the data scheduling order of each batch data, a batch data planning scheduling record is established.
[0091] Specifically, the batch data planning scheduling record can include the data amount of each batch data and the data scheduling order of each batch data.
[0092] The method for data scheduling provided by the embodiment divides each to-be-allocated data into a plurality of batch data, determines a scheduling sequence of each batch data, and generates a batch data planning scheduling record for a data scheduler to query when scheduling the batch data.
[0093] In some embodiments of the present application, the process of determining the data volume scheduling ratio of the next batch data of the current batch data in the pre-established batch data planning scheduling record is introduced by comparing the first unit time with the second unit time at which the previous batch data of the current batch data is desensitized, and the process can be divided into the following two cases:
[0094] First, if the first unit time is greater than the second unit time at which the previous batch data of the current batch data is desensitized, it is determined that the data volume scheduling ratio of the next batch data of the current batch data in the existing batch data planning scheduling record is a first ratio.
[0095] Specifically, when the first unit time is greater than the second unit time at which the previous batch data of the current batch data is desensitized, it can be indicated that the system load is increasing, and there is a risk of high storage load. The determined first ratio can be a ratio of a substantial reduction in the next batch data.
[0096] The first ratio can be customized, for example, the first ratio is 50%.
[0097] In this case, the process of adjusting the next batch data according to the data volume scheduling ratio corresponding to the next batch data to obtain to-be-scheduled batch data is introduced, and the process can include:
[0098] When the data volume scheduling ratio corresponding to the next batch data is the first ratio, a first part of data in the next batch data is selected as to-be-scheduled batch data, and the data volume of the first part of data is the product of the data volume of the next batch data and the first ratio.
[0099] Specifically, the first part of data in the next batch data can be randomly selected.
[0100] For example, if the next batch data contains 100 data, when the first ratio is 50%, 100*50% = 50 data in the 100 data of the next batch data can be randomly selected as to-be-scheduled batch data.
[0101] It can be understood that when the data quantity scheduling ratio corresponding to the next batch of data is the first ratio, it means that all data of the next batch of data cannot be scheduled to the data desensitization processor at one time due to the prevention of high system load, so that other data in addition to the first part of data will be left in the next batch of data, such as 100-50=50 data left, which can be used as load deferred execution data and can be scheduled after the first part of data is scheduled, and the load deferred execution data can be used as the next batch of data of the first part of data.
[0102] Secondly, if the first unit time is not greater than the second unit time at which the previous batch of data of the current batch of data is desensitized, it is determined that the data quantity scheduling ratio of the next batch of data of the current batch of data in the existing batch data planning scheduling record is the second ratio.
[0103] Specifically, when the first unit time is greater than the second unit time at which the previous batch of data of the current batch of data is desensitized, it can be indicated that the system load is normal or not high, and the determined first ratio can be a ratio of small adjustment or no adjustment of the next batch of data.
[0104] The second ratio can be customized, and the second ratio can be greater than the first ratio, for example, the second ratio is 95% or 100%.
[0105] In this case, the process of adjusting the next batch of data according to the data quantity scheduling ratio corresponding to the next batch of data to obtain the to-be-scheduled batch data according to the above step S130 is introduced, which can include:
[0106] When the data quantity scheduling ratio corresponding to the next batch of data is the second ratio, the second part of data in the next batch of data is selected as the to-be-scheduled batch data, and the data quantity of the second part of data is the product of the data quantity of the next batch of data and the second ratio.
[0107] Specifically, the second part of data in the next batch of data can be randomly selected.
[0108] For example, the next batch of data contains 100 data, and when the second ratio is 95%, 100*95%=95 data can be randomly selected from the 100 data of the next batch of data as the to-be-scheduled batch data.
[0109] The remaining data of the next batch of data can be the next batch of data after the second part of data in the scheduling sequence, or the remaining data can be added to the next batch of data of the current batch of data, and then the remaining data can be scheduled together when the next batch of data is scheduled.
[0110] For example, the next batch of data contains 100 data, and when the second ratio is 100%, all data of the next batch of data can be directly used as the batch of data to be scheduled.
[0111] The method for scheduling data provided by the embodiment can compare the relationship between the first unit time and the second unit time, analyze whether the current load of the system has the risk of too high load, schedule the batch of data according to the normal scheduling ratio if not, and schedule the original batch of data in a halving manner if yes, thereby ensuring the efficient desensitization of the entire system.
[0112] In some embodiments of the present application, the process of scheduling the batch of data to be scheduled mentioned in the above embodiments is introduced, and the process can include:
[0113] The batch of data to be scheduled is added to the existing task queue, so that the data desensitization processor obtains the batch of data to be scheduled from the task queue.
[0114] Specifically, the data scheduler and the data desensitization processor can arrange the scheduling of tasks through the task queue, the data scheduler can add the batch of data to be scheduled to the task queue, and the data desensitization processor can obtain the batch of data to be desensitized from the task queue.
[0115] The method for scheduling data provided by the embodiment can add the batch of data to be scheduled to the task queue, so that the data desensitization processor can access the task queue and obtain the batch of data to be desensitized, thereby avoiding direct access of the data scheduler, and improving the running efficiency of the entire system.
[0116] In some embodiments of the present application, the process of obtaining the first unit processing time of the current batch of data for desensitization processing mentioned in the above embodiments is introduced, and the process can include:
[0117] S1, obtaining the data amount and total processing time of the current batch of data for desensitization processing.
[0118] Specifically, the data amount and total processing time of the current batch of data for desensitization processing can be obtained from the static queue for monitoring the data desensitization state of the data desensitization processor.
[0119] The data desensitization processor can add the data amount information of the batch data and the information of the total processing time to the static queue when ending the processing of each batch data, so as to be obtained by the data scheduler from the static queue.
[0120] S2, taking the ratio of the total processing time to the data amount as a first unit duration for the current batch data to be desensitized.
[0121] It can be understood that the first unit duration represents the average duration of processing a unit of data of the current batch data, and then the ratio of the total processing time to the data amount can be taken as the first unit duration for the current batch data to be desensitized.
[0122] Specifically, considering that the ratio of the total processing time to the data amount can be an infinitely small number, which is not an accurate time value, and needs to be estimated, the step S2 of taking the ratio of the total processing time to the data amount as a first unit duration for the current batch data to be desensitized can include:
[0123] S21, determining the ratio of the total processing time to the data amount.
[0124] It can be understood that the ratio of the total processing time to the data amount can be an infinitely small number, and then the ratio can be kept to a preset length of decimal places, for example, to ten decimal places.
[0125] S22, rounding the thousandth place of the ratio, keeping the estimated value before the thousandth place of the ratio, and taking the duration corresponding to the estimated value as the first unit duration for the current batch data to be desensitized.
[0126] For example, if the ratio of the total processing time to the data amount is 0.450023(s), the first unit duration can be 0.45(s) after rounding to the thousandth place, and if the ratio of the total processing time to the data amount is 0.455023(s), the first unit duration can be 0.46(s) after rounding to the thousandth place.
[0127] In addition, to ensure that the first unit duration is compared with the second unit duration for security, the first unit duration can be estimated to be slightly larger than the actual value.
[0128] Specifically, when there is a non-zero in each place after the hundredth place of the ratio, the hundredth place of the first unit duration can be added by 1 and the places after the hundredth place can be omitted, and the final estimated value is the first unit duration for the current batch data to be desensitized.
[0129] For example, the first unit duration is 0.450023 (s), and since there is a non-zero in each of the percentiles after the percentile (5), the final first unit duration is 0.46 (s).
[0130] The method for data scheduling provided by the embodiment calculates the ratio of the data volume of the current batch of data to be desensitized to the total processing time thereof, rounds the ratio, or processes based on the security analysis of the second unit duration, to obtain a more accurate and higher security coefficient first unit duration.
[0131] The device for implementing data scheduling provided by the embodiment of the present application is described below. The device for implementing data scheduling described below can be referred to in correspondence with the method for implementing data scheduling described above.
[0132] Referring to Figure 3 , Figure 3 The device structure diagram for implementing data scheduling disclosed by the embodiment of the present application is shown.
[0133] As Figure 3 shown, the device can include:
[0134] The unit duration acquisition unit 11 is configured to acquire a first unit duration for desensitization processing of the current batch of data when a data scheduling signal is received.
[0135] The scheduling ratio determination unit 12 is configured to compare the first unit duration with a second unit duration for desensitization processing of a previous batch of data of the current batch of data stored in advance, to determine a data volume scheduling ratio of a next batch of data of the current batch of data in a batch data planning scheduling record established in advance.
[0136] The data to be scheduled determination unit 13 is configured to adjust the next batch of data according to the data volume scheduling ratio corresponding to the next batch of data, to obtain a batch of data to be scheduled.
[0137] The data to be scheduled scheduling unit 14 is configured to schedule the batch of data to be scheduled.
[0138] Optionally, the device further includes:
[0139] The first scheduling record establishment unit is configured to divide each data to be allocated according to a directory address to which the data to be allocated belongs, to obtain a plurality of batches of data, and to determine a data volume of each batch of data.
[0140] The second scheduling record establishment unit is configured to determine a data scheduling sequence of each batch of data in the plurality of batches of data.
[0141] The third scheduling record establishing unit is configured to establish a batch data planning scheduling record according to the data amount of each batch data and the data scheduling sequence of each batch data in the plurality of batch data.
[0142] Optionally, the scheduling proportion determining unit 12 comprises:
[0143] The first proportion determining unit is configured to determine that the data amount scheduling proportion of the next batch data of the current batch data in the existing batch data planning scheduling record is a first proportion if the first unit time length is greater than a second unit time length at which the previous batch data of the current batch data is subjected to desensitization processing and is stored in advance.
[0144] The second proportion determining unit is configured to determine that the data amount scheduling proportion of the next batch data of the current batch data in the existing batch data planning scheduling record is a second proportion if the first unit time length is not greater than the second unit time length at which the previous batch data of the current batch data is subjected to desensitization processing and is stored in advance.
[0145] Optionally, the to-be-scheduled data determining unit 13 comprises:
[0146] The first proportion multiplying unit is configured to select a first part of data in the next batch data as to-be-scheduled batch data when the data amount scheduling proportion corresponding to the next batch data is the first proportion, the data amount of the first part of data being a result of multiplication of the data amount of the next batch data and the first proportion.
[0147] Optionally, the apparatus further comprises:
[0148] The load data determining unit is configured to determine data other than the first part of data in the next batch data as load deferred execution data, the load deferred execution data being data whose scheduling sequence is after the first part of data.
[0149] Optionally, the to-be-scheduled data determining unit 13 comprises:
[0150] The second proportion multiplying unit is configured to select a second part of data in the next batch data as to-be-scheduled batch data when the data amount scheduling proportion corresponding to the next batch data is the second proportion, the data amount of the second part of data being a result of multiplication of the data amount of the next batch data and the second proportion.
[0151] Optionally, the to-be-scheduled data scheduling unit 14 comprises:
[0152] The queue data adding unit is configured to add the to-be-scheduled batch data to an existing task queue, so that a data desensitization processor for data desensitization obtains the to-be-scheduled batch data from the task queue.
[0153] Optionally, the unit time obtaining unit comprises:
[0154] a processing information obtaining unit, configured to obtain a data volume of the current batch data for desensitization processing and a total processing time;
[0155] a unit time calculating unit, configured to take a ratio of the total processing time to the data volume as a first unit time of the current batch data for desensitization processing.
[0156] Optionally, the unit time calculating unit comprises:
[0157] a time result determining unit, configured to determine the ratio of the total processing time to the data volume;
[0158] a time result refining unit, configured to round off the ratio to the nearest ten-thousandth, retain an estimated value before the ten-thousandth of the ratio, and take a time corresponding to the estimated value as the first unit time of the current batch data for desensitization processing.
[0159] The data scheduling apparatus provided by the embodiments of the present application can be applied to a data scheduling device, such as a terminal, a mobile phone, a computer, etc. Optionally, Figure 4 A hardware structure block diagram of the data scheduling device is shown, and reference is made to Figure 4 The hardware structure of the data scheduling device can comprise at least one processor 1, at least one communication interface 2, at least one memory 3 and at least one communication bus 4.
[0160] In the embodiments of the present application, the number of the processor 1, the communication interface 2, the memory 3 and the communication bus 4 is at least one, and the processor 1, the communication interface 2 and the memory 3 complete communication with each other through the communication bus 4.
[0161] The processor 1 can be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application, etc.
[0162] The memory 3 can contain a high-speed RAM memory, and can also include a non-volatile memory, etc., such as at least one disk memory.
[0163] The memory stores a program, and the processor can invoke the program stored in the memory, and the program is used for:
[0164] Upon receiving a data-to-be-scheduled signal, a first unit time of the current batch data for desensitization processing is obtained.
[0165] comparing the first unit time length with a second unit time length at which the previous batch data of the current batch data is desensitized, determining a data amount scheduling ratio of next batch data of the current batch data in a batch data planning scheduling record established in advance;
[0166] adjusting the next batch data according to the data amount scheduling ratio corresponding to the next batch data, obtaining to-be-scheduled batch data, and scheduling the to-be-scheduled batch data.
[0167] Optionally, the refinement function and the extension function of the program can refer to the description above.
[0168] The embodiment of the application further provides a storage medium which can store a program suitable for processor execution, and the program is used for:
[0169] obtaining a first unit time length at which current batch data is desensitized when a data to-be-scheduled signal is received;
[0170] comparing the first unit time length with a second unit time length at which the previous batch data of the current batch data is desensitized, determining a data amount scheduling ratio of next batch data of the current batch data in a batch data planning scheduling record established in advance;
[0171] adjusting the next batch data according to the data amount scheduling ratio corresponding to the next batch data, obtaining to-be-scheduled batch data, and scheduling the to-be-scheduled batch data.
[0172] Optionally, the refinement function and the extension function of the program can refer to the description above.
[0173] Finally, it needs to be noted that, in this document, relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or sequence between the entities or operations. Moreover, the term “comprises”, “includes” or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitation, the element defined by the statement “comprises a” does not exclude the presence of another identical element in the process, method, article or device including the element.
[0174] The various embodiments described in this specification are intended to be combinable unless otherwise indicated herein. The various embodiments described in this specification are described in the progressions noted, with each embodiment emphasizing different aspects over others, and the various embodiments can be combined as desired, with reference to each other as appropriate.
[0175] The above description of disclosed embodiments is intended to be illustrative and not restrictive. Many modifications of these embodiments by one having ordinary skill in the art, using the principles and novel features disclosed herein, will be within the scope of the application. The scope of the application, therefore, is not to be limited to the above described embodiments but is to be accorded the widest scope consistent with the principles and novel features described herein.
Claims
1. A method of scheduling data, characterized by, The application comprises the following steps: When receiving the data scheduling signal, a first unit time for desensitizing the current batch data is obtained; The first unit time is compared with a second unit time for desensitizing the previous batch data of the current batch data, and the data amount scheduling ratio of the next batch data of the current batch data in the previously established batch data planning scheduling record is determined; The next batch data is adjusted according to the data amount scheduling ratio corresponding to the next batch data, and the adjusted batch data is obtained, and the adjusted batch data is scheduled; The establishment process of the batch data planning scheduling record comprises the following steps: According to the directory address of each to-be-allocated data, the to-be-allocated data is divided into several batch data, and the data amount of each batch data is determined; The data scheduling sequence of each batch data in the several batch data is determined; According to the data amount of each batch data and the data scheduling sequence of each batch data, the batch data planning scheduling record is established.
2. The method of claim 1, wherein, The first unit time is compared with a second unit time for desensitizing the previous batch data of the current batch data, and the data amount scheduling ratio of the next batch data of the current batch data in the previously established batch data planning scheduling record is determined, comprising: If the first unit time is greater than the second unit time for desensitizing the previous batch data of the current batch data, the data amount scheduling ratio of the next batch data of the current batch data in the existing batch data planning scheduling record is determined as a first ratio; If the first unit time is not greater than the second unit time for desensitizing the previous batch data of the current batch data, the data amount scheduling ratio of the next batch data of the current batch data in the existing batch data planning scheduling record is determined as a second ratio.
3. The method of claim 2, wherein, The next batch data is adjusted according to the data amount scheduling ratio corresponding to the next batch data, and the adjusted batch data is obtained, comprising: When the data amount scheduling ratio corresponding to the next batch data is the first ratio, a first part of data in the next batch data is selected as the to-be-scheduled batch data, and the data amount of the first part of data is the product of the data amount of the next batch data and the first ratio.
4. The method of claim 3, wherein, After selecting the first part of data in the next batch data as the to-be-scheduled batch data, further comprising: The data in the next batch data except the first part of data is determined as load deferred execution data, and the load deferred execution data is data with a scheduling sequence after the first part of data.
5. The method of claim 2, wherein, The next batch data is adjusted according to the data amount scheduling ratio corresponding to the next batch data, and the adjusted batch data is obtained, comprising: When the data quantity scheduling ratio corresponding to the next batch of data is a second ratio, a second part of data in the next batch of data is selected as the batch of data to be scheduled, and the data quantity of the second part of data is a result of multiplication of the data quantity of the next batch of data and the second ratio.
6. The method of claim 1, wherein, Scheduling the batch of data to be scheduled, comprising: adding the batch of data to be scheduled to an existing task queue, so that a data desensitization processor used for data desensitization acquires the batch of data to be scheduled from the task queue.
7. The method of claim 1, wherein, obtaining a first unit processing time length of the current batch of data for desensitization processing, comprising: obtaining a data quantity of the current batch of data for desensitization processing and a total processing time; taking a ratio of the total processing time to the data quantity as the first unit time length of the current batch of data for desensitization processing.
8. The method of claim 7, wherein, taking a ratio of the total processing time to the data quantity as the first unit time length of the current batch of data for desensitization processing, comprising: determining the ratio of the total processing time to the data quantity; rounding a thousandth of the ratio, retaining an estimated value before the thousandth of the ratio, and taking a time length corresponding to the estimated value as the first unit time length of the current batch of data for desensitization processing.
9. An apparatus for data scheduling, the apparatus comprising: comprising: a unit time length obtaining unit, configured to, when a data to be scheduled signal is received, obtain a first unit time length of a current batch of data for desensitization processing; a scheduling ratio determining unit, configured to compare the first unit time length with a second unit time length of a previous batch of data of the current batch of data for desensitization processing, which is stored in advance, and determine a data quantity scheduling ratio of a next batch of data of the current batch of data in a batch data planning scheduling record established in advance; a to-be-scheduled data determining unit, configured to adjust the next batch of data according to the data quantity scheduling ratio corresponding to the next batch of data, to obtain a batch of data to be scheduled; a to-be-scheduled data scheduling unit, configured to schedule the batch of data to be scheduled; a first scheduling record establishing unit, configured to divide each to-be-allocated data according to a directory address to which the to-be-allocated data belongs, to obtain a plurality of batches of data, and determine a data quantity of each batch of data; a second scheduling record establishing unit, configured to determine a data scheduling sequence of each batch of data in the plurality of batches of data; a third scheduling record establishing unit, configured to establish a batch data planning scheduling record according to the data quantity of each batch of data in the plurality of batches of data and the data scheduling sequence of each batch of data.
10. An apparatus for data scheduling, the apparatus comprising: comprise a memory and a processor; the memory is configured to store a program; the processor is configured to execute the program to implement each step of the method for data scheduling according to any one of claims 1-8.
11. A readable storage medium, having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement each step of the method for data scheduling according to any one of claims 1-8. The computer program is executed by the processor to implement each step of the method for data scheduling according to any one of claims 1-8.
Citation Information
Patent Citations
Batch data transmission method and device
CN102571739A
Batch data desensitization method
CN111324908A