Data set equalization processing method and device, electronic equipment and storage medium
By dynamically adjusting the data amount of the data set in multimodal training to achieve equalization, the problem of uneven data parallel computing in multimodal training is solved, and the computing efficiency and resource utilization are improved.
Patent Information
- Application Number
- CN202411856245.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-16
AI Technical Summary
In multimodal training, due to the essential differences in processing methods of images and videos, the data sizes after ViT encoding are inconsistent, which in turn leads to load imbalance problems in parallel computing, forming the so-called "bubble" phenomenon, reducing the utilization rate of computing resources and training efficiency.
By obtaining the data sizes of multiple pending multimodal data, calculating the average value of the original data set as the reference value, comparing the data with the reference value, determining the maximum heap and the minimum heap, and gradually adjusting the data distribution in the original data set by dynamically exchanging the amount of data in the maximum heap and the minimum heap, until all data volumes reach equalization.
The data volume balance in the data center is realized, computing efficiency and resource utilization are improved, efficient execution of parallel computing tasks is ensured, and computing bottlenecks caused by uneven data volume is avoided.
Smart Images

Figure CN120011804A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to technical fields such as deep learning, multimodal training, and distributed large model training optimization, and especially to a data set balanced processing method, device, electronic device, and storage medium. Background Art
[0002] In the field of multimodal training, a common practice is to introduce image and video coding models, such as Vision Transformer (ViT), before large deep learning models. However, due to the essential differences in the way images and videos are processed, this difference causes the data encoded by ViT to show significant differences in size. This inconsistency in data size further causes the problem of load imbalance in the parallel computing process, that is, the amount of data in some computing paths is much larger than that in other paths, thus forming the so-called "bubble" phenomenon during training. These bubbles not only reduce the utilization of computing resources, but also affect the efficiency and performance of training, resulting in the overall training effect failing to meet expectations. Therefore, how to effectively solve the problem of data parallel computing imbalance in multimodal training has become a key challenge to improve training performance and efficiency. Summary of the invention
[0003] The present disclosure provides a data set balancing processing method, device, electronic device and storage medium.
[0004] According to one aspect of the present disclosure, a method for balancing a data set is provided, the method comprising:
[0005] Obtaining the data size of multiple multimodal data to be processed to obtain the original data set;
[0006] Calculate the average value of the original data set to obtain a benchmark value;
[0007] Calculate the data in the original data set and the benchmark value, and determine the maximum heap and the minimum heap according to the calculation results;
[0008] The amount of data in the original data set is dynamically adjusted by exchanging the amount of data in the maximum heap and the amount of data in the minimum heap, so that the amount of data in the original data set is balanced.
[0009] According to another aspect of the present disclosure, a data set balancing processing device is provided, comprising:
[0010] An acquisition module is used to acquire the data size of a plurality of multimodal data to be processed to obtain an original data set;
[0011] A first calculation module, used to calculate the average value of the original data set to obtain a reference value;
[0012] A second calculation module, used for calculating the data in the original data set and the reference value, and determining a maximum heap and a minimum heap according to the calculation results;
[0013] The balancing module is used to dynamically adjust the amount of data in the original data set by exchanging the amount of data in the maximum heap and the minimum heap, so that the amount of data in the original data set is balanced.
[0014] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0015] at least one processor; and
[0016] a memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any method in any of the above technical solutions.
[0018] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any one of the methods described in the above technical solutions.
[0019] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program, wherein the computer program implements any one of the methods described in the above technical solutions when executed by a processor.
[0020] The present disclosure provides a method, device, equipment and storage medium for balancing data sets. The present disclosure collects and composes an original data set containing multiple sizes of multimodal data to be processed, and then calculates the average value of these data as a benchmark value. Next, each data in the original data set is compared with the benchmark value, and the data is allocated to the maximum heap and the minimum heap according to the difference. Finally, by dynamically exchanging the amount of data in the maximum heap and the minimum heap, the data distribution in the original data set is gradually adjusted until all data amounts are balanced. This method achieves data volume balance in the data set, which not only improves computing efficiency, but also optimizes resource utilization and ensures efficient execution of parallel computing tasks.
[0021] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0023] Figure 1 is a schematic diagram of steps of a data set balancing processing method in an embodiment of the present disclosure;
[0024] Figure 2 yes Figure 1 Schematic diagram of the process of step S101;
[0025] Figure 3 yes Figure 1 Flow chart of step S102;
[0026] Figure 4 yes Figure 1 Schematic diagram of the process of step S103;
[0027] Figure 5 yes Figure 1 Flow chart of step S104;
[0028] Figure 6 yes Figure 5 A flow chart of step S502;
[0029] Figure 7 yes Figure 5 Another flowchart of step S502;
[0030] Figure 8 A principle block diagram of a data set equalization processing device in an embodiment of the present disclosure;
[0031] Fig. 9 It is a block diagram of an electronic device used to implement the data set equalization processing method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0032] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0033] The present disclosure provides a method for balancing a data set. Figure 1 As shown, Figure 1 : is a schematic diagram of the steps of a data set balancing processing method in an embodiment of the present disclosure. The method can be applied to a server side. The method includes:
[0034] Step S101, obtaining the data size of a plurality of multimodal data to be processed to obtain an original data set.
[0035] Specifically, it is first necessary to collect the data volume information of the multimodal data to be processed distributed on different computing nodes or devices. These data volumes refer to the data scale that will participate in the computing task on each node or device. For multimodal data, multimodal data may include text data, image data, and video data, etc. These data volumes are taken as elements, and then arranged in a certain order to form a data structure called "original data set". The original data set is the basis for recording and comparing the amount of multimodal data to be processed by each computing node. It provides an intuitive reference for subsequent data processing, which is conducive to the subsequent comparison and analysis of the data volume of different nodes, and then the balanced distribution of data volume can be achieved. That is, this step is the prerequisite for achieving load balancing, and lays the foundation for subsequent optimization of computing efficiency and reasonable allocation of resources.
[0036] In this way, by obtaining the data size of multiple multimodal data to be processed and forming a data set, it is possible to centrally manage and monitor the data load on each computing node, providing an accurate data basis for subsequent load balancing and resource optimization. That is, this step can clearly understand the data processing capabilities of each node in the entire system, thus laying a solid foundation for achieving balanced data distribution and optimizing computing resource utilization, and helping to improve the overall computing efficiency and performance of the system.
[0037] Step S102, calculating the average value of the original data set to obtain a reference value.
[0038] Specifically, calculating the average value of the original data set and obtaining the benchmark value means that in the data processing scheme, it is first necessary to perform mathematical calculations on the amount of multimodal data to be processed on each computing node (i.e., the original data set) collected to obtain the arithmetic mean of all data volumes. This average value is called the benchmark value, which represents the expected average state of the data volume on all computing nodes. The specific implementation process can be to add up all the values in the original data set and divide it by the number of values in the data set, and the result is the benchmark value. This benchmark value is then used as a reference standard for measuring and adjusting the amount of data at each node, so as to determine which nodes have too much data and which nodes have insufficient data by comparing with the benchmark value, and then guide the subsequent data redistribution to achieve load balancing. This process is a key step in achieving balanced data distribution in the entire system, ensuring that computing tasks can be fairly and reasonably distributed among the nodes.
[0039] In this way, by calculating the average value of the original data set, a reasonable data volume target, namely the benchmark value, can be determined, which represents the ideal average data processing volume of all computing nodes in the system. This benchmark value serves as a reference point for subsequent balancing operations, helping to identify and adjust nodes with excessive or insufficient data volume, thereby effectively distributing the load, optimizing resource utilization, and improving overall computing efficiency. It can be seen that this method ensures that in multi-node parallel processing scenarios, each node can evenly bear computing tasks, avoiding the situation where some nodes are overloaded while other nodes are idle, and ultimately achieving more stable and efficient data processing.
[0040] Step S103, calculating the data in the original data set and the reference value, and determining the maximum heap and the minimum heap according to the calculation results.
[0041] Specifically, the Max Heap and the Min Heap are two special binary heap data structures, which follow different sorting properties: in the Max Heap, the value of each parent node is always greater than or equal to the value of its child node, which means that the top of the heap (root node) is the maximum value in the entire heap, which is convenient for fast access and deletion of the maximum element; on the contrary, in the Min Heap, the value of each parent node is always less than or equal to the value of its child node, so that the top of the heap is the minimum value in the heap, which is convenient for fast access and deletion of the minimum element. These two heap structures are widely used in various algorithms, such as sorting algorithms, priority queue implementations, etc. They ensure fast access to specific values and maintenance of heap properties through effective data organization.
[0042] In the embodiment of the present disclosure, the data in the original data set is calculated with the benchmark value, and when the maximum heap and the minimum heap are determined according to the calculation results, the difference between the amount of multimodal data to be processed of each computing node (i.e., the value in the original data set) and the average value (benchmark value) calculated previously is first determined. This process is achieved by comparing each data point in the original data set with the benchmark value: if a data point is greater than the benchmark value, the difference (i.e., the excess part) is put into the maximum heap, indicating that the amount of data of the node exceeds the average expected value; on the contrary, if the data point is less than the benchmark value, the difference (i.e., the insufficient part) is put into the minimum heap, indicating that the amount of data of the node is lower than the average expected value. The maximum heap and the minimum heap store positive and negative differences respectively, providing a basis for subsequent load balancing operations, so that data exceeding the average value can be redistributed to those nodes below the average value to achieve overall data volume balance. This method effectively utilizes the data structure characteristics of the maximum heap and the minimum heap to dynamically adjust the data distribution and optimize the resource allocation in parallel computing.
[0043] In this way, by calculating the data in the original data set with the benchmark value and determining the maximum heap and minimum heap based on the calculation results, the imbalance of data volume on each computing node can be effectively identified and quantified. This calculation method can not only accurately locate which nodes have data volumes exceeding the average (forming the maximum heap) and which nodes have data volumes below the average (forming the minimum heap), but also provide clear guidance for subsequent data redistribution, ensuring that the implementation of data load balancing strategies is more efficient and accurate. Ultimately, this method can significantly improve the overall performance and stability of the parallel processing system, reduce computing bottlenecks caused by uneven data volume by optimizing resource allocation, and achieve a more even distribution of computing load.
[0044] Step S104 , dynamically adjusting the amount of data in the original data set by exchanging the amount of data in the maximum heap and the amount of data in the minimum heap, so that the amount of data in the original data set is balanced.
[0045] Specifically, since the maximum heap is used to store nodes with excess data that exceeds the benchmark value (i.e., the average data volume), and the minimum heap is used to store nodes with insufficient data that is lower than the benchmark value. By dynamically taking out the data volume from the maximum heap (i.e., from the node with larger data volume) and allocating it to the data volume in the minimum heap (i.e., the node with smaller data volume), the data volume difference in the system can be gradually reduced. Specifically, if the top element of the maximum heap (i.e., the node with the largest data volume) can fully meet the needs of the top element of the minimum heap (i.e., the node with the least data volume), then the data volume is directly transferred; if the element of the maximum heap cannot fully meet the element of the minimum heap, then after transferring the necessary data volume, the remaining data volume is put back into the maximum heap; conversely, if there is still a remaining demand for the element of the minimum heap after the transfer, then this part of the demand will also be put back into the minimum heap. This process is repeated until both the maximum heap and the minimum heap are empty, at which time the data volume in the original data set reaches a balanced state, ensuring that each node in the system carries a data volume close to the average value, thereby optimizing the overall data processing capability and resource utilization.
[0046] In this way, by dynamically redistributing the amount of data, the data load on each computing node is balanced, thereby improving the efficiency and performance of parallel processing and reducing the situation where some nodes are overloaded and other nodes are idle due to uneven data volume. Ultimately, this dynamic adjustment mechanism helps to improve the stability and response speed of the overall system and ensure that computing tasks can be executed evenly and efficiently on all nodes.
[0047] The present disclosure provides a method, device, equipment and storage medium for balancing data sets. The present disclosure collects and composes an original data set containing multiple sizes of multimodal data to be processed, and then calculates the average value of these data as a benchmark value. Next, each data in the original data set is compared with the benchmark value, and the data is allocated to the maximum heap and the minimum heap according to the difference. Finally, by dynamically exchanging the amount of data in the maximum heap and the minimum heap, the data distribution in the original data set is gradually adjusted until all data amounts are balanced. This method achieves a balanced amount of data in the data set, which not only improves computing efficiency, but also optimizes resource utilization and ensures efficient execution of parallel computing tasks.
[0048] In some optional embodiments, see Figure 2 , Figure 2 yes Figure 1 Flow diagram of step S101. Step S101, obtaining the data size of a plurality of multimodal data to be processed to obtain an original data set, including:
[0049] Step S201, obtaining multimodal data output by the image and video coding model, and calculating the data size of the output multimodal data to obtain a calculation result.
[0050] Specifically, multimodal data refers to data that contains multiple types of information, such as images and videos. In the field of deep learning, these data need to be encoded for model processing. Image and video coding models (such as VisionTransformer) refer to models that can convert raw image and video data into a format suitable for machine learning model input. The implementation process of this solution includes two main steps: first, obtain the encoded multimodal data from the image and video coding model, which is the output result after model processing; second, quantify the size of these output data, that is, measure the size or number of each data set, and obtain a specific calculation result. This result can be used to evaluate the distribution of data and provide a basis for subsequent data equalization processing.
[0051] Step S202, calculating the standard deviation of the above calculation results, if the standard deviation exceeds a preset threshold, the data size of the corresponding multimodal data is used as the original data set.
[0052] Specifically, "standard deviation" is a key statistical indicator to measure the degree of discreteness of data distribution, which indicates the degree of deviation between the values in the data set and the average value. In this scheme, the standard deviation of the set of data sizes of multimodal data output by the image and video coding model (i.e., the calculation result) is first calculated. If this standard deviation exceeds a predetermined threshold, it indicates that the data is unevenly distributed among different modalities, and the threshold is a benchmark for judging whether the data is balanced. In this case, the data size set of these multimodal data whose standard deviation exceeds the threshold is regarded as the original data set, which means that these data will be used for further analysis and processing, so as to take corresponding measures to adjust and optimize the data distribution and ensure the balance of the data during the multimodal training process. That is, the embodiment of the present disclosure identifies and marks the data sets that need special processing by evaluating the relationship between the standard deviation of the data size and the preset threshold to achieve the goal of data equalization.
[0053] In this way, by obtaining the multimodal data output by the image and video coding model and calculating the data size of these data, detailed calculation results about the data distribution can be obtained. Furthermore, by calculating the standard deviation of these results, the discreteness of the data can be quantified. If the standard deviation exceeds the preset threshold, it indicates that there is a significant imbalance in the data between different modalities. At this time, these data sizes are used as the original data set, so that these unbalanced data can be accurately identified and optimized. That is, this method can actively monitor and respond to data imbalance problems, thereby improving the efficiency and accuracy of data processing, optimizing the subsequent multimodal training process, and ensuring the balance and effectiveness of model training.
[0054] In some optional embodiments, see Figure 3 , Figure 3 yes Figure 1 Flow diagram of step S102. Step S102, calculating the average value of the original data set to obtain a reference value, includes:
[0055] Step S301, sum up all the data in the original data set to obtain the total data volume.
[0056] Specifically, the "original data set" refers to a set of data collected and determined for analysis in the previous steps. "Accumulation" refers to the process of adding all the values in these data sets one by one. In the process of implementing the solution, it is necessary to first determine the original data set, which may include data collected from multiple nodes or multiple time periods. Then, by performing an accumulation operation, that is, summing each data point in the data set, the total data volume of the entire data set is obtained. This total data volume provides the total value of all individual data points in the data set, which can be used for subsequent data analysis, such as calculating averages, benchmarks, or performing other statistical analysis, thereby providing a quantitative basis for further data management and decision-making. The accumulation operation combines all data points in the original data set into a total data volume, laying the foundation for subsequent data processing and analysis.
[0057] Step S302, calculating the ratio between the total data volume and the number of data in the original data set to obtain a reference value.
[0058] Specifically, "total data volume" refers to the total amount obtained by adding up all the data in the original data set in the previous step, while "number of data" refers to the number of data points contained in the data set. The ratio between these two values is calculated, that is, the total data volume is divided by the number of data to obtain the average value of each data point, which is called the "benchmark value". The benchmark value provides a measure to evaluate the average performance of each data point in the original data set relative to the whole, and is an important reference indicator in subsequent data analysis and processing.
[0059] In this way, by accumulating all the data in the original data set, the total data volume is obtained; by calculating the ratio between the total data volume and the number of data in the original data set, an ideal data distribution standard, namely the benchmark value, can be determined. This benchmark value represents the amount of data that each node should process if the data volume is evenly distributed on all nodes under ideal conditions. This method helps to quickly identify the imbalance of the current data distribution, provides precise guidance for subsequent data redistribution, and ensures the fairness and efficiency of data processing. In this way, resource utilization can be optimized and the overall performance of the system can be improved, especially in parallel computing and distributed systems, which helps to avoid the problem of overloading some nodes while other node resources are idle, and achieves more balanced and efficient data processing.
[0060] In some optional embodiments, see Figure 4 , Figure 4 yes Figure 1 Flow diagram of step S103. Step S103, calculating the data in the original data set and the reference value, and determining the maximum heap and the minimum heap according to the calculation results, including:
[0061] Step S401, calculating the difference between each data in the original data set and the reference value to obtain a difference data set.
[0062] Specifically, first determine the benchmark value, which is the sum of all data points in the original data set divided by the number of data points. Then, for each data point in the original data set, calculate the difference between it and the benchmark value, which is the result of subtracting the benchmark value from each data point. These differences are collected to form a new data set, called the "difference data set." This difference data set reflects the degree of deviation of each data point from the overall average state and can be used for subsequent data analysis and processing.
[0063] Step S402: Divide the differences in the difference data set into a maximum heap or a minimum heap according to a preset value.
[0064] Specifically, the "preset value" is a predetermined threshold used to distinguish whether a data point is above or below the average. After obtaining the difference data set, the difference between each data point and the benchmark value is calculated, and then these differences are compared with the preset value. If a difference is greater than the preset value, it indicates that the value of the data point is higher than the benchmark value, and this difference is placed in the "max heap"; on the contrary, if the difference is less than the preset value, it indicates that the value of the data point is lower than the benchmark value, and this difference is placed in the "min heap". The maximum heap and the minimum heap are two special data structures used to store larger and smaller values, respectively, for efficient data management and processing. In this way, data points can be effectively classified according to their relative size to the average value, which facilitates subsequent data processing and analysis.
[0065] In this way, by calculating the difference between each data in the original data set and the benchmark value, a difference data set is obtained, and the difference in the difference data set is divided into a maximum heap or a minimum heap according to the preset value. This division method enables the nodes that exceed the average value (maximum heap) to provide additional data for the nodes that are below the average value (minimum heap), thereby effectively redistributing data. This process not only optimizes the distribution of data, but also improves the overall computing efficiency and resource utilization, ensures the load balance of the system, avoids the problem of some nodes being overloaded while other nodes are idle, and ultimately achieves more efficient and stable data processing.
[0066] In some optional embodiments, the difference values in the difference data set are divided into a maximum heap or a minimum heap according to a preset value, including:
[0067] Extract the difference values greater than the preset value from the difference data set to form a maximum heap;
[0068] Differences less than a preset value are extracted from the difference data set to form a minimum heap.
[0069] Specifically, the "difference data set" refers to a data set that contains the difference between each data in the original data set and the benchmark value (that is, the average value of all data). After obtaining the difference data set. Then, according to a pre-set threshold, two types of differences are separated from the difference data set: those differences greater than the preset value are selected to form a "maximum heap", which is used to identify nodes whose data volume exceeds the average value; and the differences less than the preset value are formed into a "minimum heap", which is used to identify nodes whose data volume does not reach the average value. The construction of the maximum heap and the minimum heap is to make effective data volume adjustments in subsequent steps, so that nodes with larger data volumes can transfer appropriate amounts of data to nodes with smaller data volumes to achieve balanced data load in the entire system. This process ensures the strategic nature of data redistribution and improves the efficiency and fairness of data processing.
[0070] In this way, by sorting the differences in the difference data set, extracting the differences greater than the preset value from the difference data set to form a maximum heap, and extracting the differences less than the preset value from the difference data set to form a minimum heap, the data load in the system can be effectively classified and prioritized. This method allows nodes with an average load (maximum heap) and nodes with a load below the average (minimum heap) to be clearly distinguished, providing clear guidance for subsequent data redistribution. In this way, it can be ensured that nodes with a large amount of data can appropriately reduce their load, while nodes with a small amount of data can obtain additional data, thereby achieving load balancing for the entire system, improving computing efficiency and resource utilization, and ensuring system stability and response speed.
[0071] In some optional embodiments, see Figure 5 , Figure 5 yes Figure 1 Flow diagram of step S104. Step S104, dynamically adjusting the amount of data in the original data set by exchanging the amount of data in the maximum heap and the minimum heap so that the amount of data in the original data set is balanced, includes:
[0072] Step S501, obtaining a maximum value in a maximum heap and a minimum value in a minimum heap, and obtaining a first value and a second value respectively.
[0073] Specifically, a "max heap" is a binary heap in which the value of each node is greater than or equal to the value of its child node, so the top element of the heap is the maximum value in the entire heap; in contrast, a "min heap" is another binary heap in which the value of each node is less than or equal to the value of its child node, so that the top element of the heap is the minimum value in the heap. In the process of implementing the embodiment of the present disclosure, the two heaps are first identified and constructed, wherein the maximum heap contains all data point difference values above the average value, and the minimum heap contains all data point difference values below the average value. Then, the top element of the heap is extracted from the maximum heap to obtain the "first value", that is, the maximum difference; the top element of the heap is extracted from the minimum heap to obtain the "second value", that is, the minimum difference. These two values represent the maximum deviation and minimum deviation from the benchmark value in the data set, respectively, and provide key reference values for further data processing and analysis.
[0074] Step S502 , dynamically allocating part of the data in the original data set corresponding to the first value to the data in the original data set corresponding to the second value according to the data volume of the first value and the data volume of the second value, until all data volumes in the original data set are balanced.
[0075] Specifically, after determining the first value and the second value, a portion of the data is dynamically transferred from the original data set corresponding to the first value with a larger data volume to the data set corresponding to the second value with a smaller data volume. This transfer is performed based on the difference in the data volume of the two, with the purpose of reducing the difference between them. This process continues until the data volume in all original data sets reaches a balanced state, that is, the data is evenly distributed in each data set without obvious deviation. In this way, the problem of load imbalance in the data set can be effectively solved, the data processing efficiency can be improved, and the use of resources can be optimized.
[0076] In this way, the maximum value in the maximum heap and the minimum value in the minimum heap are obtained to obtain the first value and the second value respectively, and then according to the data volume of the first value and the data volume of the second value, part of the data in the original data set corresponding to the first value is dynamically allocated to the original data set corresponding to the second value until all the data volumes in the original data set are balanced. This method effectively balances the data distribution in the entire system by identifying and adjusting the nodes with the largest and smallest data volumes, and dynamically allocating the excess data volume from the maximum heap node to the minimum heap node. This can significantly improve the processing efficiency and resource utilization of the system, ensure that each node can evenly undertake computing tasks, avoid the problem of some nodes being overloaded while other nodes have idle resources, and ultimately achieve more efficient and stable data processing.
[0077] In some optional embodiments, see Figure 6 , Figure 6 yes Figure 5A flow chart of step S502. Step S502, dynamically allocating part of the data in the original data set corresponding to the first value to the data in the original data set corresponding to the second value according to the data volume of the first value and the data volume of the second value, until all the data volumes in the original data set are balanced, includes:
[0078] Step S601, if the data volume of the first value is greater than the data volume of the second value, extract the same data volume as the second value from the original data set data corresponding to the first value, and assign it to the original data set data corresponding to the second value to obtain an updated data set data.
[0079] Specifically, if the amount of data of the first value exceeds the amount of data of the second value, this adjustment step will be performed: an amount of data equal to the amount of data of the second value is extracted from the original data set corresponding to the first value with a larger amount of data, and it is redistributed to the original data set corresponding to the second value with a smaller amount of data. The purpose of this process is to reduce the difference between the two sets of data, so that the amount of data of the second value increases, while the amount of data of the first value decreases accordingly. Through this data redistribution, an updated data set is finally obtained, in which the amount of data in each part is more balanced. In short, this feature achieves a balanced distribution of data within the data set and optimizes the overall data distribution by extracting data from a large set to supplement a small set.
[0080] Step S602, calculating the difference between the data volume of the first value and the data volume of the second value to obtain the remaining data volume of the first value.
[0081] Specifically, the difference between the two is determined by subtracting the amount of data in the "second value" from the amount of data in the "first value". This difference represents the amount of data remaining in the "first value" relative to the "second value" after data redistribution. In short, this feature accurately calculates the amount of data that exceeds or remains in the first value after considering the data balancing requirements through a simple subtraction operation, providing a quantitative basis for subsequent data adjustments.
[0082] Step S603: Add the remaining data amount of the first value to the maximum heap, re-obtain the maximum value in the maximum heap and the minimum value in the minimum heap, and obtain a new first value and a new second value.
[0083] Specifically, during the adjustment process, if there is a difference in the amount of data, the remaining amount of data in the "first value" (i.e., the amount of data that exceeds the "second value") is added to the maximum heap, which is a data structure in which the amount of data stored is the largest of all the data to be processed. Subsequently, the maximum value (the updated maximum amount of data) is extracted from the maximum heap, and the minimum value (the updated minimum amount of data) is extracted from the minimum heap, and these two values become the "new first value" and "new second value", respectively. This process ensures that sets with larger data volumes and sets with smaller data volumes can be re-evaluated and adjusted to achieve a balanced distribution of data. In short, this feature determines a new data volume metric by adding the data volume difference back to the maximum heap and updating the values of the maximum heap and the minimum heap, thereby promoting a more balanced data distribution between data sets.
[0084] Step S604: reallocate part of the data in the data set corresponding to the new first value to the data set corresponding to the new second value until all data in the data set is balanced.
[0085] Specifically, a certain amount of data is extracted from the data set corresponding to the "new first value" and assigned to the data set corresponding to the "new second value" in order to reduce the difference between the two. This process is repeated until the data volume of all data sets reaches a balanced state, that is, there is no obvious set with too large or too small data volume, thereby ensuring load balancing of the entire data set. In short, this feature gradually achieves a balanced distribution of data volume within the entire data set by dynamically transferring data from sets with larger data volumes to sets with smaller data volumes, thereby optimizing the overall distribution and processing efficiency of data.
[0086] The disclosed embodiment achieves efficient allocation and optimization of resources by dynamically adjusting nodes with unbalanced data volumes. Specifically, when the data volume (first value) of one node is greater than that of another node (second value), the excess data volume is transferred from the node of the first value to the node of the second value to reduce the difference between them. Subsequently, by calculating the new difference in the data volume of the two and updating the maximum heap and the minimum heap, it is possible to continuously identify and adjust the nodes with the largest and smallest data volumes until the data volumes of all nodes in the entire system are balanced. This method not only improves the efficiency of data processing, but also ensures the stability and response speed of the system, avoids the problem that some nodes have performance degradation due to data overload, while other nodes have idle resources due to insufficient data, and ultimately achieves load balancing and maximum utilization of resources.
[0087] In some optional embodiments, see Figure 7 , Figure 7 yes Figure 5Another flow chart of step S502. Step S502, dynamically allocating part of the data in the original data set corresponding to the first value to the data in the original data set corresponding to the second value according to the data amount of the first value and the data amount of the second value, until all the data amounts in the original data set are balanced, includes:
[0088] Step S701, if the data volume of the first value is less than the data volume of the second value, extract the same data volume as the first value from the original data set data corresponding to the first value, and allocate it to the original data set data corresponding to the second value to obtain the updated data set data.
[0089] Specifically, when the data volume of the "first value" is less than that of the "second value", an amount of data equal to the "first value" is extracted from the original data set corresponding to the "first value" and redistributed to the data set corresponding to the "second value". This process is to balance the difference between the two data sets so that the reduced part of the data set of the "first value" can be supplemented to the data set of the "second value", thereby achieving a balanced data volume between the data sets. In short, this feature extracts a corresponding amount of data from the data set with a smaller data volume and distributes it to the data set with a larger data volume to achieve a balanced data state in the updated data set.
[0090] Step S702, calculating the difference between the data volume of the second value and the data volume of the first value to obtain the data volume of the second value that differs.
[0091] Specifically, the difference between the amount of data of the second value and the amount of data of the first value is calculated, that is, the amount of data of the "first value" is subtracted from the amount of data of the "second value", so as to obtain the difference between them. This result is called the "amount of data of the second value that differs". This difference reflects the degree of difference in the amount of two data sets and is an important indicator for evaluating data balance. That is, by calculating the difference between the amount of data of the two data sets, a specific value representing the difference between them is obtained, which provides a basis for subsequent data adjustment and balance.
[0092] Step S703, adding the data amount of the second value that differs to the minimum heap, re-obtaining the maximum value in the maximum heap and the minimum value in the minimum heap, to obtain a new first value and a new second value.
[0093] Specifically, after determining the data volume difference between the "first value" and the "second value", this difference value, that is, the "data volume of the second value that differs", will be added to the "minimum heap", which is a data structure that stores smaller values. Subsequently, the largest value (representing the data set with the largest data volume) is extracted from the "maximum heap", and the smallest value (representing the data set with the smallest data volume) is extracted from the "minimum heap", and these two values become the "new first value" and the "new second value", respectively. This process allows the system to continuously identify and adjust the data sets with the largest and smallest data volumes to achieve dynamic balancing of data. That is, by adding the data volume difference to the minimum heap and re-obtaining values from the maximum heap and the minimum heap, the first and second values used for subsequent balancing processing are updated.
[0094] Step S704: reallocate part of the data in the data set corresponding to the new first value to the data set corresponding to the new second value until all data in the data set is balanced.
[0095] Specifically, the part of the data that exceeds the balancing requirement in the data set corresponding to the "new first value" is transferred to the data set corresponding to the "new second value", which can reduce the load of the data set with a larger data volume and increase the load of the data set with a smaller data volume. This process will continue until the data volume of all data sets is adjusted to a balanced state, that is, there is no obvious data set with too large or too small data volume. That is, by continuously extracting data from the set with a larger data volume and allocating it to the set with a smaller data volume until all data volumes in the entire data set are balanced, the overall distribution and processing efficiency of the data are optimized.
[0096] In this way, by accurately controlling the transfer of data from nodes with larger data volumes (second value) to nodes with smaller data volumes (first value), dynamic balancing of the system load is achieved. Specifically, it first transfers all the data volumes of nodes with smaller data volumes to nodes with larger data volumes, then calculates the difference between the two and re-adds this difference to the minimum heap. By continuously updating the maximum heap and the minimum heap and redistributing data, the data volumes of all nodes in the system are balanced. This method not only improves the efficiency of data processing, but also ensures the stability and response speed of the system, avoiding the problem of performance degradation of some nodes due to data overload and idle resources of other nodes due to insufficient data, and ultimately achieving load balancing and maximum utilization of resources.
[0097] In order to facilitate the overall understanding of the technical solution of the present application, an example is given below: assuming that there are 4 computing devices (such as GPUs), and the batch_size (data size) they are currently processing is 10, 15, 20, and 25, respectively. These data are combined into a data set to obtain an original data set, namely [10, 15, 20, 25]. First, the average value of the original data set is calculated to obtain a reference value (b_avg). By calculating the average value of these 4 batch_sizes: b_avg = (10 + 15 + 20 + 25) / 4 = 17.5, the reference value is 17.5.
[0098] Calculate the difference between each number in the original data set and the reference value to obtain the difference data set A1, that is, [10-17.5, 15-17.5, 20-17.5, 25-17.5], and get [-7.5, -2.5, 2.5, 7.5].
[0099] Construct the maximum heap H1 and minimum heap H2 according to the difference data set: put the positive and negative values in the difference data set A1 into the maximum heap H1 and the minimum heap H2 respectively, and obtain the maximum heap H1 (storing positive difference values): [7.5, 2.5]; the minimum heap H2 (storing negative difference values): [-7.5, -2.5].
[0100] Next, swap the top elements of the heap and adjust: swap the top of the maximum heap H1 (maximum value 7.5) and the top of the minimum heap H2 (minimum value -7.5): device 25 (original batch_size is 25) sends part of the data to device 10 (original batch_size is 10), so that the batch_size difference between the two is reduced by 7.5. After the swap, update the difference data set A1, and reconstruct the maximum heap H1 and the minimum heap H2: get the new difference data set A1: [0, -2.5, 2.5, 0]. Updated maximum heap H1: [2.5], updated minimum heap H2: [-2.5]. Swap the top of the maximum heap H1 (2.5) and the top of the minimum heap H2 (-2.5) again: device 20 (original batch_size is 20) sends part of the data to device 15 (original batch_size is 15), so that the batch_size difference between the two is reduced by 2.5. After the exchange, the difference data set A1 is updated, and the maximum heap H1 and minimum heap H2 are rebuilt: the new difference data set A1: [0,0,0,0]. Both the maximum heap H1 and the minimum heap H2 are empty. After the above steps, the batch_size of all computing devices is adjusted to be close to the mean value of 17.5, specifically [17.5,17.5,17.5,17.5], achieving balanced distribution of computing load.
[0101] The above example simply shows how to dynamically adjust the batch_size on each computing device to make the load in parallel computing more balanced, thereby improving the overall computing efficiency.
[0102] The following describes an apparatus embodiment of the present application, which can be used to execute the data set balancing processing method in the above-mentioned embodiment of the present application. For details not disclosed in the apparatus embodiment of the present application, please refer to the above-mentioned embodiment of the data set balancing processing method of the present application.
[0103] The present disclosure also provides a data set balancing processing device 800, such as Figure 8 As shown, including:
[0104] An acquisition module 801 is used to acquire the data size of a plurality of multimodal data to be processed to obtain an original data set;
[0105] A first calculation module 802 is used to calculate the average value of the original data set to obtain a reference value;
[0106] The second calculation module 803 is used to calculate the data in the original data set and the reference value, and determine the maximum heap and the minimum heap according to the calculation results;
[0107] The balancing module 804 is used to dynamically adjust the amount of data in the original data set by exchanging the amount of data in the maximum heap and the amount of data in the minimum heap, so that the amount of data in the original data set is balanced.
[0108] In some optional embodiments, the first calculation module 802 calculates the average value of the original data set to obtain the reference value, including:
[0109] Accumulate all the data in the original data set to obtain the total data volume;
[0110] The ratio between the total data volume and the number of data in the original data set is calculated to obtain the benchmark value.
[0111] In some optional embodiments, the second calculation module 803 calculates the data in the original data set and the reference value, and determines the maximum heap and the minimum heap according to the calculation results, including:
[0112] Calculate the difference between each data in the original data set and the benchmark value to obtain a difference data set;
[0113] Divide the differences in the difference data set into a maximum heap or a minimum heap according to the preset value.
[0114] In some optional embodiments, the second calculation module 803 divides the difference values in the difference data set into a maximum heap or a minimum heap according to a preset value, including:
[0115] Extract the difference values greater than the preset value from the difference data set to form a maximum heap;
[0116] Differences less than a preset value are extracted from the difference data set to form a minimum heap.
[0117] In some optional embodiments, the balancing module 804 dynamically adjusts the amount of data in the original data set by exchanging the amount of data in the maximum heap and the minimum heap so that the amount of data in the original data set is balanced, including:
[0118] Obtain the maximum value in the maximum heap and the minimum value in the minimum heap to obtain a first value and a second value respectively;
[0119] According to the data volume of the first value and the data volume of the second value, part of the data in the original data set corresponding to the first value is dynamically allocated to the original data set corresponding to the second value until all data volumes in the original data set are balanced.
[0120] In some optional embodiments, the balancing module 804 dynamically allocates part of the data in the original data set corresponding to the first value to the original data set corresponding to the second value according to the data amount of the first value and the data amount of the second value, until all data amounts in the original data set are balanced, including:
[0121] If the data volume of the first value is greater than the data volume of the second value, extracting the same amount of data as the data volume of the second value from the original data set data corresponding to the first value, and assigning it to the original data set data corresponding to the second value, so as to obtain updated data set data;
[0122] Calculate the difference between the data volume of the first value and the data volume of the second value to obtain the remaining data volume of the first value;
[0123] Add the remaining data of the first value to the maximum heap, re-obtain the maximum value in the maximum heap and the minimum value in the minimum heap, and obtain a new first value and a new second value;
[0124] Part of the data in the data set corresponding to the new first value is reallocated to the data set corresponding to the new second value until all data amounts in the data set are balanced.
[0125] In some optional embodiments, the balancing module 804 dynamically allocates part of the data in the original data set corresponding to the first value to the original data set corresponding to the second value according to the data amount of the first value and the data amount of the second value, until all data amounts in the original data set are balanced, including:
[0126] If the data volume of the first value is less than the data volume of the second value, extracting the same amount of data as the data volume of the first value from the original data set data corresponding to the first value, and assigning it to the original data set data corresponding to the second value, so as to obtain updated data set data;
[0127] Calculate the difference between the data volume of the second value and the data volume of the first value to obtain the data volume of the second value that differs;
[0128] The data amount of the second value that differs is added to the minimum heap, and the maximum value in the maximum heap and the minimum value in the minimum heap are re-obtained to obtain a new first value and a new second value;
[0129] Part of the data in the data set corresponding to the new first value is reallocated to the data set corresponding to the new second value until all data amounts in the data set are balanced.
[0130] In some optional embodiments, the acquisition module 801 acquires the data size of a plurality of multimodal data to be processed to obtain an original data set, including:
[0131] Obtaining multimodal data output by the image and video coding model, and calculating the data size of the output multimodal data to obtain a calculation result;
[0132] The standard deviation of the above calculation results is calculated. If the standard deviation exceeds a preset threshold, the data size of the corresponding multimodal data is used as the original data set.
[0133] In the technical solution disclosed herein, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0134] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0135] Fig. 9 A schematic block diagram of an example electronic device 900 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0136] like Fig. 9As shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0137] A number of components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0138] The computing unit 901 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above, such as a data set equalization processing method. For example, in some embodiments, the data set equalization processing method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the applet distribution described above may be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to perform the data set equalization processing method in any other appropriate manner (e.g., by means of firmware).
[0139] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0140] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data set equalization processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.
[0141] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0142] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0143] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0144] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0145] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0146] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A data set balancing processing method, the method comprising: Obtaining the data size of multiple multimodal data to be processed to obtain the original data set; Calculate the average value of the original data set to obtain a benchmark value; Calculate the data in the original data set and the benchmark value, and determine the maximum heap and the minimum heap according to the calculation results; The amount of data in the original data set is dynamically adjusted by exchanging the amount of data in the maximum heap and the amount of data in the minimum heap, so that the amount of data in the original data set is balanced.
2. The method according to claim 1, wherein: Calculating the average value of the original data set to obtain a reference value includes: Accumulating all data in the original data set to obtain a total data volume; The ratio between the total data volume and the number of data in the original data set is calculated to obtain the reference value.
3. The method according to claim 1, wherein: The calculating the data in the original data set and the reference value, and determining the maximum heap and the minimum heap according to the calculation results, includes: Calculate the difference between each data in the original data set and the reference value to obtain a difference data set; The difference values in the difference data set are divided into the maximum heap or the minimum heap according to a set preset value.
4. The method according to claim 3, wherein: The dividing the difference values in the difference data set into the maximum heap or the minimum heap according to the preset value includes: Extracting differences greater than the preset value from the difference data set to form a maximum heap; Differences smaller than the preset value are extracted from the difference data set to form a minimum heap.
5. The method according to any one of claims 1 to 4, wherein: The dynamically adjusting the amount of data in the original data set by exchanging the amount of data in the maximum heap and the amount of data in the minimum heap so that the amount of data in the original data set is balanced includes: Obtaining a maximum value in the maximum heap and a minimum value in the minimum heap to obtain a first value and a second value respectively; According to the data volume of the first value and the data volume of the second value, part of the data in the original data set corresponding to the first value is dynamically allocated to the original data set corresponding to the second value until all data volumes in the original data set are balanced.
6. The method according to claim 5, wherein: The dynamically allocating part of the data in the original data set corresponding to the first value to the data in the original data set corresponding to the second value according to the data amount of the first value and the data amount of the second value until all data amounts in the original data set are balanced includes: If the data volume of the first value is greater than the data volume of the second value, extracting the same amount of data as the data volume of the second value from the original data set data corresponding to the first value, and allocating the data to the original data set data corresponding to the second value, so as to obtain updated data set data; Calculate the difference between the data volume of the first value and the data volume of the second value to obtain the remaining data volume of the first value; Add the remaining data amount of the first value to the maximum heap, re-obtain the maximum value in the maximum heap and the minimum value in the minimum heap, and obtain a new first value and a new second value; Part of the data in the data set corresponding to the new first value is reallocated to the data set corresponding to the new second value until all data amounts in the data set are balanced.
7. The method according to claim 5, wherein: The dynamically allocating part of the data in the original data set corresponding to the first value to the data in the original data set corresponding to the second value according to the data amount of the first value and the data amount of the second value until all data amounts in the original data set are balanced includes: If the data volume of the first value is less than the data volume of the second value, extracting the same amount of data as the data volume of the first value from the original data set data corresponding to the first value, and allocating the data to the original data set data corresponding to the second value, so as to obtain updated data set data; Calculate the difference between the data volume of the second value and the data volume of the first value to obtain the data volume of the second value that differs; Add the data amount of the second value that differs to the minimum heap, re-obtain the maximum value in the maximum heap and the minimum value in the minimum heap, and obtain a new first value and a new second value; Part of the data in the data set corresponding to the new first value is reallocated to the data set corresponding to the new second value until all data amounts in the data set are balanced.
8. The method according to claim 1, wherein: The step of obtaining the data size of the plurality of multimodal data to be processed to obtain the original data set includes: Obtaining multimodal data output by the image and video coding model, and calculating the data size of the output multimodal data to obtain a calculation result; The standard deviation of the calculation result is calculated, and if the standard deviation exceeds a preset threshold, the data size of the corresponding multimodal data is used as the original data set.
9. A data set equalization processing device, comprising: An acquisition module is used to acquire the data size of a plurality of multimodal data to be processed to obtain an original data set; A first calculation module, used to calculate the average value of the original data set to obtain a reference value; A second calculation module, used for calculating the data in the original data set and the reference value, and determining a maximum heap and a minimum heap according to the calculation results; The balancing module is used to dynamically adjust the amount of data in the original data set by exchanging the amount of data in the maximum heap and the minimum heap, so that the amount of data in the original data set is balanced.
10. The device according to claim 9, wherein: The first calculation module calculates the average value of the original data set to obtain a reference value, including: Accumulating all data in the original data set to obtain a total data volume; The ratio between the total data volume and the number of data in the original data set is calculated to obtain the reference value.
11. The device according to claim 9, wherein: The second calculation module calculates the data in the original data set and the reference value, and determines the maximum heap and the minimum heap according to the calculation results, including: Calculate the difference between each data in the original data set and the reference value to obtain a difference data set; The difference values in the difference data set are divided into the maximum heap or the minimum heap according to a set preset value.
12. The device according to claim 11, wherein The second calculation module divides the difference values in the difference data set into the maximum heap or the minimum heap according to a preset value, including: Extracting differences greater than the preset value from the difference data set to form a maximum heap; Differences smaller than the preset value are extracted from the difference data set to form a minimum heap.
13. The device according to any one of claims 9 to 12, wherein: The balancing module dynamically adjusts the amount of data in the original data set by exchanging the amount of data in the maximum heap and the minimum heap, so that the amount of data in the original data set is balanced, including: Obtaining a maximum value in the maximum heap and a minimum value in the minimum heap to obtain a first value and a second value respectively; According to the data volume of the first value and the data volume of the second value, part of the data in the original data set corresponding to the first value is dynamically allocated to the original data set corresponding to the second value until all data volumes in the original data set are balanced.
14. The device according to claim 13, wherein: The balancing module dynamically allocates part of the data in the original data set corresponding to the first value to the original data set corresponding to the second value according to the data volume of the first value and the data volume of the second value until all data volumes in the original data set are balanced, including: If the data volume of the first value is greater than the data volume of the second value, extracting the same amount of data as the data volume of the second value from the original data set data corresponding to the first value, and allocating the data to the original data set data corresponding to the second value, so as to obtain updated data set data; Calculate the difference between the data volume of the first value and the data volume of the second value to obtain the remaining data volume of the first value; Add the remaining data amount of the first value to the maximum heap, re-obtain the maximum value in the maximum heap and the minimum value in the minimum heap, and obtain a new first value and a new second value; Part of the data in the data set corresponding to the new first value is reallocated to the data set corresponding to the new second value until all data amounts in the data set are balanced.
15. The device according to claim 13, wherein: The balancing module dynamically allocates part of the data in the original data set corresponding to the first value to the original data set corresponding to the second value according to the data volume of the first value and the data volume of the second value until all data volumes in the original data set are balanced, including: If the data volume of the first value is less than the data volume of the second value, extracting the same amount of data as the data volume of the first value from the original data set data corresponding to the first value, and allocating the data to the original data set data corresponding to the second value, so as to obtain updated data set data; Calculate the difference between the data volume of the second value and the data volume of the first value to obtain the data volume of the second value that differs; Add the data amount of the second value that differs to the minimum heap, re-obtain the maximum value in the maximum heap and the minimum value in the minimum heap, and obtain a new first value and a new second value; Part of the data in the data set corresponding to the new first value is reallocated to the data set corresponding to the new second value until all data amounts in the data set are balanced.
16. The device according to claim 9, wherein: The acquisition module acquires the data size of a plurality of multimodal data to be processed to obtain an original data set, including: Obtaining multimodal data output by the image and video coding model, and calculating the data size of the output multimodal data to obtain a calculation result; The standard deviation of the calculation result is calculated, and if the standard deviation exceeds a preset threshold, the data size of the corresponding multimodal data is used as the original data set.
17. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.
19. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.