Data processing method and device, electronic equipment and storage medium

By obtaining the equivalent value of the data to be collected from the storage device and the number of available threads, and selecting the target data for data acquisition, the problem of low thread utilization is solved, and the thread utilization rate is improved and data acquisition efficiency is improved.

CN120353604AActive Publication Date: 2025-07-22INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510828570.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-22
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

The inconsistent data volume of different storage devices in the storage system leads to a low thread utilization rate and a low thread utilization rate in the prior art.

Method used

By obtaining the equivalent value of the data to be collected by the storage device and the number of available threads, select the target data to be collected according to the trend of making the number of available threads close to the minimum value, and use the number of available threads to collect data to ensure that the number of threads is negatively correlated with the sum of the equivalent value of the target data.

Benefits of technology

It improves thread utilization, reduces thread idleness, and improves data acquisition efficiency and system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353604A_ABST
    Figure CN120353604A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device, electronic equipment and a storage medium, and relates to the technical field of data processing.The data processing method includes the steps that equivalent values of at least two pieces of to-be-collected data in at least two pieces of storage equipment are obtained, and available thread counts of the at least two pieces of storage equipment are obtained; the equivalent value is a thread count required for collecting the data to be collected; selecting at least one target to-be-collected data from the at least two to-be-collected data on the basis of the equivalent values of the at least two to-be-collected data according to the trend that the number of available threads approaches to the minimum value; the available thread count is in negative correlation with the sum of the equivalent values of the selected at least one target to-be-collected data; according to the method, the available thread count is utilized to perform data acquisition on the at least one target to-be-acquired data, so that thread idleness is reduced, and the thread utilization rate is improved. Therefore, the technical problem that the utilization rate of the threads is low can be solved, and the technical effect of improving the utilization rate of the threads is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, and particularly relates to a method and device for data processing, an electronic device, and a storage medium. Background Art

[0002] There are multiple storage devices in a storage system. To ensure the security of data in multiple storage devices, it is necessary to collect data from multiple storage devices for centralized management.

[0003] In related technologies of data processing, data collection is usually performed on storage devices one by one in a fixed order. Since the amounts of data stored in different storage devices are inconsistent, the threads consumed for data collection by different storage devices are inconsistent, resulting in low utilization rate of threads. Summary of the Invention

[0004] The present application provides a method and device for data processing, an electronic device, and a storage medium, so as to at least solve the problem of low utilization rate of threads in related technologies.

[0005] The present application provides a method for data processing, including: Obtaining equivalent values of at least two to-be-collected data in at least two storage devices, and obtaining the available thread numbers of at least two storage devices; the equivalent value is the number of threads required to collect the to-be-collected data; Based on the equivalent values of at least two to-be-collected data, selecting at least one target to-be-collected data from at least two to-be-collected data according to the trend of making the available thread number approach the minimum value; the available thread number is negatively correlated with the total sum of the equivalent values of the at least one selected target to-be-collected data; Using the available thread number to perform data collection on at least one target to-be-collected data.

[0006] The present application also provides a device for data processing, including: An obtaining unit, configured to obtain equivalent values of at least two to-be-collected data in at least two storage devices, and obtain the available thread numbers of at least two storage devices; the equivalent value is the number of threads required to collect the to-be-collected data; A selecting unit, configured to select at least one target to-be-collected data from at least two to-be-collected data based on the equivalent values of at least two to-be-collected data according to the trend of making the available thread number approach the minimum value; the available thread number is negatively correlated with the total sum of the equivalent values of the at least one selected target to-be-collected data; A collecting unit, configured to use the available thread number to perform data collection on at least one target to-be-collected data.

[0007] The present application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the steps of any of the above data processing methods when executing the computer program.

[0008] The present application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above data processing methods.

[0009] The present application also provides a computer program product including a computer program, which, when executed by a processor, implements the steps of any of the above data processing methods.

[0010] Through the present application, since the equivalent values of at least two data to be collected in at least two storage devices are obtained, and the available thread numbers of at least two storage devices are obtained; the equivalent value is the number of threads required to collect the data to be collected; based on the equivalent values of at least two data to be collected, at least one target data to be collected is selected from at least two data to be collected in a trend of making the available thread number approach the minimum value; the available thread number is negatively correlated with the total sum of the equivalent values of at least one selected target data to be collected; and the available thread number is used to perform data collection on at least one target data to be collected, the thread idleness is reduced, thereby improving the utilization rate of the thread. Therefore, the technical problem of low utilization rate of the thread can be solved, and the technical effect of improving the utilization rate of the thread can be achieved. Description of the Drawings

[0011] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0012] Figure 1 It is a schematic flowchart of a data processing method provided by an embodiment of the present application; Figure 2 It is a schematic flowchart of the entire process of a data processing method provided by an embodiment of the present application; Figure 3 It is a schematic structural diagram of a data processing device provided by an embodiment of the present application; Figure 4 It is a schematic structural diagram of another data processing device provided by an embodiment of the present application. Detailed Embodiments

[0013] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0014] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and not to describe a specific order or sequence.

[0015] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0016] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the method of data processing depends, the specific application environment architecture or specific hardware architecture will be described herein.

[0017] The embodiments of the present application provide a method of data processing. The method will be described in detail in combination with the execution process of the method of data processing.

[0018] Figure 1 It is a schematic flowchart of a method of data processing provided by an embodiment of the present application.

[0019] As Figure 1 shown, the method includes the following steps: Step 101, obtain the equivalent values of at least two pieces of data to be collected in at least two storage devices, and obtain the available number of threads of at least two storage devices; the equivalent value is the number of threads required to collect the data to be collected.

[0020] In this embodiment, the storage device can be various physical devices for storing data, such as hard disks, solid-state drives, tape libraries, etc. They are used to store different types and amounts of data to meet the data storage requirements in different business scenarios. The data to be collected refers to the data that needs to be collected and processed in the storage device, which can be files, database records, log information, etc. The equivalent value refers to the computing resources (such as the number of threads) required to collect a specific data. The equivalent value is determined according to factors such as data type, data size, and the complexity of data collection, reflecting the "workload" required to collect data. The number of threads is the computing resource unit for parallel processing of data during data collection. Each thread can execute tasks independently, and multiple threads can perform data collection simultaneously, thereby improving the efficiency of data collection. The available number of threads refers to the number of thread resources currently available for allocation in the system for data collection work. The available number of threads is limited by the system's hardware resources (such as the number of cores of the Central Processing Unit (CPU), memory, etc.).

[0021] Identify multiple storage devices included in the system, such as storage device A, storage device B, and storage device C, etc. In each storage device, identify the data that needs to be collected. For example, in storage device A, there is data a1 and a2 to be collected, and in storage device B, there is data b1 and b2 to be collected, etc., and record these data to be collected in a data list. For each data to be collected, calculate its equivalent value. The calculation method can be to comprehensively consider various factors, such as the size of the data volume, data type, and business metrics, etc. Taking the data volume as an example, if the data volume is larger, usually more time and resources are required for collection. Then, the size of the data volume of each data to be collected can be quantified first to determine the weight value it occupies in the dimension of the data volume. Similarly, for different data types, such as text data, image data, video data, etc., due to their different collection methods and complexities, corresponding weights are also assigned respectively. For business metrics, such as the urgency and importance of data, they are scored and quantified according to business requirements. Finally, the weight values in these different dimensions are weighted and calculated to obtain the equivalent value of the data to be collected. For example, for the data to be collected a1, its data volume weight is 0.4, data type weight is 0.3, and business metric weight is 0.3. After corresponding calculations, its equivalent value is 5. The system obtains the number of threads currently available for data collection tasks by means of real-time monitoring or periodically querying the task scheduling system. This number will change dynamically according to the running situation of the overall tasks in the system. For example, if there are a total of 20 threads in the system currently, and 8 of them are executing other non-data collection tasks, then the available number of threads at this time is 12.

[0022] By obtaining the equivalent value of each data to be collected and the number of available threads, it is possible to reasonably allocate threads according to the actual requirements of the data to be collected and the available situation of system resources. This avoids the problem of low thread utilization rate caused by collecting one by one in a traditional fixed order, enables the limited thread resources to be utilized more efficiently and reasonably, and improves the efficiency of the entire data collection process.

[0023] Step 102: Based on the equivalent values of at least two data to be collected, select at least one target data to be collected from the at least two data to be collected in a trend that makes the number of available threads approach the minimum value; the number of available threads is negatively correlated with the total equivalent value of the at least one target data to be collected selected.

[0024] Approaching the minimum value means that through task combination, the remaining available thread number is infinitely close to 0 (ideal state: total equivalent value = number of available threads), achieving the maximum utilization rate of thread resources. Negatively correlated means that there is an inverse relationship between the number of available threads and the total equivalent value of the selected target data to be collected. When the total equivalent value of the target data is large, the number of available threads in the system tends to be small, and vice versa.

[0025] First, calculate the equivalent value of the data to be collected in each storage device according to the aforementioned method, and at the same time obtain the current number of available threads through system monitoring or querying the task scheduling system. For example, it is calculated that the equivalent value of data A to be collected in storage device 1 is 5, and the equivalent value of data B to be collected is 3; the equivalent value of data C to be collected in storage device 2 is 4, and the equivalent value of data D to be collected is 6; the current number of available threads in the system is 10.

[0026] From all the data to be collected, find the data to be collected whose equivalent value is less than or equal to the number of available threads, and sort them in descending order of the equivalent value. If the number of available threads is 10, then in the above example, the equivalent values of all the data to be collected are less than or equal to 10, and the sorting is D (6), A (5), C (4), B (3). Start selecting from the data with the largest equivalent value. First, select data D, whose equivalent value is 6. At this time, the number of used threads is 6, and the remaining available threads are 4. Within the remaining available threads, continue to select from the remaining data to be collected the data with the largest equivalent value that does not exceed the remaining available threads. For example, at this time, the remaining available threads are 4. Among the remaining data to be collected, the equivalent value of data A, which is 5, exceeds the remaining available threads of 4 and cannot be selected; the equivalent value of data C to be collected meets the condition, so select data C to be collected. The cumulative number of used threads is 6 + 4 = 10, and the remaining available threads are 0. If there are still remaining available threads at this time, repeat this step until there are no remaining available threads or all the data to be collected have been considered. If the equivalent value of no data to be collected is less than or equal to the number of available threads, it may mean that the number of available threads is too small to meet the minimum thread requirements of a single piece of data to be collected. In this case, it is necessary to wait for the threads to finish collecting the data and then update the number of available threads so that the updated number of threads can process the remaining data to be collected.

[0027] Selecting the target data to be collected in a way that the number of available threads tends to approach the minimum value can complete as many data collection tasks as possible within the limited thread resources, avoid the idleness and waste of threads, improve the utilization rate of threads, and enhance the overall performance and efficiency of the system.

[0028] Step 103: Use the number of available threads to perform data collection on at least one target data to be collected.

[0029] The target data to be collected refers to the data determined from among numerous data to be collected after the previous selection process and that meets the requirement of being efficiently collectable under the current constraint of the number of available threads. Data collection: refers to the process of extracting information from the target data to be collected. Data collection can include multiple links such as reading, processing, storing, and analyzing. The purpose is to collect and process the data scattered in the system or on the storage device for subsequent analysis and use.

[0030] The system first obtains the current number of available threads, which can be obtained through the thread management module of the operating system. Assuming that the number of CPU cores in the current system is N and each core supports M threads, the number of available threads may be affected by factors such as system load and other applications' occupancy, and is usually N*M minus the number of threads currently in use. The system selects at least one target data to be collected according to the data collection requirements and priorities. The selection basis of the target data may include the size, complexity, storage location, access frequency, etc. of the data. For example, if some data is stored on a remote server, these data may need to be collected preferentially. The system reasonably allocates thread resources according to the current number of available threads. For each target data to be collected, the system will allocate an appropriate number of threads according to its complexity, size and priority. If the data collection task is relatively complex or contains multiple processing steps (such as decompression, large-scale data conversion), the system may allocate more threads to this task. The specific thread allocation algorithm can be based on strategies such as priority weighted round-robin method, dynamic load balancing, etc. For example, in the case of a large amount of data, the system can adopt a divide-and-conquer strategy, split the data into multiple subtasks, and allocate different threads for parallel processing. Once the thread resources are allocated, the system starts to execute the data collection task. Each thread is responsible for collecting a part of the target data, and the task execution order and dependency relationship between the threads will be managed by the scheduling algorithm. If there are dependency relationships between the collection tasks, the system will ensure that the threads are executed in the correct order to avoid conflicts and errors. During the data collection process, the system dynamically adjusts the number of available threads according to the progress of the tasks. If some threads have completed their tasks, the system can allocate the idle threads to other collection tasks that have not been completed yet, ensuring that the overall collection task is accelerated. During the collection process, if some tasks encounter bottlenecks (such as delays), the system will also appropriately reduce the number of threads for these tasks and allocate the resources to other tasks that can be processed in parallel. When all the tasks of the target data to be collected are completed, the system summarizes the collection results, stores the data in the specified storage location, and performs subsequent data processing and analysis. The end of the collection task also means the recycling of the number of available threads, and the thread resources can be used for other tasks.

[0031] By reasonably allocating the number of available threads, the system can achieve maximized parallelism during the data collection process. The efficient utilization of thread resources can significantly improve the speed of data collection, especially when dealing with a large amount of data or data with high-frequency access, and the system can complete the tasks quickly.

[0032] Through this application, since the equivalent values of at least two to-be-acquired data in at least two storage devices are obtained, and the available thread counts of at least two storage devices are obtained; the equivalent value is the number of threads required to acquire the to-be-acquired data; based on the equivalent values of at least two to-be-acquired data, at least one target to-be-acquired data is selected from at least two to-be-acquired data in a trend of making the available thread count approach the minimum value; the available thread count is negatively correlated with the total sum of the equivalent values of the at least one selected target to-be-acquired data; the available thread count is used to perform data acquisition on at least one target to-be-acquired data, reducing thread idleness, thereby improving the utilization rate of threads. Therefore, the technical problem of low thread utilization rate can be solved, and the technical effect of improving the utilization rate of threads can be achieved.

[0033] As a refinement of step 101, when executing to obtain the equivalent values of at least two to-be-acquired data in at least two storage devices, it can be implemented in but not limited to the following ways, including: obtaining the time-consuming coefficients of at least two to-be-acquired data, and obtaining the data type weights and business indicator weights of at least two to-be-acquired data; calculating the equivalent value according to the time-consuming coefficients, data type weights and business indicator weights.

[0034] The time-consuming coefficient reflects the relative amount of time consumed by the to-be-acquired data during the acquisition process, and can be used to measure the differences in acquisition time consumption of different to-be-acquired data. The larger its value, the longer the acquisition time consumption of the to-be-acquired data. The data type weight refers to the weight allocation of the importance or priority according to different types of data. Different types of data have different importance in the business process, so different weights need to be assigned to each data type. The higher the weight, the greater the impact of the data on the business. The business indicator weight refers to the weight assigned to different business indicators according to business requirements. The business indicator weight measures the contribution degree of data to business goals, and is usually related to business goals and key performance indicators. The higher the weight, the greater the impact of the data on business success. The business indicator weight reflects the importance of the to-be-acquired data to the business. According to business requirements and the use of data, different weights are assigned to different data to highlight the priority of key data in the acquisition process.

[0035] Collect the historical acquisition time consumption records of the data to be acquired, calculate their average value as the historical average acquisition time consumption; at the same time, find the median from all the historical acquisition time consumption records and determine it as the benchmark acquisition time consumption. Then, divide the historical average acquisition time consumption by the benchmark acquisition time consumption to obtain the time consumption coefficient. For example, if the historical average acquisition time consumption of a certain data to be acquired is 100 seconds and the benchmark acquisition time consumption is 80 seconds, then the time consumption coefficient is 100÷80 = 1.25. According to the common data types in the storage system, define the weights of each data type in advance. For example, assign a weight of 0.3 to text data, 0.5 to image data, 0.7 to video data, etc. According to the actual type of the data to be acquired, directly obtain the corresponding data type weight value, which is determined by the business department through evaluating and scoring according to the importance of the data to the business. For example, assign a higher weight such as 0.8 to core business data and a lower weight such as 0.3 to auxiliary business data. Substitute the obtained time consumption coefficient, data type weight, and business metric weight into the formula for calculation. First, perform an addition calculation on the data type weight and the business metric weight to obtain the addition result; then, multiply the addition result by the time consumption coefficient to obtain the equivalent value. For example, if the time consumption coefficient of a certain data to be acquired is 1.25, the data type weight is 0.5, and the business metric weight is 0.6, then the addition result is 0.5 + 0.6 = 1.1, and the equivalent value is 1.25×1.1 = 1.375.

[0036] By comprehensively considering the three key factors of the time consumption coefficient, data type weight, and business metric weight, it is possible to more accurately evaluate the actual resource amount required for acquiring different data to be acquired, avoid the inaccuracy problem caused by estimating the number of threads based on a single factor (such as the data volume), and improve the rationality of resource allocation.

[0037] As a refinement of the above embodiment, when calculating the equivalent value according to the time consumption coefficient, data type weight, and business metric weight, it can be implemented in but not limited to the following ways, including: performing an addition calculation on the data type weight and the business metric weight to obtain the addition result; performing a product calculation on the addition result and the time consumption coefficient to obtain the equivalent value.

[0038] For the sake of easy understanding, an example is provided. Suppose the data type weight of a certain data to be acquired is 0.3 (for example, text data), the business metric weight is 0.5 (for example, relatively important business data), and the time consumption coefficient is 1.25 (obtained based on the ratio of the historical average acquisition time consumption to the benchmark acquisition time consumption). Add the data type weight 0.3 and the business metric weight 0.5 to obtain the addition result 0.8. Multiply the addition result 0.8 by the time consumption coefficient 1.25 to obtain the equivalent value 1.0. Record the calculated equivalent value 1.0 in the attribute or configuration file of the data to be acquired for subsequent use as a basis for thread allocation during data acquisition.

[0039] By adding to calculate the comprehensive weight and then multiplying it by the time-consuming coefficient to obtain the equivalent value, it is possible to comprehensively consider the time consumption of data collection and the business importance, and provide a more accurate assessment of resource requirements for data collection tasks.

[0040] As a refinement of the above embodiments, when executing to obtain the time-consuming coefficients of at least two data to be collected, the following methods can be adopted but are not limited to, including: obtaining the historical average collection time of at least two data to be collected, and obtaining the benchmark collection time of at least two data to be collected; the benchmark collection time is the median of all historical collection times of at least two data to be collected; performing a quotient calculation on the historical average collection time and the benchmark collection time to obtain the time-consuming coefficient.

[0041] The historical average collection time refers to the average time-consuming value calculated through multiple historical collection records for a certain data to be collected, which reflects the average time cost required for this data in previous collection processes. The benchmark collection time refers to the value in the middle position after arranging these values in ascending order among all historical collection time records of a certain data to be collected. If the number of records is odd, it is the middle value; if it is even, the average of the two middle values is usually taken. The benchmark collection time can better represent the typical level of the data collection time-consuming and is less affected by extreme values. The time-consuming coefficient refers to the quotient of the historical average collection time and the benchmark collection time, which is used to measure the relative relationship between the data to be collected and the typical level in terms of collection time-consuming. If the time-consuming coefficient is greater than 1, it means that the data collection takes a long time; if it is less than 1, it means that the collection time is short.

[0042] After the system completes the acquisition task of the data to be acquired in each storage device each time, it will automatically record the acquisition time of each data to be acquired this time and store it in a historical acquisition time database. For example, for the data to be acquired 1 in storage device A, the acquisition time for the first time is 10 seconds, the second time is 15 seconds, and the third time is 12 seconds. Then these three acquisition times will be recorded in the database. When it is necessary to obtain the historical average acquisition time of a certain data to be acquired, the system will extract all the historical acquisition time records of this data from the database, and then add up these recorded values and divide by the total number of records. Taking the data to be acquired 1 as an example, its historical average acquisition time = (10 + 15 + 12) / 3 = 12.3 seconds. The system also obtains all the historical acquisition time records of this data to be acquired from the database and sorts them in ascending order. If the number of records is odd, such as the 5 acquisition time records sorted as [8, 10, 12, 14, 16], the reference acquisition time is the middle value of 12 seconds; if the number of records is even, such as the 4 acquisition time records sorted as [9, 11, 13, 15], the reference acquisition time is the average of the two middle values (11 + 13) / 2 = 12 seconds. After obtaining the historical average acquisition time and the reference acquisition time, divide the historical average acquisition time by the reference acquisition time to obtain the acquisition time coefficient. For the data to be acquired 1, assuming its historical average acquisition time is 12.3 seconds and the reference acquisition time is 12 seconds, then the acquisition time coefficient = 12.3 / 12 = 1.025.

[0043] Through the calculation of the historical average acquisition time and the reference acquisition time, it is possible to comprehensively and accurately understand the historical performance and typical level of the data to be acquired in terms of acquisition time, providing a reliable basis for subsequent reasonable arrangement of acquisition tasks and resource allocation.

[0044] As a refinement of step 102, when selecting at least one target data to be acquired from at least two data to be acquired according to the equivalent values of at least two data to be acquired and in the trend of making the number of available threads approach the minimum value, the following methods can be adopted but are not limited to: determining whether there is at least one target equivalent value in the equivalent values; the target equivalent value is an equivalent value less than or equal to the number of available threads; in the case where there is at least one target equivalent value in the equivalent values, select the target data to be acquired from the data to be acquired corresponding to at least one target equivalent value in descending order of the target equivalent value; wherein, the sum of the target equivalent values of the selected target data to be acquired does not exceed the number of available threads; select the target data to be acquired from the remaining data to be acquired in the trend of making the number of available threads approach the minimum value.

[0045] For ease of understanding, an example is provided. The system determines that the number of available threads is 10 based on the current running state. For multiple data to be collected, their equivalent values are as follows: Data A is 5, Data B is 3, Data C is 4, Data D is 6, and Data E is 2. The system filters out the data to be collected whose equivalent values are less than or equal to the number of available threads (10), namely Data A (5), Data B (3), Data C (4), Data D (6), and Data E (2). The system sorts the filtered data in descending order of equivalent value: Data D (6), Data A (5), Data C (4), Data B (3), Data E (2). Select the data in the sorted order one by one. First, select Data D (6), the number of used threads becomes 6, and the remaining available threads are 4. Then, try to select Data A (5), but at this time, the number of used threads will reach 11, exceeding the available threads 10, so it cannot be selected. Then, try to select Data C (4), the number of used threads becomes 10, and the remaining available threads are 0. The total target equivalent value of the target data to be collected is 6 + 4 = 10, which is exactly equal to the available threads 10. If new thread resources are released during the data collection process and the number of available threads increases, the system can select new target data to be collected from the remaining data to be collected again according to the above rules. The system continuously monitors the data collection progress and thread usage, and loops through the above steps until all data to be collected are completed.

[0046] By reasonably selecting the target data to be collected, it is ensured that under the limit of the number of available threads, as much thread resources as possible are utilized for data collection, reducing thread idleness.

[0047] In practical applications, the equivalent value of the data to be collected is dynamically adjusted, and the adjustment method can be implemented but not limited to the following methods, including: obtaining the target mode coefficient of the target performance mode of the storage device where the data to be collected is located; different performance modes correspond to different mode coefficients; performing a product calculation on the equivalent value of the data to be collected and the target mode coefficient to obtain the adjusted equivalent value, so as to perform data collection on the data to be collected based on the adjusted equivalent value.

[0048] The target performance mode refers to the performance state mode in which the storage device is currently located. Different performance modes represent the operating performance of the storage device under different conditions, such as high-speed mode, normal mode, energy-saving mode, etc. The target mode coefficient refers to the coefficient corresponding to the target performance mode of the storage device, which is used to quantify the performance differences under different performance modes and adjust the equivalent value of data collection. Different performance modes correspond to different mode coefficients. For example, the mode coefficient corresponding to the high-speed mode may be 1.2, the mode coefficient corresponding to the normal mode is 1, and the mode coefficient corresponding to the energy-saving mode is 0.8. The adjusted equivalent value refers to the value obtained by multiplying the original equivalent value of the data to be collected by the target mode coefficient, which reflects the number of thread resources actually required to collect this data under the current performance mode of the storage device.

[0049] The system has the function of detecting the performance state of the storage device. By running a performance detection program and analyzing the real-time operation data of the storage device, such as indicators like CPU usage rate, memory occupancy, data read / write speed, etc., it determines the current performance mode of the storage device. For example, if a storage device has a high CPU usage rate and a fast data read / write speed, the system determines that it is in the high-speed mode. According to the pre-set mapping relationship table between the performance mode and the mode coefficient, it finds the target mode coefficient corresponding to the current target performance mode. Assuming that the mapping relationship table stipulates that the high-speed mode corresponds to a mode coefficient of 1.2, the normal mode corresponds to a mode coefficient of 1, and the energy-saving mode corresponds to a mode coefficient of 0.8, then for a storage device in the high-speed mode, the obtained target mode coefficient is 1.2. Obtain the original equivalent value of the data to be collected, which is calculated based on the attributes of the data itself (such as data volume, data type, business indicators, etc.) without considering the performance factors of the storage device. For example, the original equivalent value of the data to be collected is 5. Perform a product calculation on the original equivalent value and the obtained target mode coefficient, that is, 5×1.2 = 6, and the adjusted equivalent value is 6. The system makes a decision on whether to perform data collection on this data to be collected based on the adjusted equivalent value, and when scheduling data collection tasks and allocating thread resources, it uses the adjusted equivalent value as an important reference basis to reasonably arrange data collection tasks and allocate thread resources.

[0050] Considering the current performance mode of the storage device, adjusting the equivalent value through the target mode coefficient enables the system to more accurately allocate thread resources for data collection tasks, avoiding uneven or unreasonable resource allocation caused by performance differences of the storage device.

[0051] As a refinement of the above embodiments, when executing to obtain the target mode coefficient of the target performance mode of the storage device where the data to be collected is located, it can be implemented in but not limited to the following ways, including: obtaining the target health degree of the storage device where the data to be collected is located; according to the pre-established mapping relationship between the health degree and the performance mode, finding the target mode coefficient corresponding to the target health degree.

[0052] The target health degree is a quantitative index reflecting the current health status of the storage device, usually obtained by comprehensively evaluating the hardware status of the storage device (such as the wear degree of the disk, the error rate of the memory, etc.) and system operation indexes (such as response latency, the frequency of error codes, etc.) through professional tools. The higher the value, the better the health status of the storage device. The mapping relationship is a pre-established corresponding rule that associates different health degrees of the storage device with corresponding performance modes, and is used to quickly determine the performance mode that the storage device should match under a specific health condition. The target mode coefficient is the coefficient corresponding to the target performance mode of the storage device, and is used to adjust the equivalent value of the data to be collected according to the performance mode of the storage device.

[0053] The system uses a professional storage device health assessment tool to regularly detect the hardware status and system operation indexes of the storage device. The hardware status includes the number of bad blocks on the disk, the error rate of the memory, etc.; the system operation indexes include the average response time, data transfer rate, etc. The assessment tool comprehensively calculates the collected indexes according to the pre-set weights and calculation rules to obtain the target health degree of the storage device. For example, the weights are set as 30% for disk wear, 20% for memory error rate, 30% for response latency, and 20% for data transfer rate. After the indexes are standardized, the target health degree is obtained as 75 by adding them according to the weights. Establish the health degree - performance mode mapping relationship: The system pre-establishes the mapping relationship between the health degree and the performance mode according to the historical data and performance test results of the storage device. For example, when the health degree is greater than 80, it corresponds to the high-speed mode; when the health degree is between 60 and 80, it corresponds to the normal mode; when the health degree is less than 60, it corresponds to the energy-saving mode. After obtaining the target health degree of the storage device, find the corresponding performance mode according to the established mapping relationship, and then determine the target mode coefficient. If the target health degree is 75, it corresponds to the normal mode according to the mapping relationship, and the target mode coefficient is 1. Use the obtained target mode coefficient to adjust the equivalent value of the data to be collected, and the new equivalent value is used to guide the thread allocation and scheduling of the data collection task. For example, for data with an original equivalent value of 5, the adjusted equivalent value in the normal mode is still 5×1 = 5, and the system arranges the data collection task reasonably according to this value.

[0054] Through the target health degree assessment and the mapping relationship, accurately obtain the target mode coefficient, and then reasonably adjust the equivalent value, so that the system can accurately allocate resources and schedule tasks according to the actual performance status of the storage device.

[0055] For the convenience of better understanding the whole process of data processing, as Figure 2 shown Figure 2It is a schematic flowchart of the entire process of data processing provided by the embodiments of this application. Task classification and quantification module: 1. Acquisition resource (i.e., data to be acquired) definition: Storage device: The storage device is regarded as the top-level resource. For the multi-device management software, each added storage device is a storage resource. Resource type: The subordinate resources of the storage device, corresponding to the specific resources in the three business scenarios of block, file, and object storage, such as basic clusters, storage pools, disks, ports, volumes in the block scenario, file systems and shared directories in the file scenario, and buckets in the object scenario. Business metrics: Specific metrics in different business scenarios, such as capacity, performance, alarms, etc.

[0056] Based on the above classification, it is now defined that an "acquisition resource" is formed by combining a storage device - resource type - business resource, resulting in an independent acquisition resource identifier. For example, Storage A - Volume - Performance, Storage A - Cluster - Alarm, and Storage B - Volume Performance are all independent acquisition resources. The "acquisition resource" mentioned later refers to the concept defined by splitting here.

[0057] 2. Fixed classification benchmark: Based on the actual business scenario, fixed equivalent definitions are now given for the resource types and business metrics in the acquisition resources. Resource type weight (RW: Resource Weight), the resource type weight represents the acquisition complexity differences at different resource levels (such as clusters, storage pools, volumes, file systems). Examples are as follows: RW of Volume = 4.0, RW of Pool = 1.0, RW of FileSystem = 3.0. Business metric weight (IW: Indicator Weight). The indicator type weight represents the acquisition time differences for different metrics (such as performance, capacity). For example, IW of the performance metric = 3.0, IW of the capacity metric = 1.0.

[0058] 3. Dynamic Grading Benchmark: During the actual data collection process, due to the different magnitudes of different resources, there will be significant differences in the corresponding collection durations. For example, if Storage A has 1000 volumes and Storage B has 30 volumes, then there are obvious differences in the collection consumption for the same volume-performance collection. Therefore, in the calculation of collection consumption, a dynamic grading strategy is introduced on the basis of a fixed grading benchmark. The calculation logic of this strategy is as follows: Historical Duration Record: Using the sliding window algorithm, the system independently records the historical average duration (unit: minute) for each collection resource. For example: Collection Resource 1 (Storage A - Volume - Performance): Average duration is 10 minutes. Collection Resource 2 (Storage A - Storage Pool - Performance): Average duration is 2 minutes. Duration Normalization: Take the median of the historical durations of all collection tasks in the current system as the benchmark duration, and calculate the standard duration coefficient of the collection resource: CT = Average Duration / Benchmark Duration. Now assume the base duration is 2, then the CT of the above Collection Resource 1 is 5, and the CT of Collection Resource 2 is 1. Equivalent Value Calculation: Finally, the thread equivalent value result for each collection resource is: Thread Equivalent = CT × (RW + IW). Combining the above, the thread equivalent value of Collection Resource 1 can be finally calculated as 5 × (4 + 3) = 35 equivalents, while that of Collection Resource 2 is 1 × (1 + 3) = 4 equivalents.

[0059] Intelligent Combined Scheduling Module: As Figure 2 shown, when the collection task starts, the program will start processing according to the following logic: while there are idle resources in the thread pool: 1. Select the task with the largest equivalent value from the queue; 2. If the equivalent value of this task ≤ the remaining number of threads (i.e., the available number of threads): Add it to the current task package; After adding the task with the largest equivalent value to the current task package, the remaining number of threads is equal to the remaining number of threads before adding the task with the largest equivalent value minus the equivalent value of the task with the largest equivalent value; 3. Otherwise: Skip this task and check the task with the second largest equivalent value. In actual execution, the following optimization logic processing needs to be carried out: 1. Prioritize combining multiple tasks of the same storage device to reduce connection overhead; 2. If there is a task whose equivalent value is higher than the total number of threads, execute it separately; The actual execution example is as follows: Now assume the total thread resources are 80, and there are four collection resources, namely: Collection Resource A: 35; Collection Resource B: 28; Collection Resource C: 20; Collection Resource D: 15; The corresponding task combination process: Select Task A (35) → Remaining 45; Select Task B (28) → Remaining 17; Select Task D (15) → Remaining 2 (insufficient to execute Task C); Total utilization rate: (35 + 28 + 15) / 80 = 97.5%.

[0060] Dynamic Policy Engine: Considering the importance and necessity of the normal operation of the storage service, the concept of storage dynamic policy is added on the above basis. For the storage managed by multiple devices, the acquisition policy can be configured, including high performance, standard performance, low performance, and dynamic floating. Configuration description: High performance, this storage allows configuration on the standard basis, and the maximum over-allocation ratio is 1.1 to 1.5 times the number of physical threads (default 1.2). Reduce the sampling interval to obtain more performance data. Standard performance, no additional processing is done according to the normal calculated standard equivalent. Low performance, this storage reduces the thread occupancy ratio on the standard basis: 40% - 80% (default 60%), increases the sampling interval, and reduces the number of collected performance data. Dynamic floating, linked to the health function of the system, and dynamically adjusted according to preset conditions and linked to the system health function. This application does not provide more explanations for the storage health function. The health branch of the storage ranges from 0 to 100 and is affected by multiple dimensions such as storage alarms, performance indicators, and hardware status, which is a visual manifestation of the storage operation state. Floating adjustment rules: Switch trigger conditions (taking storage A as an example): High performance, when the system is in the default window period (such as 1:00 - 5:00 in the morning). Manually trigger emergency collection (such as fault diagnosis). Standard performance: The current mode is high performance, and the high-performance window period is exited. The health returns to the healthy range (>=75). Low performance: The storage health is lower than the healthy range (<75), and the storage device triggers an overload alarm. 3. The relationship between the dynamic policy and the scheduling module. The equivalent of the acquisition resource will be processed twice according to the policy of the corresponding storage, and this is the actual equivalent value of the acquisition resource at this time. Its actual execution process is as follows: The mode controller determines the mode according to the device health (such as storage A → low-performance mode); Adjust the equivalent base: Corrected equivalent = Original equivalent × Mode coefficient. Example of mode coefficient: High-performance mode: 1.2, standard mode: 1.0, low-power mode: 0.6. The dynamic packer uses the corrected equivalent for task sorting and packing. For example, the original equivalent of acquisition resource 1 mentioned above is 35. At this time, the storage is in the low-performance mode and the coefficient is 0.6, then the final corrected equivalent of acquisition resource 1 is 21.

[0061] By dynamically analyzing the actual resource consumption characteristics of different collection tasks and intelligently adjusting the thread allocation strategy, this application effectively solves the core pain points of low resource utilization and unbalanced collection efficiency in traditional storage monitoring systems. Compared with the fixed polling mechanism, this solution can automatically identify high-load tasks and preferentially allocate resources, reducing thread idleness and task queuing. At the same time, in high-concurrency scenarios, it actively restricts the resource occupancy of non-critical tasks to avoid interfering with the performance of the storage device itself. By intelligently combining collection tasks of different magnitudes, the system can fully exploit the hardware potential and significantly improve the processing throughput while ensuring data integrity. The adaptive mode control mechanism enables the system to flexibly handle complex scenarios such as business peaks and equipment failures. It can not only quickly complete the collection of critical data in emergency situations but also actively degrade when the device load is high to ensure business continuity.

[0062] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0063] An embodiment of this application also provides a data processing device. Figure 3 As shown in the structural schematic diagram of a data processing device provided by an embodiment of this application, Figure 3 it includes: An acquisition unit 21, configured to acquire the equivalent values of at least two pieces of data to be collected in at least two storage devices, and acquire the available thread numbers of at least two storage devices; the equivalent value is the number of threads required to collect the data to be collected. A selection unit 22, configured to select at least one target data to be collected from at least two pieces of data to be collected based on the equivalent values of at least two pieces of data to be collected, in a trend that makes the available thread number approach the minimum value; the available thread number is negatively correlated with the total equivalent value of the at least one selected target data to be collected. A collection unit 23, configured to use the available thread number to collect data for at least one target data to be collected.

[0064] Through the present application, since the equivalent values of at least two to-be-acquired data in at least two storage devices are obtained, and the available thread numbers of at least two storage devices are obtained; the equivalent value is the number of threads required to acquire the to-be-acquired data; based on the equivalent values of at least two to-be-acquired data, at least one target to-be-acquired data is selected from at least two to-be-acquired data in a trend of making the available thread number approach the minimum value; the available thread number is negatively correlated with the sum of the equivalent values of the selected at least one target to-be-acquired data; the available thread number is used to perform data acquisition on at least one target to-be-acquired data, reducing thread idleness, thereby improving the utilization rate of threads. Therefore, the technical problem of low utilization rate of threads can be solved, and the technical effect of improving the utilization rate of threads can be achieved.

[0065] Further, in a possible implementation manner of this embodiment, as Figure 4 shown, the obtaining unit 21 includes: An obtaining module 211, configured to obtain the time-consuming coefficients of at least two to-be-acquired data, and obtain the data type weights and service index weights of at least two to-be-acquired data; A calculating module 212, configured to calculate the equivalent value according to the time-consuming coefficient, the data type weight, and the service index weight.

[0066] Further, in a possible implementation manner of this embodiment, as Figure 4 shown, the calculating module 212 is further configured to: Perform a summation calculation on the data type weight and the service index weight to obtain a summation result; Perform a product calculation on the summation result and the time-consuming coefficient to obtain the equivalent value.

[0067] Further, in a possible implementation manner of this embodiment, as Figure 4 shown, the obtaining module 211 is further configured to: Obtain the historical average acquisition time of at least two to-be-acquired data, and obtain the benchmark acquisition time of at least two to-be-acquired data; the benchmark acquisition time is the median of all historical acquisition times of at least two to-be-acquired data; Perform a quotient calculation on the historical average acquisition time and the benchmark acquisition time to obtain the time-consuming coefficient.

[0068] Further, in a possible implementation manner of this embodiment, as Figure 4 shown, the selecting unit 22 is further configured to: Determine whether there is at least one target equivalent value among the equivalent values; the target equivalent value is an equivalent value that is less than or equal to the available thread number; When there is at least one target equivalent value in the equivalent values, the target to-be-collected data is selected from the to-be-collected data corresponding to at least one target equivalent value in descending order of the target equivalent value; wherein, the total target equivalent value of the selected target to-be-collected data does not exceed the number of available threads; The target to-be-collected data is selected from the remaining to-be-collected data in a trend that makes the number of available threads approach the minimum value.

[0069] Further, in a possible implementation manner of this embodiment, as Figure 4 shown, the apparatus further includes: The obtaining unit 21 is further configured to obtain a target mode coefficient of a target performance mode of the storage device where the to-be-collected data is located; different performance modes correspond to different mode coefficients; The calculation unit 24 is configured to perform a product calculation on the equivalent value of the to-be-collected data and the target mode coefficient to obtain an adjusted equivalent value, so as to perform data collection on the to-be-collected data based on the adjusted equivalent value.

[0070] Further, in a possible implementation manner of this embodiment, as Figure 4 shown, the obtaining unit 21 is further configured to: Obtain the target health degree of the storage device where the to-be-collected data is located; According to the pre-established mapping relationship between the health degree and the performance mode, find the target mode coefficient corresponding to the target health degree.

[0071] For the description of the features in the corresponding embodiment of the data processing apparatus, reference may be made to the relevant description in the corresponding embodiment of the data processing method, which will not be elaborated here one by one.

[0072] An embodiment of the present application further provides an electronic device, including a memory and a processor. The memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments of data processing.

[0073] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps in any one of the above method embodiments of data processing when running.

[0074] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk or optical disc, and other various media that can store computer programs.

[0075] Embodiments of the present application also provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the method embodiments of the above data processing are implemented.

[0076] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the method embodiments of the above data processing are implemented.

[0077] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0078] The above has introduced in detail a data processing method, apparatus, electronic device, and storage medium provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method for data processing, characterized in that, Including: Obtaining equivalent values of at least two pieces of data to be collected in at least two storage devices, and obtaining the available thread counts of the at least two storage devices; The equivalent value is the number of threads required to collect the data to be collected; Based on the equivalent values of at least two pieces of data to be collected, selecting at least one target data to be collected from the at least two pieces of data to be collected in a trend of making the available thread count approach the minimum value; The available thread count is negatively correlated with the total sum of the equivalent values of the selected at least one target data to be collected; Using the available thread count to perform data collection on the at least one target data to be collected.

2. The method for data processing according to claim 1, wherein The obtaining equivalent values of at least two pieces of data to be collected in at least two storage devices includes: Obtaining the time-consuming coefficients of the at least two pieces of data to be collected, and obtaining the data type weights and service index weights of the at least two pieces of data to be collected; Calculating the equivalent value according to the time-consuming coefficient, the data type weight, and the service index weight.

3. The method for data processing according to claim 2, wherein The calculating the equivalent value according to the time-consuming coefficient, the data type weight, and the service index weight includes: Performing a summation calculation on the data type weight and the service index weight to obtain a summation result; Performing a product calculation on the summation result and the time-consuming coefficient to obtain the equivalent value.

4. The method for data processing according to claim 2, wherein The obtaining the time-consuming coefficients of the at least two pieces of data to be collected includes: Obtaining the historical average collection time of the at least two pieces of data to be collected, and obtaining the benchmark collection time of the at least two pieces of data to be collected; the benchmark collection time is the median of all historical collection times of the at least two pieces of data to be collected; Performing a quotient calculation on the historical average collection time and the benchmark collection time to obtain the time-consuming coefficient.

5. The method for data processing according to claim 1, characterized in that, The selecting at least one target data to be collected from the at least two pieces of data to be collected based on the equivalent values of at least two pieces of data to be collected in a trend of making the available thread count approach the minimum value includes: Determining whether there is at least one target equivalent value among the equivalent values; the target equivalent value is an equivalent value less than or equal to the available thread count; In the case where there is the at least one target equivalent value among the equivalent values, selecting the target data to be collected from the data to be collected corresponding to the at least one target equivalent value in descending order of the target equivalent value; wherein, the total sum of the target equivalent values of the selected target data to be collected does not exceed the available thread count; Selecting the target data to be collected from the remaining data to be collected in a trend of making the available thread count approach the minimum value.

6. The method for data processing according to claim 1, characterized in that The method further includes: Obtaining the target mode coefficient of the target performance mode of the storage device where the data to be collected is located; different performance modes correspond to different mode coefficients; Performing a product calculation on the equivalent value of the data to be collected and the target mode coefficient to obtain an adjusted equivalent value, so as to perform data collection on the data to be collected based on the adjusted equivalent value.

7. The method for data processing according to claim 6, wherein The obtaining the target mode coefficient of the target performance mode of the storage device where the data to be collected is located includes: Obtaining the target health degree of the storage device where the data to be collected is located; According to the pre-established mapping relationship between the health degree and the performance mode, search for the target mode coefficient corresponding to the target health degree.

8. A data processing device, characterized in that, It includes: An acquisition unit, configured to acquire the equivalent values of at least two pieces of data to be acquired in at least two storage devices, and acquire the available thread numbers of the at least two storage devices; The equivalent value is the number of threads required to acquire the data to be acquired. A selection unit, configured to select at least one target data to be acquired from the at least two pieces of data to be acquired based on the equivalent values of the at least two pieces of data to be acquired, in a trend of making the available thread number approach the minimum value; The available thread number is negatively correlated with the total sum of the equivalent values of the at least one selected target data to be acquired. An acquisition unit, configured to use the available thread number to perform data acquisition on the at least one target data to be acquired.

9. An electronic device, characterized in that, It includes: A memory, configured to store a computer program; A processor, configured to implement the steps of the data processing method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the data processing method according to any one of claims 1 to 7 when executed by a processor.

Citation Information

Patent Citations

  • System and method for monitoring and scheduling of middleware threads

    CN103428272A

  • Embedded PLC (Programmable Logic Controller) engine implementation method and engine

    CN107315385A

  • Data processing method and device in multi-thread environment

    CN110058940A

  • Thread resource scheduling method and device, storage medium and electronic equipment

    CN114237895A

  • Batch data processing method and apparatus, computer device and storage medium

    WO2020228177A1