Big data cleaning processing method based on artificial intelligence
By introducing real-time adjustment factors and grouping mechanisms in big data cleaning processing, dynamically calculate the active coefficient and adjust resource allocation, the problem that traditional resource allocation methods are difficult to adapt to dynamic changes in tasks is solved, and the efficiency and resource utilization of big data cleaning are improved.
Patent Information
- Application Number
- CN202510282162.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-13
AI Technical Summary
The traditional big data cleaning task resource allocation method is difficult to adapt to dynamic changes in the task, resulting in oversupply or insufficient resources in certain stages of the task execution process, affecting the task execution efficiency and success rate.
Using artificial intelligence-based big data cleaning and processing methods, we use real-time adjustment of factors and grouping big data tasks, accurately calculate the activity coefficient, and dynamically adjust resource allocation to ensure that resources match the highest priority cloud computing resources.
It improves the efficiency of big data cleaning, ensures efficient utilization of resources, can be applicable to big data cleaning tasks of different types and properties, and realizes the transformation from static priority adjustment to dynamic priority adjustment.
Smart Images

Figure CN120144926A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data cleaning, and particularly to a big data cleaning and processing method based on artificial intelligence. Background Art
[0002] In the big data era, data cleaning is an indispensable part of the data analysis and mining process. With the explosive growth of the data volume, big data cleaning tasks have become increasingly complex and huge, and the demand for computing resources has also increased day by day. However, traditional resource allocation methods are often based on static task characteristics and preset resource allocation strategies, and it is difficult to adapt to the dynamically changing requirements of big data cleaning tasks. In big data processing systems, resource allocation is usually based on the initial characteristics of tasks and the estimated execution time. This method works well in scenarios where task characteristics are relatively stable and resource requirements are predictable. However, in big data cleaning tasks, the characteristics and resource requirements of tasks often change continuously with the inflow of data and the progress of the cleaning process. For example, some big data cleaning tasks may require a large amount of computing resources at the beginning to process complex data conversion and cleaning rules, while only less resources may be needed in the subsequent stage for simple data verification and storage.
[0003] For example, Chinese Patent Application No. 202210564357.7 discloses a big data cleaning task processing method and a cloud computing system based on artificial intelligence, which uses the activity coefficient of big data cleaning tasks to determine the first candidate cloud computing resource group, and then matches the highly adaptable candidate cloud computing resources from it, and generates a more adaptable big data cleaning strategy for the highly adaptable candidate cloud computing resources, optimizing the matching scheme of cloud computing resources and improving the big data cleaning efficiency of cloud computing resources at the same time.
[0004] In existing patent documents, static resource allocation methods are difficult to cope with this dynamic change. On the one hand, it may lead to an excess or shortage of resources in some stages of task execution, resulting in waste or bottleneck problems of resources. On the other hand, it cannot dynamically adjust resource allocation according to the real-time needs and priorities of tasks, so that high-priority tasks and dependent tasks may not be able to obtain sufficient resources in time, thus affecting the execution efficiency and success rate of the overall task. Summary of the Invention
[0005] The present application provides a big data cleaning and processing method based on artificial intelligence. By introducing a real-time adjustment factor and grouping big data tasks, the calculation of the activity coefficient is made more accurate, and the cleaning efficiency is improved at the same time.
[0006] The present application provides a big data cleaning and processing method based on artificial intelligence, including: S101. Collect big data cleaning tasks, obtain the characteristics of the big data cleaning tasks, and group the big data cleaning tasks according to the obtained characteristics of the big data cleaning tasks, into a high request volume group, a medium request volume group, and a low request volume group; S102. Collect the historical request volume data of the grouped big data cleaning tasks, and use data analysis methods to calculate the basic activity coefficient; S103. Collect the current request volume of each group of big data cleaning tasks in real time, calculate the historical average request volume according to the historical request volume data of the big data cleaning tasks, and calculate the real-time adjustment factor according to the current request volume and the historical average request volume; S104. Calculate the final activity coefficient according to the calculated basic activity coefficient and the real-time adjustment factor, obtain the allocable cloud computing resource group corresponding to the big data cleaning task according to the final activity coefficient, match the highest cloud computing resource according to the priority of the cloud computing resources in the obtained allocable cloud computing resource group, generate a resource allocation request, send the generated resource allocation request to the management server and generate a cleaning strategy, and the priority of the cloud computing resources is calculated based on the cloud computing resource portrait corresponding to the big data cleaning task.
[0007] Preferably, the formula for calculating the historical average request volume is: , where is the real-time request volume at the current time point t; W is the length of the time window; is the historical average request volume within the time window W, and N is the number of data points within the time window W.
[0008] Preferably, the formula for the real-time adjustment factor is: , where represents the real-time adjustment factor; represents the difference between the current request volume and the historical average request volume; is a normalization factor, where is the global average request volume of all grouped tasks over a long period of time, and this factor is used to adjust the difference in the basic activity levels between different grouped tasks; α is an adjustment factor used to control the influence degree of the normalization term on the real-time adjustment factor, where α > 0; A is a preset smoothing coefficient used to adjust the sensitivity of the real-time adjustment factor to the change in the request volume, where 0 < A < 1, and it is preset through expert experience.
[0009] Preferably, the final activity coefficient = basic activity coefficient real-time adjustment factor, and the allocable cloud computing resource group is a set of cloud computing resources reserved or allocated for the big data cleaning task in the cloud computing resource group.
[0010] Preferably, the calculation method for the priority of the cloud computing resources further includes: S201. Set the activity coefficient threshold according to the calculated final activity coefficient. When the final activity coefficient of the big data cleaning task exceeds the set activity coefficient threshold, arrange the priorities from high to low according to the magnitude of the excess over the set activity coefficient threshold. S202. According to the obtained priorities of the big data cleaning tasks, use the recognition algorithm to identify the feature data in the priorities, and insert the identified big data cleaning tasks containing the feature data into the front end of the big data cleaning task queue in real time, and adjust the priorities of the big data cleaning tasks in real time. S203. Allocate cloud computing resources according to the total amount of cloud computing resources and the adjusted priorities of the big data cleaning tasks.
[0011] Preferably, the method for allocating cloud computing resources further includes: S301. Construct a task dependency framework diagram according to the high-request volume group, medium-request volume group, and low-request volume group in step S101. S302. According to the constructed task dependency framework diagram, adjust the resource allocation strategy, and test the adjusted resource allocation strategy using the efficiency index. If the test fails, dynamically re-allocate the cloud computing resources.
[0012] Preferably, the dependency framework diagram includes nodes and directed edges. The nodes represent specific big data cleaning tasks. The nodes include the name and description information of the big data cleaning tasks. Draw directed edges between the nodes to represent the dependency relationship between tasks. The direction of the directed edge represents the execution order of the big data cleaning tasks, that is, the task pointed to by the arrow depends on the task at the starting point of the arrow.
[0013] Preferably, the method for dynamically re-allocating cloud computing resources: S401. Calculate the efficiency index according to the completion time, resource utilization rate, and task success rate of the big data cleaning task. S402. Compare the calculated efficiency index with the preset efficiency threshold. If the efficiency index is greater than or equal to the efficiency threshold, there is no need to adjust the resource allocation strategy. If the efficiency index is less than the efficiency threshold, then the resource allocation strategy needs to be adjusted, that is, dynamically re-allocate the cloud computing resources.
[0014] Preferably, use the weighted average method to calculate the resource utilization rate. The calculation formula is: , where U represents the resource utilization rate, which is an indicator used to measure the overall resource utilization efficiency of the system or device; represents the utilization rate of the i-th type of resource, which is calculated according to specific monitoring indicators and reflects the utilization degree of this resource in the current state; It represents the importance of the i-th resource in the overall resource utilization, i.e., the weight, which is set according to historical data, expert experience, or system requirements, and is a coefficient used to adjust the influence of different resources in the comprehensive calculation; It represents the sum of the products of the utilization rates of all resources and their corresponding weights, reflecting the overall resource utilization considering the weights; It represents the sum of the weights of all resources, which is used for normalizing the sum of products.
[0015] Preferably, resource dynamic reallocation means that during the task execution, there are situations of resource idleness or inefficient use. The situations of resource idleness or inefficient use are due to changes in task priorities, asynchronous task execution progress, or inaccurate initial resource allocation, resulting in underutilization.
[0016] One or more technical solutions provided in this application have at least the following technical effects or advantages: By introducing real-time adjustment factors and grouping big data tasks, the calculation of the activity coefficient becomes more accurate, while improving the cleaning efficiency. Eventually, the calculation of the activity coefficient is more dynamic and accurate, can reflect the changes in task request volume in real time. Through grouping and real-time adjustment, the calculation of the final activity coefficient is more in line with the actual task situation, improving the accuracy, can adjust the activity coefficient in real time according to the changes in task request volume, enhancing the flexibility of the system. A more accurate final activity coefficient helps to allocate resources more reasonably, improve the cleaning efficiency, can be applied to different types and natures of big data cleaning tasks, and has a wide range of adaptability; Real-time adjustment of the priorities of big data cleaning tasks ensures that important, active, or tasks with specific characteristics obtain sufficient computing resources, improving the efficiency and effect of big data processing. Since high-priority tasks can quickly obtain computing resources, the overall task processing time will be shortened. By dynamically adjusting the resource allocation weights, it ensures that resources are utilized more efficiently. Since the task processing speed is accelerated, the system throughput will increase accordingly. For tasks with characteristic data, they can be inserted and processed in real time, improving the system response speed; It realizes the transformation from static priority adjustment to dynamic priority adjustment, which is more flexible and adaptable; Through the characteristic data recognition mechanism, it can quickly respond to the needs of specific tasks, improving the system response speed and flexibility; By dynamically adjusting the resource allocation weights, it ensures the efficient utilization of resources; By considering the dependencies between tasks, optimizing the resource allocation strategy, and introducing parallel processing, the overall task processing efficiency is improved, bottleneck problems in resource allocation are avoided, tasks can obtain resources and be completed in a timely manner according to the order of dependencies, task delays caused by improper resource allocation are reduced. Through parallel processing, the parallel processing capabilities of computing resources are fully utilized, and the task completion time is shortened. The optimization strategy considering task dependencies can allocate resources more reasonably, reducing the task waiting time. Assuming the average task waiting time under the original strategy is W1 and the average task waiting time under the optimized strategy is W2, the waiting time is reduced by (W1 - W2). Parallel processing further enhances the processing efficiency. Assuming that it takes T time to complete a task under the original strategy and only T / N time (N is the number of tasks processed in parallel) through parallel processing under the optimized strategy, the processing efficiency is increased by N times. Considering the reduction in waiting time and the improvement in processing efficiency comprehensively, the overall task processing efficiency has been significantly improved; By dynamically adjusting the cloud computing resource allocation strategy, the execution efficiency of big data cleaning tasks is optimized. By calculating the efficiency index and comparing it with the efficiency threshold, the effectiveness of the current resource allocation strategy can be quantitatively evaluated, and it can be objectively judged whether the resource allocation strategy needs to be adjusted. Through the resource dynamic reallocation mechanism, the resource allocation can be optimized in real time, improving the resource utilization efficiency and task execution efficiency, ensuring that high-priority tasks and dependent tasks can obtain sufficient resources in a timely manner. The introduction of the resource dynamic reallocation mechanism enables resources to always be used by the tasks that need them most, further improving the resource utilization efficiency and task execution efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic flowchart of a method for processing big data cleaning based on artificial intelligence according to the present invention; Figure 2 It is a schematic flowchart of a method for calculating the priority of cloud computing resources in an embodiment of the present invention; Figure 3 It is a schematic flowchart of a method for allocating cloud computing resources in an embodiment of the present invention; Figure 4 It is a schematic flowchart of a method for dynamically reallocating cloud computing resources in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0018] To facilitate the understanding of the present invention, the present application will be described more comprehensively with reference to the relevant drawings; the preferred embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein; on the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0019] It should be noted that the terms "vertical", "horizontal", "upper", "lower", "left", "right" and similar expressions used in this article are for illustrative purposes only and do not represent the only implementation.
[0020] Unless otherwise defined, all technical and scientific terms used in this article have the same meaning as commonly understood by those skilled in the technical field to which this invention belongs; the terms used in the specification of this invention are only for the purpose of describing specific implementations and are not intended to limit this invention; the term "and / or" used in this article includes any and all combinations of one or more related listed items.
[0021] Example 1: Figure 1 It is a schematic flowchart of a big data cleaning and processing method based on artificial intelligence according to an embodiment of the present invention, including: S101, collect big data cleaning tasks, obtain the characteristics of the big data cleaning tasks, and group the big data cleaning tasks according to the obtained characteristics of the big data cleaning tasks into a high request volume group, a medium request volume group, and a low request volume group; Furthermore, use a table to organize the basic information of the big data cleaning tasks to be processed. The characteristics include data type, processing complexity, and historical performance. The data type includes structured data (tabular data in a database), semi-structured data (data in JSON or XML format), and unstructured data (such as text, images, and audio). The big data cleaning tasks are divided into a structured data group, a semi-structured data group, and an unstructured data group according to the data type; analyze the processing complexity of each big data cleaning task. The processing complexity includes the difficulty of data conversion, the number and complexity of cleaning rules. The big data cleaning tasks are divided into a simple processing group, a medium processing group, and a complex processing group according to the processing complexity of the big data cleaning tasks; collect the historical cleaning records of the big data cleaning tasks, and divide the big data cleaning tasks into a high request volume group, a medium request volume group, and a low request volume group according to the historical request volume of each big data cleaning task. For the above three grouping dimensions and criteria, this application selects to group the big data cleaning tasks according to the historical request volume of each big data cleaning task, set a unique identifier for each group, and establish a detailed file for each group, including information such as the group name, unique identifier, task list included, and grouping criteria. As new tasks are added and old tasks are completed, the grouping information is updated regularly to maintain its accuracy and timeliness.
[0022] S102, collect the historical request volume data of the grouped big data cleaning tasks, and use data analysis methods to calculate the basic activity coefficient; Specifically, clean the collected historical request volume data, remove outliers, and fill in missing data. Since outliers are caused by system failures, incorrect records, etc., use the linear regression method. Use the processed historical request volume data in the linear regression method. The historical request volume of each group of tasks is regarded as the dependent variable, and time is regarded as the independent variable. The linear regression algorithm analyzes the relationship between the dependent variable and the independent variable, fits a linear equation, and calculates the basic activity coefficient of each group of tasks according to the linear regression equation.
[0023] S103. Real-time collect the current request volume of each big data cleaning task, calculate the historical average request volume according to the historical request volume data of the big data cleaning task, and calculate the real-time adjustment factor according to the current request volume and the historical average request volume. Furthermore, use the request volume data in the recent period of time to calculate the historical average request volume, and calculate the real-time adjustment factor accordingly. The formula for calculating the historical average request volume is: , where is the real-time request volume at the current time point t; W is the length of the time window (for example, the request volume data in the past hour, day, or week); is the historical average request volume within the time window W, N is the number of data points within the time window W. Calculate the real-time adjustment factor according to the current request volume of each big data cleaning task collected in real time and the calculated historical average request volume. The formula for the real-time adjustment factor is: , where represents the real-time adjustment factor; represents the difference between the current request volume and the historical average request volume; is a normalization factor, where is the global average request volume of all grouped tasks within a long period of time (for example, the historical average request volume in the past month or year). This factor is used to adjust the difference in the basic activity levels between different grouped tasks; α is an adjustment factor used to control the influence degree of the normalization term on the real-time adjustment factor (α > 0); A is a preset smoothing coefficient used to adjust the sensitivity of the real-time adjustment factor to the change in the request volume (0 < A < 1), and it is preset through expert experience.
[0024] S104. Calculate the final activity coefficient according to the calculated basic activity coefficient and the real-time adjustment factor. Obtain the allocable cloud computing resource group corresponding to the big data cleaning task according to the final activity coefficient. When the highest cloud computing resource is matched according to the priority of the cloud computing resources in the obtained allocable cloud computing resource group, generate a resource allocation request, send the generated resource allocation request to the management server and generate a cleaning strategy. The priority of the cloud computing resources is calculated based on the cloud computing resource profile corresponding to the big data cleaning task.
[0025] Specifically, adjust the basic activity coefficient according to the calculated real-time adjustment factor to obtain the final activity coefficient. The final activity coefficient = basic activity coefficient For the real-time adjustment factor, based on the calculated final activity coefficient, determine the allocable cloud computing resource group corresponding to the big data cleaning task. The allocable cloud computing resource group is a set of cloud computing resources reserved or allocated for the big data cleaning task in the cloud computing resource group. The final activity coefficient has a positive feedback adjustment on the quantity of allocable cloud computing resources, that is, the higher the activity coefficient, the more likely the quantity of allocable resources is. Traverse the task allocation list, which contains cloud computing resource information sorted based on the priority of cloud computing resources. The priority of cloud computing resources is calculated based on the cloud computing resource profile corresponding to the big data cleaning task, reflecting factors such as resource performance, availability, and cost-effectiveness. In the allocable cloud computing resource group, match the cloud computing resource with the highest priority according to the priority of the task allocation list. When the cloud computing resource with the highest priority is successfully matched in the allocable cloud computing resource group, generate a resource allocation request. This resource allocation request contains the specific cloud computing resource information requested for allocation, as well as the relevant identifiers and requirements of the big data cleaning task. Send the generated resource allocation request to the management server. The management server is the core system responsible for cloud computing resource management and allocation, capable of processing resource requests and performing corresponding resource allocation. After receiving the resource allocation request, the management server will further obtain the big data cleaning strategy related to the big data cleaning task. The big data cleaning strategy contains key information such as the specific execution plan, required resources, and time requirements of the big data cleaning task, and is an important basis for executing the big data cleaning task.
[0026] The technical solutions in the embodiments of the present application at least have the following technical effects or advantages: By introducing the real-time adjustment factor and grouping the big data tasks, the calculation of the activity coefficient is made more accurate, and at the same time, the cleaning efficiency is improved. The calculation of the final activity coefficient is more dynamic and accurate, and can reflect the change of the task request volume in real time. Through grouping and real-time adjustment, the calculation of the final activity coefficient is more in line with the actual task situation, improving the accuracy. It can adjust the activity coefficient in real time according to the change of the task request volume, enhancing the flexibility of the system. The more accurate final activity coefficient helps to allocate resources more reasonably, improve the cleaning efficiency, and can be applied to different types and natures of big data cleaning tasks, with a wide range of adaptability.
[0027] Embodiment 2: Based on determining the priority according to the final activity coefficient in Embodiment 1, the adjustment of the priority is static. In this embodiment, the priority is dynamically adjusted through the final activity coefficient and the cut-in processing, improving the utilization rate of cloud computing resources and shortening the task processing time.
[0028] Such as Figure 2As shown, the method for calculating the priority of cloud computing resources further includes: S201. Set an active coefficient threshold according to the calculated final active coefficient. When the final active coefficient of the big data cleaning task exceeds the set active coefficient threshold, arrange the priorities from high to low according to the magnitude of the excess over the set active coefficient threshold.
[0029] Furthermore, based on the total amount of cloud computing resources in the system, including key resources such as CPU, memory, storage, and network bandwidth, traverse the historical task data to understand the average demand of the tasks, that is, the consumption of system resources by the tasks under normal circumstances. Set the active coefficient threshold according to the total amount of system resources and the average demand of the tasks. This active coefficient threshold can not only reflect the activity level of the tasks but also ensure the reasonable allocation of system resources. Compare the calculated final active coefficient of the big data cleaning task with the set active coefficient threshold. When the final active coefficient of the big data cleaning task exceeds the set active coefficient threshold, it indicates that the task is currently in a highly active state and has a high demand for system resources. Arrange the big data cleaning tasks that exceed the active coefficient threshold according to the magnitude of the excess over the active coefficient threshold. That is, the big data cleaning task with the largest excess over the set active coefficient threshold has the highest priority and is ranked at the front, and then arranged in order according to the amount of excess over the set active coefficient threshold. Since the system load and task requirements are dynamically changing, the active coefficient threshold also needs to be adjusted regularly. Set a time period, such as daily, weekly, or monthly, to re-evaluate the system load and task requirements. According to the evaluation results, adjust the active coefficient threshold in a timely manner to ensure the stability and efficiency of the system. Apply the set active coefficient threshold to the system to ensure that the priority promotion mechanism can operate normally. Conduct a comprehensive test on the system to verify whether the setting of the active coefficient threshold is reasonable. Continuously monitor the running state of the system, collect data such as task processing time and resource utilization rate, regularly analyze the monitoring data, evaluate the setting effect of the active coefficient threshold, and optimize and adjust as needed.
[0030] S202. Use an identification algorithm to identify the characteristic data in the priority according to the obtained priority of the big data cleaning task, and insert the identified big data cleaning task containing the characteristic data into the front end of the big data cleaning task queue in real time to adjust the priority of the big data cleaning task in real time; Specifically, in the big data cleaning task, feature data is data that can indicate the urgency, importance of the task, or has an abnormal data pattern. For the abnormal data pattern, it refers to a pattern different from the regular data in the big data cleaning task data. Suddenly appearing abnormally high trading volume or abnormally low user activity both represent abnormal patterns. By setting up a scanner to scan the big data cleaning tasks that enter the priority queue, when a new task is added to the queue, the scanner immediately captures it and prepares to traverse. The scanner starts to traverse the captured task data one by one. For each task, the scanner reads the values of all its fields for subsequent identification of feature data. The traversal process can be implemented through loops or iterations. For example, a for loop can be used to traverse the task list, or an iterator can be used to access tasks one by one. There is a regular database and a flag bit set in the scanner. During the traversal, the scanner compares the scanned big data cleaning task data with the built-in regular database. Once a difference is found, the flag bit in the scanner marks it. According to the abnormal degree of the data, the data with a high abnormal degree is marked as urgent, and other abnormal data follows in order. The marked abnormal data is inserted into the front end of the task queue for processing in real time, thereby adjusting the priority of the big data cleaning task in real time.
[0031] S203. Allocate cloud computing resources according to the total amount of cloud computing resources and the adjusted priority of the big data cleaning task; Furthermore, clarify the total amount of currently available cloud computing resources in the system, including but not limited to computing resources (such as CPU, memory), storage resources, and network resources. According to the adjusted priority of the big data cleaning task, allocate cloud computing resources. Set a basic resource allocation weight for the lowest priority task, set it to 1. When the system resources are tight, the lowest priority task will receive the least resource allocation to ensure the normal operation of higher priority tasks. Set a higher resource allocation weight for the highest priority task, set it to 10. When key tasks require it, the system can give priority to meeting their resource needs to ensure the smooth progress of the tasks. Set a series of intermediate weight values between the lowest and highest weights (between 1 and 10) to meet the needs of tasks with different priorities. Regularly monitor the change situation of the priority of the big data cleaning tasks in the system. When the system priority changes, adjust the resource allocation weight in a timely manner. According to the quantitative relationship of the real-time resource allocation amount, calculate the amount of computing resources that each task should obtain, and allocate the corresponding computing resources to each task.
[0032] The technical solutions in the embodiments of the present application at least have the following technical effects or advantages: ensuring that important, active, or tasks with specific characteristics obtain sufficient computing resources, improving the efficiency and effectiveness of big data processing. Since high-priority tasks can quickly obtain computing resources, the overall task processing time will be shortened. By dynamically adjusting the resource allocation weights, it is ensured that resources are utilized more efficiently. Since the task processing speed is accelerated, the system throughput will increase accordingly. For tasks with characteristic data, they can be inserted and processed in real time, improving the system's response speed; realizing the transformation from static priority adjustment to dynamic priority adjustment, being more flexible and adaptable; through the characteristic data recognition mechanism, it can quickly respond to the needs of specific tasks, improving the system's response speed and flexibility; by dynamically adjusting the resource allocation weights, the efficient utilization of resources is ensured.
[0033] Embodiment 3: Based on the resource allocation strategies in Embodiment 1 and Embodiment 2, which mainly allocate cloud computing resources according to the activity coefficient and do not consider the relationships between tasks, the present application optimizes the resource allocation strategy by constructing a task dependency graph and analyzing the input-output relationships between each group of tasks.
[0034] As Figure 3 shown, the method for allocating cloud computing resources further includes: S301, constructing a task dependency framework graph according to the high-request volume group, medium-request volume group, and low-request volume group in step S101; Specifically, the dependency relationships of tasks within a group: For the high-request-volume group, use Visio to list all tasks in the high-request-volume group one by one, assign a unique identifier to each task. For each task in the high-request-volume group, through consulting relevant experts and task executors, analyze the dependency relationships between the prerequisite tasks and subsequent tasks of each task in detail. The dependency relationship means that one task (subsequent task) must wait for another task (prerequisite task) to complete before it can start. Record the tasks with dependency relationships. The record includes all prerequisite tasks and subsequent tasks of the task. According to the recorded dependency relationships, construct the dependency chain of the high-request-volume group. Start from the task without dependencies (i.e., the task without prerequisite tasks), and gradually construct the dependency chain until all tasks are included. According to the dependency chain, determine the execution order of tasks within the high-request-volume group to ensure that prerequisite tasks are executed before subsequent tasks to meet the requirements of the dependency relationships. Use the topological sorting algorithm to automatically determine the execution order of tasks. The same method is adopted for the medium-request-volume group and the low-request-volume group to construct the dependency chains of the medium-request-volume group and the low-request-volume group respectively; The dependency relationships between tasks of different groups: Analyze whether there are associations between tasks of different groups, that is, whether the tasks of one group need to depend on the completion of the tasks of another group before they can be executed. For tasks with cross-group dependencies, sort out their dependency relationships, clarify which group's tasks are prerequisite tasks and which group's tasks are subsequent tasks, record the cross-group dependency relationships, integrate the dependency chains within the high-request-volume group, the medium-request-volume group, and the low-request-volume group. After integrating the intra-group dependency relationships, add the cross-group dependency relationships to form a task dependency framework diagram. Use the graphical tool Visio to draw the task dependency framework diagram. Use Visio to draw nodes, where the nodes represent specific big data cleaning tasks. The nodes include the names and description information of the big data cleaning tasks. Draw directed edges between the nodes to represent the dependency relationships between tasks. The direction of the directed edge represents the execution order of the big data cleaning tasks, that is, the task pointed to by the arrow depends on the task at the starting point of the arrow. The cross-group dependency relationships are also represented by directed edges to ensure that the task dependency relationships between different groups are correctly displayed.
[0035] S302. According to the constructed task dependency framework diagram, adjust the resource allocation strategy; Further, according to the constructed task dependency framework diagram, adjust the priority of the big data cleaning tasks, move the prerequisite tasks in the task dependency block diagram to a higher level so that the prerequisite tasks can obtain cloud computing resources preferentially, enabling the prerequisite tasks to be completed in a timely manner, thereby freeing up resources for subsequent tasks. For complex big data cleaning tasks, when the length of the dependency chain or the number of dependency relationships between tasks exceeds the preset thresholds respectively, it indicates that the task is a complex big data cleaning task. When the deadline of the big data cleaning task is less than the preset time threshold, the task is an urgent big data cleaning task. Increase the priorities of the complex big data cleaning tasks and the urgent big data cleaning tasks, and then optimize and adjust the cloud computing resource allocation strategy.
[0036] According to the task dependency framework diagram, identify the tasks that have no dependency relationships or whose dependency relationships have been satisfied (i.e., the prerequisite tasks have been completed), and process the identified tasks simultaneously, that is, these identified tasks are in a parallel relationship. Allocate the big data cleaning tasks that can be processed in parallel to different cloud computing resources to make full use of the parallel processing capabilities of the cloud computing resources, and monitor the execution status of the parallel processing tasks in real time. According to the monitoring results, adjust the resource allocation strategy according to the actual situation.
[0037] The technical solutions in the embodiments of the present application at least have the following technical effects or advantages: By considering the dependency relationships between tasks, optimizing the resource allocation strategy, and introducing parallel processing, the processing efficiency of the overall tasks is improved, bottleneck problems in resource allocation are avoided, tasks can obtain resources and be completed in a timely manner according to the order of dependency relationships, reducing task delays caused by improper resource allocation. Through parallel processing, the parallel processing capabilities of computing resources are fully utilized, shortening the task completion time. The optimization strategy considering task dependency relationships can allocate resources more reasonably, reducing task waiting time. Assume that the average task waiting time under the original strategy is W1, and the average task waiting time under the optimized strategy is W2, then the waiting time is reduced by (W1 - W2). Parallel processing further improves the processing efficiency. Assume that it takes T time to complete a task under the original strategy, and only T / N time (N is the number of tasks processed in parallel) is required through parallel processing under the optimized strategy, then the processing efficiency is increased by N times. Considering the reduction in waiting time and the improvement in processing efficiency comprehensively, the overall task processing efficiency is significantly improved.
[0038] Embodiment 4: Based on the above Embodiments 1 to 3, allocate cloud computing resources according to the adjusted cloud computing resource strategy, and clean the big data cleaning tasks according to the allocated resources. In this embodiment, use the efficiency index to test the efficiency of processing the big data cleaning tasks before. If not, further optimize the resource allocation strategy by dynamically reallocating resources.
[0039] Such as Figure 4Method for dynamically redistributing cloud computing resources as shown below: S401. Calculate an efficiency index based on the completion time, resource utilization rate, and task success rate of the big data cleaning task; Specifically, the efficiency index is a quantitative indicator for measuring the execution efficiency of the big data cleaning task. The efficiency index is based on the completion time, resource utilization rate, and task success rate of the big data cleaning task. The task completion time records the total time required for the task to start and complete. The resource utilization rate usually represents the degree of consumption of system resources during the execution of the task, including multiple aspects such as CPU utilization rate, memory utilization rate, and disk I / O utilization rate. The weighted average method is used to calculate the resource utilization rate, and the calculation formula is: , where U represents the resource utilization rate, which is an indicator for measuring the overall resource utilization efficiency of the system or device; represents the utilization rate of the i-th resource, which is calculated based on specific monitoring indicators and reflects the utilization degree of the resource in the current state; represents the importance of the i-th resource in the overall resource usage, that is, the weight, which is set according to historical data, expert experience, or system requirements, and is a coefficient for adjusting the influence of different resources in the comprehensive calculation; represents the sum of the products of the utilization rates of all resources and their corresponding weights, reflecting the overall resource utilization situation considering the weights; represents the sum of the weights of all resources, which is used to normalize the above product sum to ensure that the value of the comprehensive resource utilization rate U is within a reasonable range; The task success rate is based on the criteria or conditions for task success, collects the results of task execution, and evaluates the task success rate. The weighted average method is used to calculate the task success rate, and the calculation formula is: , where S represents the task success rate, which is an indicator for measuring the task completion efficiency; represents the number of tasks actually successfully completed, which is the count of tasks successfully completed during task execution; represents the total number of tasks, which is the count of all tasks during task execution. Calculate the efficiency index based on the obtained task completion time, resource utilization rate, and task success rate. The calculation formula is: , where represents the reciprocal of the task completion time. The shorter the task completion time, the larger its reciprocal, and the greater the contribution to the efficiency index; represents the reciprocal of the resource utilization rate. The lower the resource utilization rate (i.e., the more efficient the resource utilization), the larger its reciprocal, and the greater the contribution to the efficiency index; S represents the task success rate. The higher the task success rate, the greater the contribution to the efficiency index; respectively represent the weights of the task completion time, resource utilization rate, and task success rate, which are set according to historical data, expert experience, or system requirements.
[0040]
[0040] In S402, compare the calculated efficiency index with a preset efficiency threshold. If the efficiency index is greater than or equal to the efficiency threshold, there is no need to adjust the resource allocation strategy. If the efficiency index is less than the efficiency threshold, then the resource allocation strategy needs to be adjusted, that is, perform dynamic reallocation of cloud computing resources; Furthermore, collect and understand the experience and opinions of experts in related fields. Experts usually have in-depth understanding of the system or process, and their experience can provide valuable reference for setting the efficiency threshold. At the same time, clarify the specific requirements of the system, including performance requirements, resource limitations, operation objectives, etc. Based on expert experience and system requirements, set an efficiency threshold. This efficiency threshold is a critical value that can reflect the optimization degree of the resource allocation strategy. Compare the calculated efficiency index with the set efficiency threshold. If the efficiency index is greater than or equal to the efficiency threshold, it indicates that the current resource allocation strategy meets the requirements and can continue to be maintained or further optimized. If the efficiency index is less than the efficiency threshold, it indicates that there are problems with the current resource allocation strategy, and perform dynamic reallocation of cloud computing resources.
[0041]
[0041] Dynamic resource reallocation means that during the task execution process, there will always be some resources that are idle or inefficiently used. These resources may not be fully utilized due to reasons such as changes in task priorities, asynchronous task execution progress, or inaccurate initial resource allocation. To improve the overall resource utilization efficiency, the system needs to reallocate these idle or inefficiently used resources to those tasks that currently need more resources. The process of resource reallocation includes identifying idle or inefficiently used resources, determining tasks that need more resources, and implementing the reallocation of resources. During this process, the system needs to monitor the resource usage situation in real time, identify which resources are in an idle or inefficient state through data analysis, and at the same time evaluate the resource requirements of each task to ensure that resources can be accurately allocated to the tasks that need them most. Dynamic resource reallocation is a continuous process that requires continuous monitoring, comparison, adjustment, and optimization. The system needs to monitor the changes in the efficiency index in real time, compare it with the efficiency threshold in a timely manner, and adjust the resource allocation strategy according to the comparison results.
[0042] The technical solutions in the embodiments of the present application at least have the following technical effects or advantages: By dynamically adjusting the cloud computing resource allocation strategy, the execution efficiency of the big data cleaning task is optimized. By calculating the efficiency index and comparing it with the efficiency threshold, the effectiveness of the current resource allocation strategy can be quantitatively evaluated, and it can be objectively judged whether the resource allocation strategy needs to be adjusted. Through the resource dynamic reallocation mechanism, the resource allocation can be optimized in real time, improving the resource utilization efficiency and the task execution efficiency, ensuring that high-priority tasks and dependent tasks can obtain sufficient resources in a timely manner. The introduction of the resource dynamic reallocation mechanism enables the resources to be always used by the tasks that most need them, further improving the resource utilization efficiency and the task execution efficiency.
[0043] The above is only the preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A big data cleaning and processing method based on artificial intelligence, characterized in that: include: S101, collecting big data cleaning tasks, obtaining characteristics of the big data cleaning tasks, and grouping the big data cleaning tasks according to the obtained characteristics of the big data cleaning tasks into a high request volume group, a medium request volume group, and a low request volume group; S102, collecting historical request volume data of the grouped big data cleaning tasks, and using a data analysis method to calculate a basic activity coefficient; S103, collecting the current request volume of each group of big data cleaning tasks in real time, calculating the historical average request volume according to the historical request volume data of the big data cleaning tasks, and calculating the real-time adjustment factor according to the current request volume and the historical average request volume; S104, calculate the final activity coefficient according to the calculated basic activity coefficient and the real-time adjustment factor, obtain the allocatable cloud computing resource group corresponding to the big data cleaning task according to the final activity coefficient, match the highest cloud computing resource in the obtained allocatable cloud computing resource group according to the priority of the cloud computing resources, generate a resource allocation request, send the generated resource allocation request to the management server and generate a cleaning strategy, the priority of the cloud computing resource is calculated based on the cloud computing resource portrait corresponding to the big data cleaning task.
2. The big data cleaning processing method based on artificial intelligence according to claim 1, characterized in that: The formula for calculating the historical average request volume is: ,in, is the real-time request volume at the current time point t; W is the length of the time window; is the historical average request volume in time window W, and N is the number of data points in time window W.
3. The big data cleaning and processing method based on artificial intelligence as claimed in claim 1, characterized in that: The real-time adjustment factor formula is: ,in, represents the real-time adjustment factor; Indicates the difference between the current request volume and the historical average request volume; is a normalization factor, where is the global average request volume of all grouped tasks over a long period of time, which is used to adjust the basic activity level differences between different grouped tasks; α is an adjustment factor used to control the influence of the normalization term on the real-time adjustment factor; A is a preset smoothing coefficient, which is used to adjust the sensitivity of the real-time adjustment factor to the change in request volume and is preset through expert experience.
4. The big data cleaning processing method based on artificial intelligence as claimed in claim 1, characterized in that: Final activity coefficient = basic activity coefficient The real-time adjustment factor,the allocatable cloud computing resource group is a set of cloud computing resources,reserved or allocated for the big data cleaning task in the cloud computing resource,group.
5. The big data cleaning processing method based on artificial intelligence as claimed in claim 1, characterized in that: The calculation method of cloud computing resource priority also includes: S201, according to the calculated final activity coefficient, an activity coefficient threshold is set, and when the final activity coefficient of the big data cleaning task exceeds the set activity coefficient threshold, the priorities are arranged from high to low according to the magnitude of the excess of the set activity coefficient threshold; S202, according to the obtained priority of the big data cleaning task, use the recognition algorithm to identify the characteristic data in the priority, insert the identified big data cleaning task containing the characteristic data into the front end of the big data cleaning task queue in real time, and adjust the priority of the big data cleaning task in real time; S203: Allocate cloud computing resources according to the total amount of cloud computing resources and the adjusted priority of the big data cleaning task.
6. The big data cleaning processing method based on artificial intelligence as claimed in claim 5, characterized in that: The method of allocating cloud computing resources also includes: S301, constructing a task dependency framework diagram according to the high request volume group, the medium request volume group and the low request volume group in step S101; S302, adjusting the resource allocation strategy according to the constructed task dependency framework diagram, and testing the adjusted resource allocation strategy using the efficiency index, and dynamically reallocating the cloud computing resources if the strategy fails the test.
7. The big data cleaning processing method based on artificial intelligence as claimed in claim 6, characterized in that: The dependency framework diagram includes nodes and directed edges. The nodes represent specific big data cleaning tasks. The nodes include the name and description information of the big data cleaning tasks. Directed edges are drawn between the nodes to indicate the dependency relationship between the tasks. The direction of the directed edges indicates the execution order of the big data cleaning tasks, that is, the task pointed to by the arrow depends on the task at the starting point of the arrow.
8. The big data cleaning method based on artificial intelligence as claimed in claim 6, characterized in that: Methods for dynamically reallocating cloud computing resources: S401, calculating an efficiency index according to the completion time, resource utilization rate, and task success rate of the big data cleaning task; S402, compare the calculated efficiency index with the preset efficiency threshold. If the efficiency index is greater than or equal to the efficiency threshold, there is no need to adjust the resource allocation strategy. If the efficiency index is less than the efficiency threshold, the resource allocation strategy needs to be adjusted, that is, the cloud computing resources are dynamically reallocated.
9. The big data cleaning method based on artificial intelligence as claimed in claim 8, characterized in that: The resource utilization rate is calculated using the weighted average method. The calculation formula is: , where U represents resource utilization, which is an indicator used to measure the overall resource utilization efficiency of a system or device; It represents the utilization rate of the i-th resource, which is calculated based on the specific monitoring indicators and reflects the utilization degree of the resource in the current state; It represents the importance of the i-th resource in the overall resource usage, i.e., the weight, which is set according to historical data, expert experience, or system requirements. It is a coefficient used to adjust the influence of different resources in the comprehensive calculation; It represents the sum of the products of the utilization rates of all resources and their corresponding weights, reflecting the overall resource utilization after considering the weights; Represents the sum of the weights of all resources and is used to normalize the product sum.
10. The big data cleaning method based on artificial intelligence as claimed in claim 8, characterized in that: Dynamic resource reallocation refers to the situation where resources are idle or inefficiently used during task execution. Idle or inefficient resource use is caused by changes in task priorities, asynchronous task execution progress, or inaccurate initial resource allocation, resulting in failure to fully utilize resources.
Citation Information
Patent Citations
Big data cleaning task processing method based on artificial intelligence and cloud computing system
CN114844901A
Cited By
Air traffic control monitoring method and system based on multi-modal data processing
CN120783595A