Data processing method and device for big data platform

By dynamically adjusting task priority and resource allocation and optimizing time segments and storage paths, the problem of inflexible resource allocation in data processing on big data platforms is solved, and processing efficiency and system stability are improved.

CN120216141AInactive Publication Date: 2025-06-27XIAMEN MAIGE INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510343683.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-22
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology lacks a strategy to dynamically adjust the resource occupation and task timeliness in the data processing process of big data platforms, resulting in inflexible resource allocation and insufficient adaptability, which affects processing efficiency and data access efficiency.

Method used

By extracting the time parameters of the task, resource occupancy parameters and processing complexity parameters, calculating weight values, dynamically adjusting the weight matching task status, building a task priority matrix, and decomposing the task into time segments, correlating the time segments and resource distribution, optimizing the distribution path and storage load balancing, and adjusting the resource allocation of the task link.

Benefits of technology

It improves the accuracy of task scheduling and the efficiency of resource matching, reduces execution delay, improves task distribution and storage management strategies among nodes, significantly improves the consistency and efficiency of multi-stage task processing, and enhances the real-time and stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216141A_ABST
    Figure CN120216141A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data management, in particular to a data processing method and device for a big data platform, and the method comprises the following steps: extracting a time parameter, a resource occupation parameter and a processing complexity parameter of a task, calculating a weighted value of time and resources, integrating the complexity parameter, and dynamically adjusting a weight ratio to match a task state. And identifying a priority numerical value, and constructing a task priority matrix. According to the invention, through dynamic adjustment of task parameters, the new scheme improves the accuracy of task scheduling and the efficiency of resource matching, the execution delay is reduced through optimized time slice division and resource allocation, the task distribution among nodes is improved, the storage management strategy improves the data access efficiency through adjustment of paths and distribution of loads, and in addition, the data access efficiency is improved. The dynamic optimization strategy of the task chain significantly improves the continuity and efficiency of multi-stage task processing, and enhances the real-time performance and stability of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data management, and in particular, to a data processing method and device for a big data platform. Background Art

[0002] The technical field of data management includes technical methods and system designs related to data collection, storage, management, processing, and analysis. The core content of this field covers the design of efficient data storage structures, the integrated management of multi-source data, the indexing and retrieval mechanisms for complex data, data transmission and sharing technologies, and the guarantee of consistency and integrity during the data processing process. By introducing distributed databases, cloud storage architectures, and data lake technologies, the technical field of data management supports efficient operations on large-scale structured, unstructured, and semi-structured data. Especially in the big data environment, it focuses on the real-time performance, scalability, and fault tolerance of data processing. This field also involves data security and privacy protection technologies to ensure the integrity and reliability of data during its life cycle management.

[0003] Among them, the data processing method for a big data platform refers to a technical solution for the efficient processing and standardized management of multi-source heterogeneous data involved in a big data platform. Specifically targeting the complexity problems existing in the processes of big data collection, preprocessing, cleaning, storage, and analysis, this patent theme proposes to manage data in regions through a distributed storage architecture and to use rule-based data cleaning technologies to filter and adjust noise data and abnormal data. The data preprocessing link includes automatic identification and conversion operations of data types, and the parsing and reorganization of multi-dimensional data through a logical optimization model to ensure that the degree of data structuring meets the subsequent analysis requirements. The data analysis part combines a specific indexing mechanism and a partition query strategy to quickly process and classify large-scale data through an adaptive calculation method, thereby realizing the synchronous management and efficient processing of multiple data streams within the platform.

[0004] The prior art mainly relies on static strategies in task status response, lacking the ability to dynamically adjust between resource occupation and task timeliness, resulting in inflexible resource allocation. In task decomposition and time segment division, the failure to construct dynamic associations based on the multi-dimensional attributes of tasks leads to insufficient adaptability between resource distribution and task time segments. In terms of node task distribution, the node resource carrying capacity and collaboration intensity are not effectively utilized, resulting in limited optimization of the distribution path and increased communication costs. In fragmented storage, there is a lack of a path adjustment strategy based on call frequency and load distribution, with poor storage balance, affecting data access efficiency. In the multi-stage collaboration of task chains, there is a lack of dynamic adjustment of resource allocation and link optimization strategies, with significant link bottleneck problems, affecting the overall task execution efficiency. These problems lead to obvious limitations in the adaptability and efficiency of the prior art in a dynamic data processing environment. Summary of the Invention

[0005] The object of the present invention is to solve the disadvantages existing in the prior art, and to propose a data processing method and device for a big data platform.

[0006] To achieve the above object, the present invention adopts the following technical solutions: A data processing method for a big data platform, comprising the following steps:

[0007] S1: Extract the time parameter, resource occupancy parameter, and processing complexity parameter of the task, calculate the weighted value of time and resources, integrate the complexity parameter, dynamically adjust the weight ratio to match the task status, identify the priority value, and construct a task priority matrix;

[0008] S2: Based on the priority matrix, extract the priority value and time length, decompose the task into time segments, establish a mapping between the time segments and the priority values, combine the resource requirements and status parameters, associate the time segments with the resource distribution, integrate and analyze the task time distribution data, and generate a time segment distribution table;

[0009] S3: Based on the time segment distribution table, extract the time segments and resource requirement data, analyze the node resource carrying capacity, match the cooperation intensity, integrate the task path and communication cost parameters, optimize the distribution path, allocate the task to the appropriate node, record the mapping between the task and the node, and generate a node task allocation diagram;

[0010] S4: Based on the node task allocation diagram, analyze the node task distribution data and the storage fragment call record, analyze the fragment call frequency and path length, adjust the storage path distribution, reorganize the correspondence between the nodes and the fragments, optimize the storage load balance, integrate the fragment position relationship, and generate an optimized structure for data fragment storage;

[0011] S5: Based on the distribution information of the optimized structure for data fragment storage, analyze the resource requirements and time slice distribution of the task chain, integrate the resource allocation ratio and link configuration parameters, adjust the multi-stage task cooperation logic, optimize the resource allocation of the task link, and generate a distributed task chain cooperation table.

[0012] As a further solution of the present invention, the task priority matrix includes a priority weight distribution, a dynamic weight adjustment result, and a comprehensive task score; the time segment distribution table includes a time segment set, a mapping relationship between the priority value and the time segment, and a resource requirement status table; the node task allocation diagram includes a node cooperation relationship table, a task allocation path diagram, and communication cost distribution data; the optimized structure for data fragment storage includes an optimized storage path distribution scheme, a mapping relationship between the fragments and the nodes, and a call frequency statistical table; the distributed task chain cooperation table includes task chain stage division, resource allocation parameters, and a link cooperation relationship table.

[0013] As a further solution of the present invention, the specific steps for extracting the time parameter, resource occupancy parameter, and processing complexity parameter of the task, calculating the weighted value of time and resources, integrating the complexity parameter, dynamically adjusting the weight ratio to match the task status, identifying the priority value, and constructing the task priority matrix are as follows:

[0014] S101: Based on the time parameter, resource occupancy parameter, and processing complexity parameter of the task, calculate the weighted value of time and resources for each task according to the weight distribution of each parameter in the task status, dynamically update the weight ratio for changes in the task status, record the relationship between task time and resource consumption, and generate a task time and resource weighted value table;

[0015] S102: Based on the task time and resource weighted value table, perform a normalization process on the complexity parameter of the task, combine the time and resource weighted value, establish the state impact factor for each task, summarize the matching relationship between the weighted value and the complexity parameter, optimize the difference in the impact value between task states, and generate a task state impact factor list;

[0016] S103: Based on the task state impact factor list, perform a one-by-one comparison and analysis on the task priority values, adjust the relationship of the priority parameters according to the weight gradient, arrange the task priorities in descending order, establish and analyze the mapping rule between the task state and the priority, and generate a task priority matrix.

[0017] As a further solution of the present invention, based on the priority matrix, extract the priority value and time length, decompose the task into time segments, establish the mapping between the time segment and the priority value, combine the resource requirements and status parameters, associate the time segment with the resource distribution, integrate and analyze the task time distribution data, and the specific steps for generating a time segment distribution table are as follows:

[0018] S201: Based on the priority matrix, extract the priority value and time length parameter of each task, classify the priority values, decompose and disassemble multiple time segments according to the task time, calculate the priority value corresponding to the time segment, analyze the distribution characteristics, associate the task priority with the time segment relationship, and generate a time segment priority mapping table;

[0019] S202: Based on the time segment priority mapping table, combine the resource requirement parameter of the task, extract the resource consumption of each time segment, analyze the distribution pattern of the time segment and the resource usage, calculate the resource matching value of the time segment according to the task resource usage rule, and obtain the time segment resource distribution data;

[0020] S203: Based on the time segment resource distribution data, summarize the resource distribution characteristics of time segments and tasks, match the coordination between task time distribution and resources, adjust the resource ratio rules of time segments, construct a classification structure of time distribution parameters, and generate a time segment distribution table.

[0021] As a further solution of the present invention, the specific calculation formula for the priority value corresponding to the time segment is:

[0022]

[0023] Among them, P s1 represents the priority value corresponding to the time segment, P k1 represents the priority value of the kth task, T k1 represents the time length parameter of the kth task, q1 represents the total number of tasks within the time segment, represents the sum of the time lengths of all tasks within the time segment.

[0024] As a further solution of the present invention, based on the time segment distribution table, extract the time segment and resource demand data, analyze the node resource carrying capacity, match the cooperation intensity, integrate the task path and communication cost parameters, optimize the distribution path, allocate tasks to the appropriate nodes, record the mapping between tasks and nodes, and the specific steps for generating the node task allocation diagram are as follows:

[0025] S301: Based on the time segment distribution table, extract the resource demand data of each time segment, classify and summarize the resource demands, calculate the carrying capacity value of the node resources, match the corresponding relationship between the resource demands and the node carrying values, establish the adaptability score of the time segment on the node, and generate a node resource adaptability table;

[0026] S302: Based on the node resource adaptability table, combine the time segment cooperation intensity parameter and the node communication cost parameter, analyze the task path distribution, adjust the matching of the path connection and the communication cost between nodes, optimize the distribution mode of the task path resources, construct and analyze the relationship between the task distribution path and the communication cost, and generate a task distribution path diagram;

[0027] S303: Based on the task distribution path diagram, extract the node carrying capacity and task path resource data, and allocate tasks to the appropriate nodes according to the time segment distribution, record the distribution relationship of the time segment on the node, and generate a node task allocation diagram.

[0028] As a further solution of the present invention, the specific formula for calculating the carrying capacity value is:

[0029]

[0030] Among them, C r1Represents the node resource carrying capacity value, R i3 Represents the resource demand in the i-th time segment, T i4 Represents the time length parameter of the i-th time segment, and q2 represents the total number of time segments. Represents the sum of the time lengths of all time segments.

[0031] As a further solution of the present invention, based on the node task allocation graph, analyze the node task distribution data and storage fragment call records, analyze the fragment call frequency and path length, adjust the storage path distribution, reorganize the correspondence between nodes and fragments, optimize the storage load balance, integrate the fragment position relationship, and the specific steps for generating the optimized structure of data fragment storage are as follows:

[0032] S401: Based on the node task allocation graph, extract the node task distribution data and storage fragment call records, classify and sort out the fragment call frequency of each node, analyze the change in the length of the call path, summarize the correlation between the call frequency and the path length, and generate a fragment call and path analysis table;

[0033] S402: Based on the fragment call and path analysis table, compare the distribution characteristics of the fragment call frequency and the path length, adjust the distribution of the node storage path, redefine the correspondence between fragments and nodes, identify the adjustment result of the storage path, and generate a storage path adjustment distribution map;

[0034] S403: Based on the storage path adjustment distribution map, analyze the distribution state of the node storage load, rematch the position relationship between nodes and fragments, summarize the balanced distribution characteristics between fragment storages, and integrate the fragment position relationship between nodes to generate an optimized structure of data fragment storage.

[0035] As a further solution of the present invention, based on the distribution information of the optimized structure of data fragment storage, analyze the task chain resource requirements and time slice distribution, integrate the resource allocation ratio and link configuration parameters, adjust the multi-stage task cooperation logic, optimize the task link resource allocation, and the specific steps for generating a distributed task chain cooperation table are as follows:

[0036] S501: Based on the distribution information of the optimized structure of data fragment storage, extract the task chain resource requirement parameters and time slice distribution characteristics, classify and summarize the resource occupation and time distribution patterns of the task chain, perform data integration based on the correlation of resource occupation parameters, analyze the phased resource requirement characteristics of the task chain, and generate a task chain resource and time distribution table;

[0037] S502: Based on the task chain resource and time distribution table, conduct data matching analysis on the resource allocation ratio and link configuration parameters, summarize the adaptability of resource allocation among links, adjust the link resource occupancy distribution rules, optimize the coordination relationship between links and resources, and generate a link resource allocation and configuration diagram;

[0038] S503: Based on the link resource allocation and configuration diagram, classify and integrate the collaboration logic of multi-stage tasks, analyze the resource distribution characteristics among task links, adjust the matching of collaboration intensity and resource distribution, summarize the distributed resource scheduling relationship of the task chain, and generate a distributed task chain collaboration table.

[0039] A data processing system for a big data platform, comprising:

[0040] The task priority module extracts the time parameter, resource occupancy parameter, and processing complexity parameter of the task, calculates the weighted value of time and resources, integrates the complexity parameter, dynamically adjusts the weight ratio to match the task state, and constructs a task priority matrix;

[0041] The time segment module extracts the priority value and time length based on the priority matrix, decomposes the task into time segments, establishes the mapping between the time segments and the priority values, combines the resource requirements and status parameters, associates the time segments with the resource distribution, and generates a time segment distribution table;

[0042] The node allocation module extracts the time segment and resource requirement data based on the time segment distribution table, analyzes the node resource carrying capacity, matches the collaboration intensity, integrates the task path and communication cost parameters, allocates the task to the appropriate node, and generates a node task allocation diagram;

[0043] The storage optimization module analyzes the node task distribution data and storage fragment call records based on the node task allocation diagram, analyzes the fragment call frequency and path length, adjusts the storage path distribution, reorganizes the correspondence between nodes and fragments, optimizes the storage load balance, and generates a data fragment storage optimization structure;

[0044] The link collaboration module analyzes the task chain resource requirements and time slice distribution based on the distribution information of the data fragment storage optimization structure, integrates the resource allocation ratio and link configuration parameters, adjusts the multi-stage task collaboration logic, and generates a distributed task chain collaboration table.

[0045] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0046] In the present invention, through the dynamic adjustment of task parameters, the new solution improves the accuracy of task scheduling and the efficiency of resource matching. The optimized time slice division and resource allocation reduce the execution delay and improve the task distribution among nodes. The storage management strategy improves the data access efficiency by adjusting the path and distributing the load. In addition, the dynamic optimization strategy of the task chain significantly enhances the coherence and efficiency of multi-stage task processing, and strengthens the real-time performance and stability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0048] Figure 1 It is a schematic flowchart of the steps of the present invention;

[0049] Figure 2 It is a flowchart of step S1 of the present invention;

[0050] Figure 3 It is a flowchart of step S2 of the present invention;

[0051] Figure 4 It is a flowchart of step S3 of the present invention;

[0052] Figure 5 It is a flowchart of step S4 of the present invention;

[0053] Figure 6 It is a flowchart of step S5 of the present invention;

[0054] Figure 7 It is a system module diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] The following will describe the technical solutions in the present invention with reference to the drawings.

[0056] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly, the use of the word "example" is intended to present concepts in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.

[0057] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same. "Of", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same.

[0058] In the embodiments of the present invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meanings they express are the same.

[0059] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0060] Please refer to Figure 1 , a data processing method for a big data platform, including the following steps:

[0061] S1: Extract the time parameter, resource occupancy parameter, and processing complexity parameter of the task, calculate the weighted value of time and resources, integrate the complexity parameter, dynamically adjust the weight ratio to match the task status, identify the priority value, and construct a task priority matrix;

[0062] S2: Based on the priority matrix, extract the priority value and time length, decompose the task into time segments, establish a mapping between the time segments and the priority values, combine the resource requirements and status parameters, associate the time segments with the resource distribution, integrate and analyze the task time distribution data, and generate a time segment distribution table;

[0063] S3: Based on the time segment distribution table, extract the time segments and resource requirement data, analyze the node resource carrying capacity, match the cooperation intensity, integrate the task path and communication cost parameters, optimize the distribution path, allocate the task to the appropriate node, record the mapping between the task and the node, and generate a node task allocation diagram;

[0064] S4: Based on the node task allocation diagram, analyze the node task distribution data and the storage fragment call record, analyze the fragment call frequency and path length, adjust the storage path distribution, reorganize the corresponding relationship between the node and the fragment, optimize the storage load balance, integrate the fragment position relationship, and generate an optimized structure for data fragment storage;

[0065] S5: Based on the distribution information of the optimized structure for data fragment storage, analyze the resource requirements and time slice distribution of the task chain, integrate the resource allocation ratio and link configuration parameters, adjust the multi-stage task cooperation logic, optimize the resource allocation of the task link, and generate a distributed task chain cooperation table.

[0066] The task priority matrix includes the priority weight distribution, the dynamic weight adjustment result, and the comprehensive task score. The time segment distribution table includes the time segment set, the mapping relationship between the priority value and the time segment, and the resource requirement status table. The node task allocation graph includes the node collaboration relationship table, the task allocation path graph, and the communication cost distribution data. The data fragment storage optimization structure includes the storage path distribution optimization scheme, the mapping relationship between the fragments and the nodes, and the call frequency statistics table. The distributed task chain collaboration table includes the task chain stage division, the resource allocation parameters, and the link collaboration relationship table.

[0067] Please refer to Figure 2 , the specific steps of S1 are as follows:

[0068] S101: Based on the time parameter, resource occupancy parameter, and processing complexity parameter of the task, calculate the weighted time and resource value of each task according to the weight distribution of each parameter in the task state, dynamically update the weight ratio to compare the task state changes, record the relationship between task time and resource consumption, and generate a task time and resource weighted value table;

[0069] Calculate the weighted time and resource value of the task according to the formula

[0070]

[0071] Calculate the weighted value of each task.

[0072] In the formula, W t represents the weighted time and resource value of the task, T i represents the time parameter of the task, R i represents the resource occupancy parameter of the task, C i represents the processing complexity parameter of the task, n represents the total number of tasks, is the sum of all task time parameters.

[0073] The task time parameter is directly measured through the start and end times recorded in the task log. The resource occupancy parameter can be collected through the system resource monitoring tool during the task execution (such as CPU occupancy rate, memory occupancy, etc.). The processing complexity parameter is quantified by the complexity index of the task operation steps. The index standard comes from the operation manual and the analysis after the task module is disassembled. The calculation formula is the square root of the number of operation steps.

[0074] For example, for three tasks, the time parameters are 5, 10, and 15 (unit: minutes) respectively, the resource occupancy parameters are 30, 40, and 50 (unit: resource units), and the processing complexity parameters are 2, 3, and 4. Substitute into the formula:

[0075]

[0076] W t= 10 + 40 + 100 = 150;

[0077] The calculated task weighted value is 150. The results show the weight distribution of the time and resource consumption of each task in the overall task. Through this value, the priority or scheduling strategy between tasks can be further analyzed.

[0078] S102: Based on the task time and resource weighted value table, standardize the complexity parameters of the tasks, combine the time and resource weighted values, establish the state impact factor of each task, summarize the matching relationship between the weighted value and the complexity parameters, optimize the difference in the impact values between task states, and generate a list of task state impact factors;

[0079] Normalize the task complexity parameters according to the ratio of the resource occupancy parameters of each task. The resource occupancy parameter data is obtained by parsing the resource allocation record file in the operation log. Analyze the maximum resource occupancy and average occupancy of each task in the file record, recalculate the values of the complexity parameters of each task through the normalization formula, and establish a mapping relationship with the task time and resource weighted values. During the mapping process, call the historical record analysis function in the database to extract the resource consumption data of similar tasks, and conduct a matching comparison analysis with the time weighted data to form an association impact factor matrix between tasks, and finally generate a list of task state impact factors.

[0080] S103: Based on the list of task state impact factors, conduct a comparison and analysis of the task priority values item by item, adjust the relationship of the priority parameters according to the weight gradient, arrange the task priorities in descending order, establish and analyze the mapping rules between the task state and the priority, and generate a task priority matrix;

[0081] The comparison and analysis of the priority values are carried out by dynamically adjusting the parameter relationship between the impact factor and the state weight. The task priority values in the historical task priority allocation table are used as the initial comparison benchmark for parameter calls, and the priority gradient parameters are adjusted through the task completion time and resource release time statistically recorded in the real-time task execution status. During the gradient adjustment process, the real-time monitored task priorities are reordered, the priority matrix is arranged in descending order, and at the same time, it is combined with the comparison analysis function in the priority historical adjustment record to generate the historical mapping relationship of the priority data, forming a task priority matrix.

[0082] Please refer to Figure 3 , the specific steps of S2 are as follows:

[0083] S201: Based on the priority matrix, extract the priority value and time length parameter of each task, classify the priority values, split and disassemble multiple time segments according to the task time, calculate the priority values corresponding to the time segments, analyze the distribution characteristics, associate the task priority with the time segment relationship, and generate a time segment priority mapping table;

[0084] The calculation formula for the priority value corresponding to the time segment is specifically as follows:

[0085]

[0086] Among them, P s1 represents the priority value corresponding to the time segment, and P k1 represents the priority value of the k-th task, T k1 represents the time length parameter of the k-th task, q1 represents the total number of tasks within the time segment, represents the sum of the time lengths of all tasks within the time segment.

[0087] It is assumed that the task chain contains three time segments, and the task priority values are P 11 = 3, P 21 = 5, P 31 = 2, and the time length parameters are T 11 = 10, T 21 = 20, T 31 = 15. The number of time segments is q1 = 3.

[0088] Calculate the total time length:

[0089]

[0090] Calculate the square root factor of the priority value and time distribution item by item:

[0091]

[0092] Calculate the priority eigenvalue component:

[0093] P 11 ·0.4714 = 3·0.4714 = 1.4142;

[0094] P 21 ·0.6667 = 5·0.6667 = 3.3335;

[0095] P 31 ·0.5774 = 2·0.5774 = 1.1548;

[0096] Calculate the absolute value and average value:

[0097]

[0098] The result of the priority value corresponding to the time segment is 1.9675.

[0099] The result shows the matching characteristics between the task priority distribution within a time segment and the time parameters. The numerical result reflects the priority intensity of the task distribution, providing a key reference basis for generating the time segment priority mapping table.

[0100] S202: Based on the time segment priority mapping table, combined with the resource requirement parameters of the task, extract the resource consumption of each time segment, analyze the distribution pattern of the time segment and resource usage, calculate the resource matching value of the time segment according to the task resource usage rule, and obtain the time segment resource distribution data;

[0101] The resource requirement parameters are extracted through the resource record file of the task. Among them, the resource consumption is accumulated by time segment. The specific operation is to read the resource values at each time point in the task resource usage record, and summarize them by time segment in combination with the start and end times of the time segment. When analyzing the resource distribution pattern, the resource usage mean and variance calculation methods are used to quantify the resource distribution state of the time segment. The resource matching value is obtained through the comparison and analysis of the resource requirements of each time segment with the priority mapping. In the comparison and analysis, it is necessary to call the total resource consumption in the real-time monitoring data and refine the processing in combination with the resource utilization rate in the task historical record, and finally obtain the time segment resource distribution data.

[0102] S203: Based on the time segment resource distribution data, summarize the resource distribution characteristics of the time segment and the task, match the coordination between the task time distribution and resources, adjust the resource ratio rule of the time segment, construct the classification structure of the time distribution parameters, and generate the time segment distribution table;

[0103] The resource distribution characteristics are summarized through the correlation analysis of the time segment priority data and the resource utilization rate of the task. In the summarization process, the time distribution characteristics are quantified as the average resource distribution amount in different priority time periods. Through the rule adjustment of the resource allocation strategy within the time segment, the dynamic ratio coefficient of resource utilization is divided into three categories: high load, medium load, and low load. Calculate the resource allocation requirements under each load, and combine the balance data of resource allocation in different time periods in the task historical record to construct the classification structure of the time distribution parameters and form a distribution table to match the time and resource demand characteristics of the task.

[0104] Please refer to Figure 4 , the specific steps of S3 are as follows:

[0105] S301: Based on the time segment distribution table, extract the resource requirement data of each time segment, classify and summarize the resource requirements, calculate the bearing capacity value of the node resources, match the corresponding relationship between the resource requirements and the node bearing value, establish the adaptability score of the time segment on the node, and generate the node resource adaptability table;

[0106] The specific formula for calculating the bearing capacity value is:

[0107]

[0108] Among them, C r1 represents the node resource carrying capacity value, R i3 represents the resource demand in the i-th time segment, T i4 represents the time length parameter of the i-th time segment, q2 represents the total number of time segments, represents the sum of the time lengths of all time segments.

[0109] Suppose the set data contains three time segments, and their resource demands are R 13 = 50, R 23 = 70, R 33 = 90 (unit: resource unit), and the time lengths are T 14 = 10, T 24 = 20, T 34 = 15 (unit: minute), and the total number of time segments is q2 = 3.

[0110] Calculate the sum of the time lengths:

[0111]

[0112] Calculate the square root of the time length ratio item by item:

[0113]

[0114] Calculate the adjusted value of the resource demand item by item:

[0115] R 13 ·0.4714 = 50·0.4714 = 23.57;

[0116] R 23 ·0.6667 = 70·0.6667 = 46.67;

[0117] R 33 ·0.5774 = 90·0.5774 = 51.97;

[0118] Calculate the absolute value and the node resource carrying capacity value:

[0119]

[0120] The result of the node resource carrying capacity value is 40.74 (unit: resource unit).

[0121] This result indicates the average resource carrying capacity of the node under the current time segment distribution situation, and the numerical result provides a necessary calculation basis for subsequent matching of resource demands with node carrying values and establishing time segment adaptability scores.

[0122] S302: Based on the node resource adaptation table, combined with the time segment collaboration intensity parameter and the node communication cost parameter, analyze the task path distribution, adjust the matching of path connections and the communication cost between nodes, optimize the distribution mode of task path resources, construct and analyze the relationship between the task distribution path and the communication cost, and generate a task distribution path diagram;

[0123] The time segment collaboration intensity parameter is quantified according to the adjacency degree of each time segment in the task distribution path. The node communication cost parameter is calculated through the transmission delay records in the communication logs of the distributed system. The distribution analysis of the task path is carried out based on the comparison of the collaboration intensity and the communication cost parameter of the time segment. The path connection adjustment is completed by optimizing the collaboration intensity threshold between time segments. The matching of the communication cost between paths is referenced by the resource usage mode in the time segment distribution table. By adjusting the allocation ratio of path connections, the distribution mode of task path resources tends to be balanced, and finally a task distribution path diagram is generated.

[0124] S303: Based on the task distribution path diagram, extract the node bearing capacity and task path resource data. According to the time segment distribution, allocate the tasks to the adapted nodes, record the distribution relationship of the time segments on the nodes, and generate a node task allocation diagram;

[0125] The node bearing capacity is extracted according to the records in the distributed task scheduling logs. The task path resource data is obtained by summarizing the resource distribution data of each time segment. The distribution of time segments is allocated based on the node adaptability score. The specific method is to screen the time segment and the node score table in turn, match the time segment to the optimal node, and record the resource occupancy ratio of the node where the time segment is distributed at the same time. All allocation relationships are presented in the form of a path structure diagram of the task and the node, and a node task allocation diagram is generated.

[0126] Please refer to Figure 5 , and the specific steps of S4 are as follows:

[0127] S401: Based on the node task allocation diagram, extract the node task distribution data and the storage fragment call records, classify and sort the fragment call frequencies of each node, analyze the change in the length of the call path, summarize the correlation between the call frequency and the path length, and generate a fragment call and path analysis table;

[0128] Analyze the correlation between the fragment call frequency and the path length, and calculate according to the formula

[0129]

[0130] the average correlation value of the node fragment call and the path length.

[0131] In the formula, R cRepresents the correlation value between the call frequency and the path length, F i Represents the frequency of the i-th fragment call, L i Represents the path length of the i-th call. p represents the total number of node fragment calls.

[0132] The fragment call frequency is statistically obtained through the node task distribution data and the access logs in the stored fragment call records. The path length is measured according to the actual distance of the storage path between nodes. The call frequency is calculated by counting each fragment call, and the path length is obtained from the link distance records between nodes.

[0133] For example, the frequencies of three fragment calls within a certain node are 5, 8, and 12 (unit: times) respectively, the corresponding path lengths are 10, 15, and 20 (unit: distance units), and the total number of calls is 3. Substitute into the formula:

[0134]

[0135] The calculated correlation value is 136.67. This result indicates the average correlation level between the node fragment call frequency and the path length. Subsequently, the node storage path distribution strategy can be adjusted accordingly.

[0136] S402: Based on the fragment call and path analysis table, compare the distribution characteristics of the fragment call frequency and the path length, adjust the distribution of the node storage path, redefine the correspondence between the fragment and the node, identify the adjustment results of the storage path, and generate a storage path adjustment distribution map;

[0137] The adjustment of the node storage path is achieved by analyzing the call frequency and length distribution of each path segment, combining the maximum frequency value of each path in the node fragment call record, redefining the mapping rule between the fragment and the node, recording the distribution change of the call frequency after adjusting the path, and analyzing the distribution law of the adjusted path in combination with the change trend of the call frequency. In the distribution adjustment, the communication efficiency of the node storage path is used as a constraint condition to screen and verify the path optimization scheme, and finally a storage path adjustment distribution map is generated.

[0138] S403: Based on the storage path adjustment distribution map, analyze the distribution state of the node storage load, re-match the position relationship between the node and the fragment, summarize the balanced distribution characteristics between the fragment storages, and integrate the fragment position relationship between the nodes to generate an optimized structure for data fragment storage;

[0139] The distribution status of the node storage load is quantitatively analyzed according to the load ratio of the task path resource distribution to the storage fragmentation. The positional relationship between the node and the fragmentation is dynamically adjusted according to the load analysis result. The balanced distribution characteristic among the fragmentation storages is obtained by calculating the deviation of the load mean value among the fragmentations, and integrating the load distribution by combining the change records of the fragmentation positions among the nodes. Finally, the positional distribution of the nodes and the fragmentations is presented in a graphical structure to generate an optimized structure for data fragmentation storage.

[0140] Please refer to Figure 6 , and the specific steps of S5 are as follows:

[0141] S501: Based on the distribution information of the optimized structure for data fragmentation storage, extract the resource requirement parameters of the task chain and the time slice distribution characteristics, classify and summarize the resource occupation and time distribution patterns of the task chain, perform data integration according to the relevance of the resource occupation parameters, analyze the phased resource requirement characteristics of the task chain, and generate a resource and time distribution table for the task chain;

[0142] Calculate the mode correlation value of the resource occupation and time distribution of the task chain according to the formula

[0143]

[0144] Calculate the phased characteristic value of the resource occupation and time distribution.

[0145] In the formula, A t represents the average characteristic value of the resource occupation and time distribution, U k represents the resource occupation parameter in the k-th stage, T k represents the time length in the k-th stage, and q represents the total number of time stages.

[0146] The resource occupation parameter is extracted from the task chain resource requirement record in the optimized structure for data fragmentation storage, and the time distribution parameter is read from the time slice distribution characteristic file. The time length and resource occupation amount are statistically segmented according to the phased task execution process.

[0147] For example, the task chain is divided into three stages, the resource occupation parameters in each stage are 50, 70, and 90 (unit: resource unit) respectively, the time lengths are 10, 15, and 20 (unit: minutes) respectively, and the total number of stages is 3. Substitute into the formula:

[0148]

[0149] The calculated average characteristic value of the resource occupation and time distribution of the task chain is 1116.67. This result shows the average characteristics of the resource occupation intensity and time distribution mode of the task chain, providing a basis for subsequent data integration and analysis of phased resource requirements.

[0150] S502: Based on the task chain resource and time distribution table, conduct data matching analysis on the resource allocation ratio and link configuration parameters, summarize the adaptability of resource allocation among links, adjust the link resource occupancy distribution rules, optimize the coordination relationship between links and resources, and generate a link resource allocation and configuration diagram;

[0151] The resource allocation ratio is extracted from the link resource occupancy record file. According to the resource requirement parameters of the task chain, combined with the transmission bandwidth and resource limit conditions in the link configuration parameters, adjust the link resource allocation rules. By comparing the distribution characteristics and adaptability of resource occupancy among links, optimize the coordinated distribution rules of resources among links. In the optimization of resource allocation, analyze the impact of changes in the resource allocation ratio on the link configuration efficiency, and adjust the resource distribution rules based on the coordination of link resource occupancy to generate a link resource allocation and configuration diagram.

[0152] S503: Based on the link resource allocation and configuration diagram, classify and integrate the collaboration logic of multi-stage tasks, analyze the resource distribution characteristics among task links, adjust the matching of collaboration intensity and resource distribution, and summarize the distributed resource scheduling relationship of the task chain to generate a distributed task chain collaboration table;

[0153] The collaboration logic of multi-stage tasks is classified according to the distribution characteristics of task link resources. The collaboration intensity is adjusted and optimized according to the resource matching of each task chain. The integration of collaboration logic is carried out by analyzing the correlation between the resource occupancy trend of the task chain and the link configuration parameters. The distributed resource scheduling relationship of the task chain is derived from the resource allocation table of the distributed task chain. Combining the resource distribution law of multi-stage tasks and the allocation efficiency of link resources, generate a distributed task chain collaboration table.

[0154] Please refer to Figure 7 , a data processing system for a big data platform, including:

[0155] The task priority module extracts the time parameters, resource occupancy parameters, and processing complexity parameters of the task, calculates the weighted value of time and resources, integrates the complexity parameters, dynamically adjusts the weight ratio to match the task status, and constructs a task priority matrix;

[0156] The time segment module, based on the priority matrix, extracts the priority value and time length, decomposes the task into time segments, establishes the mapping between time segments and priority values, combines the resource requirements and status parameters, correlates the time segments with resource distribution, and generates a time segment distribution table;

[0157] The node allocation module, based on the time segment distribution table, extracts the time segment and resource requirement data, analyzes the node resource carrying capacity, matches the collaboration intensity, integrates the task path and communication cost parameters, allocates the task to the appropriate node, and generates a node task allocation diagram;

[0158] Based on the node task allocation graph, the storage optimization module analyzes the node task distribution data and the storage fragment call records, analyzes the fragment call frequency and the path length, adjusts the storage path distribution, reorganizes the corresponding relationship between nodes and fragments, optimizes the storage load balance, and generates an optimized structure for data fragment storage;

[0159] Based on the distribution information of the optimized structure for data fragment storage, the link collaboration module analyzes the resource requirements and time slice distribution of the task chain, integrates the resource allocation ratio and the link configuration parameters, adjusts the multi-stage task collaboration logic, and generates a distributed task chain collaboration table.

[0160] As mentioned above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A data processing method for a big data platform, characterized in that: The following steps are involved: S1: Extract the time parameters, resource occupancy parameters and processing complexity parameters of the task, calculate the weighted values ​​of time and resources, integrate the complexity parameters, dynamically adjust the weight ratio to match the task status, identify the priority value, and construct the task priority matrix; S2: Based on the priority matrix, extract the priority value and time length, decompose the task into time segments, and establish a mapping between the time segments and priority values. Combine resource requirements and state parameters, associate the time segments with resource distribution, integrate and analyze the task time distribution data, and generate a time segment distribution table. S3: Based on the time segment distribution table, extract the time segment and resource demand data, analyze the node resource carrying capacity, match the collaboration intensity, integrate the task path and communication cost parameters, optimize the distribution path, assign tasks to the adaptation nodes, record the mapping between tasks and nodes, and generate a node task allocation diagram; S4: Based on the node task allocation graph, analyze the node task distribution data and storage fragment call records, analyze the fragment call frequency and path length, adjust the storage path distribution, reorganize the corresponding relationship between nodes and fragments, optimize storage load balancing, integrate the fragment position relationship, and generate a data fragment storage optimization structure; S5: Based on the distribution information of the data fragment storage optimization structure, analyze the task chain resource requirements and time slice distribution, integrate the resource allocation ratio and link configuration parameters, adjust the multi-stage task collaboration logic, optimize the task link resource allocation, and generate a distributed task chain collaboration table.

2. The data processing method of the big data platform according to claim 1, characterized in that: The task priority matrix includes priority weight distribution, dynamic weight adjustment results, and comprehensive task scores. The time segment distribution table includes a time segment set, a mapping relationship between priority values ​​and time segments, and a resource demand status table. The node task allocation graph includes a node collaboration relationship table, a task allocation path graph, and communication cost distribution data. The data fragment storage optimization structure includes a storage path distribution optimization plan, a fragment and node mapping relationship, and a call frequency statistics table. The distributed task chain collaboration table includes task chain stage division, resource allocation parameters, and a link collaboration relationship table.

3. The data processing method of the big data platform according to claim 1, characterized in that: The specific steps of extracting the time parameters, resource occupancy parameters and processing complexity parameters of the task, calculating the weighted values ​​of time and resources, integrating the complexity parameters, dynamically adjusting the weight ratio to match the task status, identifying the priority value, and constructing the task priority matrix are as follows: S101: Based on the time parameter, resource occupancy parameter and processing complexity parameter of the task, according to the weight distribution of each parameter in the task state, calculate the time and resource weighted value of each task, dynamically update the weight comparison task state change, record the relationship between task time and resource consumption, and generate a task time and resource weighted value table; S102: Based on the task time and resource weighted value table, standardize the complexity parameters of the task, combine the time and resource weighted values, establish the state influencing factor of each task, summarize the matching relationship between the weighted value and the complexity parameter, optimize the influence value difference between task states, and generate a task state influencing factor list; S103: Based on the task status influencing factor list, compare and analyze the task priority values ​​one by one, adjust the priority parameter relationship according to the weight gradient, arrange the task priorities in descending order, establish and analyze the mapping rules between task status and priority, and generate a task priority matrix.

4. The data processing method of the big data platform according to claim 1, characterized in that: Based on the priority matrix, the priority value and time length are extracted, the task is decomposed into time segments, and a mapping between time segments and priority values ​​is established. In combination with resource requirements and state parameters, time segments are associated with resource distribution, and task time distribution data is integrated and analyzed. The specific steps to generate a time segment distribution table are as follows: S201: Based on the priority matrix, extract the priority value and time length parameter of each task, classify the priority value, split multiple time segments according to the task time, calculate the priority value corresponding to the time segment, analyze the distribution characteristics, associate the task priority with the time segment, and generate a time segment priority mapping table; S202: Based on the time segment priority mapping table and in combination with the resource requirement parameters of the task, the resource consumption of each time segment is extracted, the distribution pattern of the time segment and resource usage is analyzed, and the resource matching value of the time segment is calculated according to the task resource usage rule to obtain the time segment resource distribution data; S203: Based on the time segment resource distribution data, summarize the resource distribution characteristics of the time segments and tasks, match the coordination of task time distribution and resources, adjust the resource allocation rules of the time segments, build a classification structure of time distribution parameters, and generate a time segment distribution table.

5. The data processing method of the big data platform according to claim 4, characterized in that: The calculation formula for the priority value corresponding to the time segment is specifically: Among them, P s1 Represents the priority value corresponding to the time segment, P k1 represents the priority value of the kth task, T k1 represents the time length parameter of the kth task, q1 represents the total number of tasks in the time segment, Represents the sum of the duration of all tasks within a time segment.

6. The data processing method of the big data platform according to claim 1, characterized in that: Based on the time segment distribution table, extract the time segment and resource demand data, analyze the node resource carrying capacity, match the collaboration intensity, integrate the task path and communication cost parameters, optimize the distribution path, assign tasks to the adaptation nodes, record the mapping between tasks and nodes, and generate the node task allocation diagram in the following specific steps: S301: Based on the time segment distribution table, extract resource demand data of each time segment, classify and summarize the resource demand, calculate the carrying capacity value of the node resource, match the corresponding relationship between the resource demand and the node carrying value, establish the adaptability score of the time segment on the node, and generate a node resource adaptation table; S302: Based on the node resource adaptation table, combined with the time segment collaboration strength parameter and the node communication cost parameter, analyze the task path distribution, adjust the matching of the path connection and the communication cost between nodes, optimize the distribution mode of the task path resources, construct and analyze the relationship between the task distribution path and the communication cost, and generate a task distribution path diagram; S303: Based on the task distribution path diagram, extract the node carrying capacity and task path resource data, distribute the tasks to the adaptation nodes according to the time segment distribution, record the distribution relationship of the time segments on the nodes, and generate a node task distribution diagram.

7. The data processing method of the big data platform according to claim 6, characterized in that: The bearing capacity value calculation formula is as follows: Among them, C r1 Represents the node resource carrying capacity value, R i3 represents the resource demand in the i-th time segment, T i4 represents the time length parameter of the i-th time segment, q2 represents the total number of time segments, Represents the sum of the duration of all time segments.

8. The data processing method of the big data platform according to claim 1, characterized in that: Based on the node task allocation graph, the specific steps of analyzing the node task distribution data and storage fragment call records, analyzing the fragment call frequency and path length, adjusting the storage path distribution, reorganizing the corresponding relationship between nodes and fragments, optimizing storage load balancing, integrating the fragment position relationship, and generating the data fragment storage optimization structure are as follows: S401: Based on the node task allocation graph, extract node task distribution data and storage fragment call records, classify and sort the fragment call frequency of each node, analyze the length change of the call path, summarize the correlation between the call frequency and the path length, and generate a fragment call and path analysis table; S402: Based on the fragment call and path analysis table, compare the distribution characteristics of the fragment call frequency and the path length, adjust the distribution of the node storage path, redefine the corresponding relationship between the fragments and the nodes, identify the adjustment result of the storage path, and generate a storage path adjustment distribution map; S403: Based on the storage path, the distribution map is adjusted, the distribution state of the node storage load is analyzed, the position relationship between the nodes and the fragments is re-matched, the balanced distribution characteristics between the fragment storage are summarized, and the fragment position relationship between the nodes is integrated to generate a data fragment storage optimization structure.

9. The data processing method of the big data platform according to claim 1, characterized in that: Based on the distribution information of the data fragment storage optimization structure, the specific steps of analyzing the task chain resource requirements and time slice distribution, integrating the resource allocation ratio and link configuration parameters, adjusting the multi-stage task collaboration logic, optimizing the task link resource allocation, and generating a distributed task chain collaboration table are as follows: S501: Based on the distribution information of the data fragment storage optimization structure, extract the task chain resource demand parameters and time slice distribution characteristics, classify and summarize the resource occupancy and time distribution patterns of the task chain, integrate data according to the correlation of the resource occupancy parameters, analyze the staged resource demand characteristics of the task chain, and generate a task chain resource and time distribution table; S502: Based on the task chain resource and time distribution table, perform data matching analysis on the resource allocation ratio and the link configuration parameters, summarize the adaptability of resource allocation between links, adjust the link resource occupancy distribution rule, and optimize the link and resource coordination relationship to generate a link resource allocation and configuration diagram; S503: Based on the link resource allocation and configuration diagram, classify and integrate the collaboration logic of multi-stage tasks, analyze the resource distribution characteristics between task links, adjust the matching of collaboration intensity and resource distribution, summarize the distributed resource scheduling relationship of the task chain, and generate a distributed task chain collaboration table.

10. A data processing system for a big data platform, characterized in that: According to a data processing method for a big data platform according to any one of claims 1 to 9, the system comprises: The task priority module extracts the time parameters, resource occupancy parameters and processing complexity parameters of the task, calculates the weighted values ​​of time and resources, integrates the complexity parameters, dynamically adjusts the weight ratio to match the task status, and constructs the task priority matrix; The time segment module extracts the priority value and time length based on the priority matrix, decomposes the task into time segments, and establishes a mapping between the time segments and the priority values. It combines the resource requirements and state parameters, associates the time segments with the resource distribution, and generates a time segment distribution table. The node allocation module extracts the time segment and resource demand data based on the time segment distribution table, analyzes the node resource carrying capacity, matches the collaboration intensity, integrates the task path and communication cost parameters, allocates the task to the adaptation node, and generates a node task allocation diagram; The storage optimization module analyzes the node task distribution data and storage fragment call records based on the node task allocation diagram, analyzes the fragment call frequency and path length, adjusts the storage path distribution, reorganizes the corresponding relationship between nodes and fragments, optimizes the storage load balance, and generates a data fragment storage optimization structure; The link collaboration module analyzes the task chain resource requirements and time slice distribution based on the distribution information of the data fragment storage optimization structure, integrates the resource allocation ratio and link configuration parameters, adjusts the multi-stage task collaboration logic, and generates a distributed task chain collaboration table.