Data processing method based on cloud computing platform and cloud computing platform
By analyzing task types on the cloud computing platform and intelligently matching server resources, and optimizing task scheduling, the cloud computing platform's inefficiency, latency and scalability problems when processing complex data processing tasks are solved, and more efficient, flexible and scalable data processing capabilities are achieved.
Patent Information
- Application Number
- CN202510459781.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-14
AI Technical Summary
When handling complex data processing tasks, cloud computing platforms face challenges such as task type diversity, unreasonable allocation of server resources, task scheduling delays, and system scalability and compatibility.
The task type is parsed through the data processing task analysis module, and the most suitable distributed data server is matched according to the task needs. The intelligent matching algorithm is used to ensure that the task is executed by the most suitable server, the task scheduling mechanism is optimized to reduce latency, and the method is designed to adapt to system expansion and compatible with different types of data processing tasks.
Improve processing efficiency, reduce processing delay, enhance system flexibility and resource utilization, and improve system scalability.
Smart Images

Figure CN119988038A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to a data processing method based on a cloud computing platform and a cloud computing platform. Background Technology
[0002] With the rapid development of information technology, the amount of data has exploded, and traditional single-machine processing methods can no longer meet the needs of efficient and fast data processing. As a new IT service model that integrates computing, storage, and network services, cloud computing platforms have become an ideal choice for processing large-scale data with their elastic resource expansion, on-demand allocation, and high availability. However, cloud computing platforms still face many challenges when processing complex data processing tasks: Diversity of task types: Data processing tasks may involve single computing tasks (such as big data analysis, image processing, etc.) or combined tasks (i.e., multiple interdependent tasks are executed in a specific order). Existing methods often lack flexible task parsing and scheduling mechanisms when processing combined tasks, resulting in low processing efficiency.
[0003] Irrational allocation of server resources: Cloud computing platforms are usually composed of various types of distributed data servers, including computing data servers, storage data servers, and comprehensive data servers. How to quickly and accurately match the most appropriate server resources according to task requirements to avoid idle or overloaded resources is a problem that needs to be solved urgently.
[0004] Task scheduling delay: During the execution of combined tasks, data transmission and state synchronization between tasks are key factors affecting the overall processing efficiency. An unreasonable task scheduling strategy will increase data transmission delay and reduce processing speed.
[0005] System scalability and compatibility: As the scale of cloud computing platforms expands, ensuring that data processing methods can efficiently adapt to new server resources and be compatible with different types of data processing tasks is an important consideration for improving system flexibility and scalability. SUMMARY OF THE INVENTION
[0006] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a data processing method based on a cloud computing platform, comprising the following steps: Step 1: The data processing task parsing module obtains the task processing type according to the data processing task sent by the client. If it is a single task, it proceeds to step 2; if it is a combined task, it proceeds to step 3; Step 2: According to the type of distributed data server required by the data processing task, the sequence of distributed data servers of the same type is matched to the corresponding distributed data server according to the sequence of distributed data servers of the same type, and the data processing task is sent to the matched distributed data server, and then the process goes to step 7; Step 3: The data processing task parsing module obtains the task processing sequence according to the combined task, and obtains the distributed data server type sequence according to the distributed data server type required by each task in the task sequence; Step 4: According to the distributed data server type sequence, the distributed data server management module obtains the distributed data server sequence corresponding to the distributed data server type; Step 5: According to the first task information in the task processing sequence and the distributed data server sequence of the corresponding distributed data server type, the first distributed data server is matched, and the distributed data server of the next task is matched according to the first distributed data server and the distributed data server sequence of the corresponding distributed data server type of the next task, until the distributed data server matching of all tasks is completed, the task processing distributed server sequence is obtained, and the task processing distributed server sequence and the combined task are packaged and sent to the first distributed data server; Step 6: After the first distributed data server completes the task, it sends the task to the next distributed data server according to the task processing distributed server sequence until all tasks are completed; Step 7: Complete data processing based on the cloud computing platform.
[0007] Furthermore, the data processing task parsing module obtains the task processing type according to the data processing task sent by the client, including: If the data processing task sent by the client is a single type of computing task, it is a single task; otherwise, it is a combined task; the single type of computing task means that all tasks in the data processing task are computing tasks of the same type.
[0008] Furthermore, the type of distributed data server required by the data processing task is used to obtain a sequence of distributed data servers of the same type, and the corresponding distributed data server is matched according to the sequence of distributed data servers of the same type, including: The types of distributed data servers required by the data processing task include computing power data servers, storage data servers and comprehensive data servers; according to the types of distributed data servers required by the data processing task, a sequence of distributed data servers of the same type is obtained; according to the set matching features and the distributed data server sequence, a matching distributed data server is obtained; the set matching features include one or more of network delay, transmission bandwidth and data processing speed.
[0009] Furthermore, the data processing task parsing module obtains a task processing sequence according to the combined task, and obtains a distributed data server type sequence according to the distributed data server type required by each task in the task sequence, including: The data processing task parsing module obtains each task and the task processing order in the combined task, obtains the task processing order sequence, and obtains the distributed data server type sequence according to the distributed data server type required by each task in the task processing order sequence.
[0010] Further, the first distributed data server is obtained by matching the first task information in the task processing sequence and the distributed data server sequence corresponding to the distributed data server type, including: According to the matching feature set for the first task, the first task information and the distributed data server sequence corresponding to the distributed data server type, the matching first distributed data server is obtained.
[0011] Further, the distributed data server sequence of the first distributed data server obtained and the corresponding distributed data server type of the next task is matched to obtain the distributed data server of the next task, including: According to the matching features set for the next task, the feature matching degree of the first distributed data server and each distributed data server in the distributed data server sequence of the corresponding distributed data server type of the next task is obtained, wherein the distributed data server corresponding to the maximum feature matching degree is the distributed data server of the next task to be matched; The feature conformity is: the feature difference between the matching feature of the first distributed data server and the distributed data server sequence of the corresponding distributed data server type of the next task and the matching feature set by the task. The smaller the difference, the greater the feature conformity.
[0012] Further, after the first distributed data server completes the task, it sends the task to the next distributed data server according to the task processing distributed server sequence, including: After the first distributed data server completes the task, it packages the first task processing result, the remaining tasks, and the task processing distributed server sequence and sends them to the next distributed data server.
[0013] The cloud computing platform is characterized by applying the data processing method based on the cloud computing platform, including a distributed data server, a distributed data server management module, a data processing task parsing module and a communication module; The distributed data server, distributed data server management module, and data processing task analysis module are respectively connected to the communication module for communication.
[0014] The beneficial effects of the present invention are: improving processing efficiency: by fine-tuning task analysis and matching with intelligent servers, ensuring that tasks are executed by the most suitable server, significantly improving data processing efficiency.
[0015] Reduce processing delay: Optimize the task scheduling mechanism, reduce the waiting time between tasks and data transmission time, and effectively reduce the overall processing delay.
[0016] Enhance system flexibility: Support flexible processing of single tasks and combined tasks to adapt to diverse data processing needs.
[0017] Improve resource utilization: Real-time monitoring of server resource status, dynamic adjustment of task allocation, avoid resource idleness or overload, and improve resource utilization.
[0018] Enhance system scalability: The design method is easy to expand and can quickly integrate new server resources to adapt to the growth of the cloud computing platform. Brief Description of the Figures
[0019] Figure 1 A flow chart of a data processing method based on a cloud computing platform; Figure 2 This is a schematic diagram of the cloud computing platform principle. Specific implementation method
[0020] The technical solution of the present invention is further described in detail below in conjunction with the accompanying drawings, but the protection scope of the present invention is not limited to the following.
[0021] The features and performance of the present invention are further described in detail below in conjunction with the embodiments.
[0022] If Figure 1 As shown, the data processing method based on the cloud computing platform includes the following steps: Step 1: The data processing task parsing module obtains the task processing type according to the data processing task sent by the client. If it is a single task, it proceeds to step 2; if it is a combined task, it proceeds to step 3; Step 2: According to the type of distributed data server required by the data processing task, the sequence of distributed data servers of the same type is matched to the corresponding distributed data server according to the sequence of distributed data servers of the same type, and the data processing task is sent to the matched distributed data server, and then the process goes to step 7; Step 3: The data processing task parsing module obtains the task processing sequence according to the combined task, and obtains the distributed data server type sequence according to the distributed data server type required by each task in the task sequence; Step 4: According to the distributed data server type sequence, the distributed data server management module obtains the distributed data server sequence corresponding to the distributed data server type; Step 5: According to the first task information in the task processing sequence and the distributed data server sequence of the corresponding distributed data server type, the first distributed data server is matched, and the distributed data server of the next task is matched according to the first distributed data server and the distributed data server sequence of the corresponding distributed data server type of the next task, until the distributed data server matching of all tasks is completed, the task processing distributed server sequence is obtained, and the task processing distributed server sequence and the combined task are packaged and sent to the first distributed data server; Step 6: After the first distributed data server completes the task, it sends the task to the next distributed data server according to the task processing distributed server sequence until all tasks are completed; Step 7: Complete data processing based on the cloud computing platform.
[0023] The data processing task parsing module obtains the task processing type according to the data processing task sent by the client, including: If the data processing task sent by the client is a single type of computing task, it is a single task; otherwise, it is a combined task; the single type of computing task means that all tasks in the data processing task are computing tasks of the same type.
[0024] According to the type of distributed data server required by the data processing task, a sequence of distributed data servers of the same type is obtained, and the corresponding distributed data server is matched according to the sequence of distributed data servers of the same type, including: The types of distributed data servers required by the data processing task include computing power data servers, storage data servers and comprehensive data servers; according to the types of distributed data servers required by the data processing task, a sequence of distributed data servers of the same type is obtained; according to the set matching features and the distributed data server sequence, a matching distributed data server is obtained; the set matching features include one or more of network delay, transmission bandwidth and data processing speed.
[0025] The data processing task parsing module obtains a task processing sequence according to the combined task, and obtains a distributed data server type sequence according to the distributed data server type required by each task in the task sequence, including: The data processing task parsing module obtains each task and the task processing order in the combined task, obtains the task processing order sequence, and obtains the distributed data server type sequence according to the distributed data server type required by each task in the task processing order sequence.
[0026] The matching of the first distributed data server according to the first task information in the task processing sequence and the distributed data server sequence corresponding to the distributed data server type to obtain the first distributed data server includes: According to the matching feature set for the first task, the first task information and the distributed data server sequence corresponding to the distributed data server type, the matching first distributed data server is obtained.
[0027] The distributed data server sequence of the first distributed data server obtained and the corresponding distributed data server type of the next task is matched to obtain the distributed data server of the next task, including: According to the matching features set for the next task, the feature matching degree of the first distributed data server and each distributed data server in the distributed data server sequence of the corresponding distributed data server type of the next task is obtained, wherein the distributed data server corresponding to the maximum feature matching degree is the distributed data server of the next task to be matched; The feature conformity is: the feature difference between the matching feature of the first distributed data server and the distributed data server sequence of the corresponding distributed data server type of the next task and the matching feature set by the task. The smaller the difference, the greater the feature conformity.
[0028] After the first distributed data server completes the task, the task is sent to the next distributed data server according to the task processing distributed server sequence, including: After the first distributed data server completes the task, it packages the first task processing result, the remaining tasks, and the task processing distributed server sequence and sends them to the next distributed data server.
[0029] If Figure 2 As shown, a cloud computing platform, applying the data processing method based on the cloud computing platform, includes a distributed data server, a distributed data server management module, a data processing task parsing module and a communication module; The distributed data server, distributed data server management module, and data processing task analysis module are respectively connected to the communication module for communication.
[0030] Specifically, the present invention provides a data processing method based on a cloud computing platform, comprising the following steps: Step 1: Task type analysis The data processing task parsing module receives the data processing tasks sent by the client, and determines whether the task type is a single task or a combined task by analyzing the task content. A single task refers to a computing task in which all operations in the task are of the same type, such as only involving data analysis; a combined task contains multiple subtasks of different types or that need to be executed in sequence.
[0031] Step 2: Single Task Processing If the task is a single task, the data processing task parsing module determines the required distributed data server type (such as computing power data server, storage data server or comprehensive data server) according to the task requirements.
[0032] The distributed data server management module selects the most suitable server from the same type of server sequence according to the server type and the set matching characteristics (such as network delay, transmission bandwidth, data processing speed, etc.).
[0033] Send the data processing task to the matched distributed data server, and directly proceed to step 7 after execution.
[0034] Step 3: Combined Task Analysis For combined tasks, the data processing task parsing module first parses out the execution order of the tasks to form a task processing sequence. According to the requirements of each subtask, the required distributed data server type is determined to form a distributed data server type sequence.
[0035] Step 4: Get the server sequence The distributed data server management module generates an available server sequence for each type of server based on the distributed data server type sequence.
[0036] Step 5: Task-Server Matching Based on the first task information in the task processing sequence and its corresponding server type sequence, combined with the set matching features, select the first distributed data server.
[0037] For subsequent tasks, the execution server of each task is determined one by one based on the matching characteristics of the current task execution server and the server in the next task server type sequence (such as considering data transmission efficiency, processing speed, etc.), forming a task processing distributed server sequence. The task processing distributed server sequence and the combined task are packaged and sent to the first distributed data server.
[0038] Step 6: Task Scheduling and Execution After the first distributed data server completes its task, it passes the task processing result and the remaining task information to the next distributed data server according to the task processing distributed server sequence, and executes in turn until all tasks are completed.
[0039] Step Seven: Task Completion After all tasks are executed, the data processing based on the cloud computing platform is completed, and the processing result is returned to the client.
[0040] Embodiment 1: Cloud Processing of a Single Data Analysis Task An e-commerce company needs to perform real-time analysis on its massive user behavior data to gain insights into user purchase preferences and optimize the recommendation algorithm. This task is a single type of data analysis task, mainly involving the aggregation, statistics, and pattern recognition of a large amount of data.
[0041] Implementation Steps: Task Type Analysis: The e-commerce company submits the data analysis task to the cloud computing platform through the client. After receiving the task, the data processing task analysis module identifies that this task is a single task and the main requirement is data analysis by parsing the task script.
[0042] Single Task Processing: The data processing task analysis module determines that a high-computing-power data server is needed to execute this task according to the task requirements. The distributed data server management module monitors the status of all current computing-power data servers, including CPU usage, available memory, network latency, etc., and generates a sequence of available servers.
[0043] An intelligent matching algorithm is used to select the optimal server from the sequence of computing-power data servers, taking into account both processing speed and network latency.
[0044] Task Execution and Completion: The data analysis task is sent to the matched optimal computing-power data server for execution. After the server completes the data analysis, it directly returns the processing result to the client, completing the entire data processing process.
[0045] Embodiment 2: Cloud Processing of a Composite Data Processing Task A research institution needs to perform preprocessing, feature extraction, and machine learning model training on a large amount of collected scientific research data. These tasks need to be executed in sequence, and each step has different requirements for computing resources and storage resources.
[0046] Implementation Steps: Task Type Analysis: The research institution submits the composite data processing task to the cloud computing platform through the client. After receiving the task, the data processing task analysis module parses that the task includes three subtasks: data preprocessing, feature extraction, and model training, and needs to be executed in sequence.
[0047] Combined task analysis: The data processing task analysis module forms a task processing sequence: data preprocessing -> feature extraction -> model training. According to the requirements of each subtask, it is determined that data preprocessing requires a storage data server, feature extraction requires a computing data server, and model training requires a comprehensive data server.
[0048] Server sequence acquisition: The distributed data server management module generates storage data server sequence, computing power data server sequence and comprehensive data server sequence according to server type requirements.
[0049] Task-server matching: According to the task processing sequence, first select the optimal storage data server for the data preprocessing task. Then, considering the data transmission efficiency and processing speed, select the optimal computing power data server for the feature extraction task. Finally, select the optimal comprehensive data server for the model training task to form a task processing distributed server sequence.
[0050] Task scheduling execution: The combined tasks and task processing distributed server sequences are packaged and sent to the first storage data server. After the storage data server completes data preprocessing, it passes the results to the computing power data server for feature extraction. After feature extraction is completed, the results are then passed to the comprehensive data server for model training. After model training is completed, the final processing results are returned to the client.
[0051] Task completed: All subtasks were successfully executed in sequence, and the scientific research institution obtained the scientific research data analysis report after preprocessing, feature extraction and model training.
Claims
1. A data processing method based on a cloud computing platform, characterized in that: The steps include: Step 1: The data processing task parsing module obtains the task processing type according to the data processing task sent by the client. If it is a single task, it proceeds to step 2; if it is a combined task, it proceeds to step 3; Step 2: According to the type of distributed data server required by the data processing task, a sequence of distributed data servers of the same type is obtained, and the corresponding distributed data server is matched according to the sequence of distributed data servers of the same type, and the data processing task is sent to the matched distributed data server, and then the process goes to step 7; Step 3: The data processing task parsing module obtains a task processing order sequence according to the combined task, and obtains a distributed data server type sequence according to the distributed data server type required by each task in the task sequence; Step 4: According to the distributed data server type sequence, the distributed data server management module obtains the distributed data server sequence corresponding to the distributed data server type; Step 5: According to the first task information in the task processing sequence and the distributed data server sequence of the corresponding distributed data server type, the first distributed data server is matched, and according to the obtained first distributed data server and the distributed data server sequence of the corresponding distributed data server type of the next task, the distributed data server of the next task is matched, until the distributed data server matching of all tasks is completed, the task processing distributed server sequence is obtained, and the task processing distributed server sequence and the combined task are packaged and sent to the first distributed data server; Step 6: After the first distributed data server completes the task, it sends the task to the next distributed data server according to the task processing distributed server sequence until all tasks are completed; Step seven: Complete data processing based on the cloud computing platform.
2. The data processing method based on the cloud computing platform according to claim 1, characterized in that: The data processing task parsing module obtains the task processing type according to the data processing task sent by the client, including: If the data processing task sent by the client is a single type of computing task, it is a single task; otherwise, it is a combined task; the single type of computing task means that all tasks in the data processing task are computing tasks of the same type.
3. The data processing method based on the cloud computing platform according to claim 2, characterized in that: The method of obtaining a sequence of distributed data servers of the same type according to the type of distributed data server required by the data processing task, and matching the corresponding distributed data server according to the sequence of distributed data servers of the same type includes: The types of distributed data servers required by the data processing tasks include computing power data servers, storage data servers and comprehensive data servers; according to the types of distributed data servers required by the data processing tasks, a sequence of distributed data servers of the same type is obtained; according to the set matching features and the distributed data server sequence, a matching distributed data server is obtained; the set matching features include one or more of network delay, transmission bandwidth and data processing speed.
4. The data processing method based on the cloud computing platform according to claim 2, characterized in that: The data processing task parsing module obtains a task processing order sequence according to the combined task, and obtains a distributed data server type sequence according to the distributed data server type required by each task in the task sequence, including: The data processing task parsing module obtains each task and the task processing order in the combined task, obtains a task processing order sequence, and obtains a distributed data server type sequence according to the distributed data server type required by each task in the task processing order sequence.
5. The data processing method based on the cloud computing platform according to claim 4 is characterized in that: The matching of obtaining the first distributed data server according to the first task information in the task processing sequence and the distributed data server sequence corresponding to the distributed data server type includes: A matched first distributed data server is obtained according to the matching feature set for the first task, the first task information, and a distributed data server sequence corresponding to the distributed data server type.
6. The data processing method based on the cloud computing platform according to claim 5, characterized in that: The step of matching the distributed data server of the next task according to the obtained first distributed data server and the distributed data server sequence of the corresponding distributed data server type of the next task comprises: According to the matching features set for the next task, the feature conformity of the matching features of the first distributed data server and each distributed data server in the distributed data server sequence of the corresponding distributed data server type of the next task is obtained, wherein the distributed data server corresponding to the maximum feature conformity is the distributed data server of the next task to be matched; The feature conformity degree is: the feature difference between the matching features of the distributed data servers in the distributed data server sequence of the first distributed data server and the corresponding distributed data server type of the next task and the matching features set for the task. The smaller the difference, the greater the feature conformity degree.
7. The data processing method based on the cloud computing platform according to claim 6, characterized in that: After the first distributed data server completes the task, the task is sent to the next distributed data server according to the task processing distributed server sequence, including: After the first distributed data server completes the task, it packages the first task processing result, the remaining tasks and the task processing distributed server sequence and sends them to the next distributed data server.
8. Cloud computing platform, characterized in that: A data processing method based on a cloud computing platform according to any one of claims 1 to 7 is applied, comprising a distributed data server, a distributed data server management module, a data processing task parsing module and a communication module; The distributed data server, distributed data server management module, and data processing task analysis module are respectively connected to the communication module for communication.
Citation Information
Patent Citations
Distributed computer cloud computing processing method
CN108847981A
Intelligent energy-saving scheduling method for data center
CN116382863A
Task scheduling method and apparatus, computer device and storage medium
WO2022247105A1
Distributed data processing method and apparatus, and electronic device
WO2025000747A1