Computing power resource scheduling method and device, computer equipment and storage medium

By obtaining computing power scheduling requests in the Hadoop cluster, selecting suitable heterogeneous computing power nodes, and allocating data processing subtasks, the problem of unsuitable computing power resource scheduling in the transformation of information technology innovation is solved, and efficient data processing and resource management are achieved.

CN121880022APending Publication Date: 2026-04-17PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-01-15
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing technologies, the computing resource scheduling method of Hadoop clusters is not suitable for heterogeneous framework computing during the transformation of information technology innovation, resulting in low data processing efficiency.

Method used

By acquiring computing power scheduling requests, determining the requesting application type, the data file to be processed, and the processing requirements, matching candidate computing power nodes are selected from the heterogeneous computing power node pool. Data processing sub-task target nodes are allocated according to the node load parameters, thereby achieving unified scheduling of heterogeneous computing power resources.

Benefits of technology

It improves data processing efficiency, supports intelligent heterogeneous computing resource management in the context of information technology innovation upgrades, and shields the complexity of heterogeneous computing resource scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121880022A_ABST
    Figure CN121880022A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a computing power resource scheduling method and device, computer equipment and a storage medium in the fields of financial science and technology and medical health. The method comprises the following steps: acquiring a computing power scheduling request, and determining a request application type, a to-be-processed data file and processing demand information according to the computing power scheduling request; determining a plurality of candidate computing power nodes matched with the request application type from a heterogeneous computing power node pool; determining at least one data processing subtask according to the to-be-processed data file and the processing demand information; and according to the node load parameters of the candidate computing power nodes, allocating a sub-task target node from all the candidate computing power nodes for each data processing sub-task, so that the sub-task target node completes the allocated data processing sub-task. Based on unified scheduling of heterogeneous computing power resources, the optimal heterogeneous computing power resources are matched for the resource demand of the computing power scheduling request, and the data processing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and more particularly to a method, apparatus, computer equipment, and storage medium for scheduling computing resources in the fields of financial technology and healthcare. Background Technology

[0002] Hadoop clusters are software frameworks capable of distributed processing of massive amounts of data, widely used in fintech and healthcare. In the financial and insurance industry, customer basic information, purchase records, claims history, and other data are scattered across multiple databases, requiring insurance companies to integrate, access, and manage this data in their daily operations. In the healthcare industry, diverse medical data, such as electronic medical records, medical images, gene sequences, and health monitoring data, are stored in different data sources. Distributed systems enable unified access, facilitating the collection and integration of medical data and providing a comprehensive data foundation for diagnosis and treatment. Hadoop clusters, through their distributed architecture, have restructured the data processing paradigm in finance and healthcare, driving the industry towards data-driven decision-making. During the transformation to domestically developed Hadoop clusters, data and tasks from non-domestic-development Hadoop clusters need to be migrated to domestically developed Hadoop clusters. Existing Hadoop clusters, using a single architecture, rely on two sets of computing and storage resources to complete a large amount of data migration and comparison work, making it difficult to effectively handle the complex scenarios of heterogeneous framework computing and reducing data processing efficiency. Summary of the Invention

[0003] Therefore, it is necessary to provide a computing resource scheduling method, apparatus, computer equipment, and storage medium to address the aforementioned technical problems, thereby solving the issues that existing computing resource scheduling methods are not applicable to heterogeneous framework computing and have low data processing efficiency.

[0004] A computing resource scheduling method, comprising: Obtain a computing power scheduling request, and determine the requesting application type, the data file to be processed, and the processing requirements based on the computing power scheduling request; Multiple candidate computing nodes matching the requested application type are identified from the heterogeneous computing node pool; Based on the data file to be processed and the processing requirements information, at least one data processing subtask is determined; Based on the node load parameters of the candidate computing power nodes, a subtask target node is allocated from all the candidate computing power nodes for each data processing subtask, so that the subtask target node completes the allocated data processing subtask.

[0005] A computing resource scheduling device, comprising: The request acquisition module is used to acquire computing power scheduling requests and determine the request application type, data file to be processed, and processing requirement information based on the computing power scheduling requests. The application type matching module is used to determine multiple candidate computing nodes that match the requested application type from the heterogeneous computing power node pool; The subtask determination module is used to determine at least one data processing subtask based on the data file to be processed and the processing requirement information; The computing power node allocation module is used to allocate a subtask target node from all the candidate computing power nodes for each data processing subtask according to the node load parameters of the candidate computing power nodes, so that the subtask target node completes the allocated data processing subtask.

[0006] A computer device includes a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein the processor implements the above-described computing resource scheduling method when executing the computer-readable instructions.

[0007] A computer-readable storage medium storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the computing resource scheduling method described above.

[0008] In the aforementioned computing power resource scheduling method, apparatus, computer equipment, and storage medium, the computing power resource scheduling method obtains a computing power scheduling request, determines the requested application type, the data file to be processed, and processing requirement information based on the computing power scheduling request; identifies multiple candidate computing power nodes matching the requested application type from a heterogeneous computing power node pool; determines at least one data processing sub-task based on the data file to be processed and the processing requirement information; and allocates a sub-task target node for each data processing sub-task from all candidate computing power nodes based on the node load parameters of the candidate computing power nodes, enabling the sub-task target node to complete the assigned data processing sub-task. This invention, based on unified scheduling of heterogeneous computing power resources, shields the complexity of heterogeneous computing power resource scheduling, matches the optimal heterogeneous computing power resources to the requested resource requirements, improves data processing efficiency, and can be applied to intelligent heterogeneous computing power resource management in domestic IT innovation upgrade scenarios. Attached Figure Description

[0009] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a schematic diagram of an application environment for a computing resource scheduling method according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating a computing resource scheduling method according to an embodiment of the present invention; Figure 3 This is a schematic diagram of a process in which the computing resource scheduling method of the present invention is applied to a Hadoop cluster. Figure 4 This is a schematic diagram of a computing resource scheduling device according to an embodiment of the present invention; Figure 5 This is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation

[0011] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0012] The computing resource scheduling method provided in this embodiment can be applied to, for example, Figure 1 In this application environment, the client communicates with the server. Clients include, but are not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0013] The computing resource scheduling method in this embodiment can be applied to the financial and insurance industry. On the one hand, risk assessment in daily operations requires comprehensive consideration of multiple data sources, such as customer credit records, health status, and claims history. This data may come from multiple data sources, including internal databases, external credit reporting agencies, and medical institutions. When a data access request is received, computing resource scheduling is needed to migrate data from external data sources to achieve unified access and data integration. On the other hand, when insurance companies replace their existing non-IT-compliant Hadoop clusters with Intel processors in IT-compliant Hadoop clusters with Hygon processors in IT-compliant Hadoop clusters during their development, heterogeneous computing resource scheduling is needed to complete a large number of data processing tasks, such as task migration, data migration, and data comparison.

[0014] The data source switching method in this embodiment can also be applied to the healthcare industry. Diverse types of medical data, such as electronic medical records, medical images, gene sequences, and health monitoring data, are stored in different data sources within a Hadoop cluster. When the Hadoop cluster is deployed on a heterogeneous chip distributed system architecture, heterogeneous computing power resource scheduling is required to access and integrate data from different data sources.

[0015] In one embodiment, such as Figure 2 As shown, a computing resource scheduling method is provided and applied to the server, including the following steps S10-S40.

[0016] S10. Obtain a computing power scheduling request, and determine the request application type, data file to be processed, and processing requirement information based on the computing power scheduling request.

[0017] In essence, a computing power scheduling request is a specific requirement used to request the server to allocate computing power nodes for data analysis and processing. Users need to process structured data in a Hadoop cluster. This can be done by inputting information through the client's interface. The client generates a computing power scheduling request based on the user's input and sends it to the server. The server parses the request to obtain the requesting application type, the data file to be processed, and the processing requirements. The requesting application type refers to the processor architecture type corresponding to the client application that needs to allocate computing power nodes. For example, a computing power scheduling request sent by a client application using an Intel processor would be classified as x86. The data file to be processed refers to the structured data file that needs to be analyzed and processed. The processing requirements information describes the methods and results of the data analysis and processing, i.e., how the computing power nodes will process the data.

[0018] S20. Identify multiple candidate computing nodes from the heterogeneous computing node pool that match the requested application type.

[0019] In essence, a heterogeneous computing node pool refers to a collection of computing resource nodes with different architectures and performance levels. These nodes provide diverse computing power support to meet the task execution needs of different application types and processing requirements. Different nodes may have different processor architectures, such as x86 or ARM. The server selects candidate computing nodes from the heterogeneous computing node pool based on the requesting application type. Candidate computing nodes are multiple computing resource nodes selected from the heterogeneous computing node pool that can meet the processing requirements of a specific computing power scheduling request. For example, when a computing power scheduling request needs to migrate data from an Intel processor, and the requesting application type is x86 and the data file to be processed has not undergone heterogeneous compilation, since the Hygon processor is also an x86 application system architecture, the Hygon processor's computing nodes can be used as candidate computing nodes.

[0020] In one embodiment, step S20, before determining multiple candidate computing nodes matching the requested application type from the heterogeneous computing power node pool, includes: S201. Perform multi-instance partitioning on the computing resources of the heterogeneous application system architecture to obtain multiple heterogeneous computing nodes; the heterogeneous application system architecture includes at least one x86 application system architecture and at least one ARM application system architecture. S202. Determine the set of all heterogeneous computing power nodes as a heterogeneous computing power node pool.

[0021] Understandably, a heterogeneous application system architecture is one that includes at least two processor systems using different architectures. For example... Figure 3 As shown, the heterogeneous application system architecture deployed in the Hadoop cluster consists of three types of heterogeneous processors: Intel processors, Kunpeng processors, and Hygon processors. The Intel and Hygon processors use the x86 architecture, while the Kunpeng processor uses the ARM architecture. Each application job corresponds to a computing power scheduling request. The application interface abstraction layer connects the resource requirements of application jobs with the unified scheduling of system resources, shielding the complexity of resource scheduling for heterogeneous computing power. The pluggable scheduling strategy engine can match the optimal heterogeneous computing power resources according to the resource requirements of application jobs. The resource scheduling process includes two stages: job selection and resource allocation. The scheduling strategy configuration can provide customized management of the pluggable scheduling strategy. Different application jobs have diverse computing power requirements; for example, x86 application jobs need to be allocated to computing power nodes after being partitioned using x86 architecture processors. Furthermore, before using the computing power resource scheduling method of this embodiment, a cross-cluster data computation comparison tool can be used to verify data consistency.

[0022] The server performs multi-instance partitioning of computing resources from Intel processors, Kunpeng processors, and Hygon processors, resulting in multiple heterogeneous computing nodes. A heterogeneous computing node refers to a computing node that accepts unified scheduling after the computing resources in a heterogeneous application system architecture have been partitioned. The server defines the set of all heterogeneous computing nodes as a heterogeneous computing node pool. At this point, computing scheduling requests corresponding to Intel processors can also be allocated to the heterogeneous computing nodes within the pool, which are derived from the partitioning of Kunpeng and Hygon processors.

[0023] This embodiment divides computing resources of different architectures into multiple instances to form a heterogeneous computing node pool. Data processing tasks can be allocated to the most suitable computing nodes based on the characteristics and requirements of different tasks. Within the heterogeneous computing node pool, computing nodes corresponding to different architectures can independently perform data migration and computation, improving stability. Furthermore, hardware with different architectures varies in cost; by constructing a heterogeneous computing node pool, a more cost-effective hardware combination can be selected based on actual needs.

[0024] In one embodiment, the heterogeneous computing power nodes include multiple first heterogeneous computing power nodes corresponding to the x86 application system architecture and multiple second heterogeneous computing power nodes corresponding to the ARM application system architecture; step S20, namely determining multiple candidate computing power nodes matching the requested application type from the heterogeneous computing power node pool, includes: S203. When the requested application type is a first application type corresponding to the x86 application system architecture, obtain the heterogeneity compatibility identifier of the computing power scheduling request; the heterogeneity compatibility identifier is determined based on the data file to be processed. S204. If the heterogeneous compatibility flag fails verification, then all idle first heterogeneous computing nodes in the heterogeneous computing node pool are identified as candidate computing nodes that match the requested application type. S205. If the heterogeneous compatibility identifier passes the verification, then all idle second heterogeneous computing nodes in the heterogeneous computing node pool are identified as candidate computing nodes that match the requested application type.

[0025] Understandably, the server categorizes the computing nodes in the heterogeneous computing node pool into two types based on their architecture: the first heterogeneous computing node and the second heterogeneous computing node. The first heterogeneous computing node refers to the computing nodes obtained by allocating computing resources from x86 architecture processors, i.e., the computing nodes corresponding to the x86 application system architecture. The second heterogeneous computing node refers to the computing nodes obtained by allocating computing resources from ARM architecture processors, i.e., the computing nodes corresponding to the ARM application system architecture. Similarly, the requesting application type is divided into a first application type corresponding to the x86 application system architecture and a second application type corresponding to the ARM application system architecture. The first application type indicates that the processor corresponding to the client application sending the computing power scheduling request uses the x86 architecture.

[0026] When determining candidate computing power nodes matching the requesting application type, the server first obtains the heterogeneity compatibility identifier of the computing power scheduling request, and then determines the candidate computing power nodes based on the heterogeneity compatibility identifier. The heterogeneity compatibility identifier is a label information used to characterize whether the computing power scheduling request can be allocated to a computing power node with a different architecture than the processor corresponding to the client application for data processing. The heterogeneity compatibility identifier is divided into verification failure and verification success. Verification failure indicates that the computing power scheduling request needs to be allocated to a computing power node with the same architecture as the processor corresponding to the client application. In this case, the server determines the first heterogeneous computing power node as a candidate computing power node matching the computing power scheduling request of the first application type; that is, data processing tasks of x86 application type need to be allocated to computing power nodes after the processor is partitioned using the x86 architecture. Verification success indicates that the computing power scheduling request needs to be allocated to a computing power node with a different architecture than the processor corresponding to the client application. In this case, the server determines the second heterogeneous computing power node as a candidate computing power node matching the computing power scheduling request of the first application type. In the heterogeneous computing node pool, each computing node is categorized into an idle state and a busy state. An idle state indicates that the node is eligible to be assigned new data processing tasks, while a busy state indicates that the node is not eligible. The server determines the idle or busy state based on the remaining resources in the node load parameters of each computing node in the heterogeneous computing node pool and a preset resource threshold. If the remaining resources in the node load parameters are greater than the preset resource threshold, the computing node is in an idle state.

[0027] This embodiment identifies matching candidate computing nodes from a heterogeneous computing node pool by recognizing heterogeneous compatibility identifiers, which can ensure application compatibility and stability and avoid operational errors caused by architectural differences.

[0028] In one embodiment, step S203, that is, before obtaining the heterogeneous compatibility identifier of the computing power scheduling request, includes: S2031. Obtain the instruction data packet from the data file to be processed, and determine whether there is heterogeneous compatible compiled code in the instruction data packet; S2032. If there is no heterogeneous compatible compiled code in the instruction data packet, the heterogeneous compatibility identifier of the computing power scheduling request is determined to be a verification failure. S2033. If the instruction data packet contains heterogeneous compatible compiled code, then the heterogeneous compatibility identifier of the computing power scheduling request is determined as verified.

[0029] Understandably, before obtaining the heterogeneity compatibility identifier of a computing power scheduling request, the server needs to perform compatibility verification on the data files to be processed in the computing power scheduling request and generate a heterogeneity compatibility identifier based on the verification result. The instruction data packet in the data files to be processed refers to a set of instructions combined to process a specific data file, used to guide the processor's computing power nodes to complete specific data processing tasks, such as data reading, computation, storage, or transmission. Heterogeneous compatibility compiled code refers to recompiled code that supports computing power nodes with different processor architectures for data processing. Some data files involve custom encryption algorithms. If a computing power scheduling request with an application type corresponding to the x86 application system architecture is assigned to an ARM architecture computing power node, an error will occur. Therefore, the custom encryption algorithm needs to be recompiled. Computing power scheduling requests that have not been recompiled will have their heterogeneity compatibility identifier fail verification and can only be assigned to x86 architecture computing power nodes. For example, the heterogeneous computing power node corresponding to the Hygon processor is identified as a candidate computing power node matching the computing power scheduling request corresponding to the Intel processor. After recompilation, the heterogeneity compatibility flag of the computing power scheduling request is verified as passed, and it can be allocated to the computing power node of the ARM architecture. The heterogeneous computing power node corresponding to the Kunpeng processor is identified as a candidate computing power node that matches the computing power scheduling request corresponding to the Intel processor.

[0030] This embodiment determines the heterogeneity compatibility identifier by examining the heterogeneous compatibility compilation code, which helps to reasonably allocate tasks on computing nodes with different architectures and improve the overall fault tolerance and stability of the system.

[0031] S30. Based on the data file to be processed and the processing requirement information, determine at least one data processing subtask.

[0032] Understandably, the server breaks down a computing power scheduling request into at least one data processing subtask based on the data file to be processed and the processing requirements. A data processing subtask refers to a task that processes a portion of the data in the data file to be processed according to the processing requirements of the computing power scheduling request. One computing power scheduling request corresponds to one data processing task, and one data processing task can be broken down into one or more data processing subtasks. For example, when the amount of data in the data file to be processed in a computing power scheduling request is too large, it is impossible to process the computing power scheduling request as a single data processing task; instead, it can be broken down into multiple data processing subtasks based on smaller data volumes.

[0033] In one embodiment, step S30, namely determining at least one data processing subtask based on the data file to be processed and the processing requirement information, includes: S301. Determine the task scenario fields and data processing indicators based on the processing requirement information; S302. The data file to be processed is split according to the task scenario field to obtain multiple sub-task data files; S303. Based on the data processing indicators and each of the sub-task data files, determine the data processing sub-task corresponding to the sub-task data file.

[0034] Understandably, the server parses the processing requirement information to obtain task scenario fields and data processing metrics. Task scenario fields are data attributes used to clearly define different processing scenarios and divide data ranges within the data processing tasks of the computing power scheduling request. Based on the task scenario fields, the server can extract a subset of data that meets specific conditions from the data file to be processed and generate data processing subtasks. Each data processing subtask is allocated a heterogeneous computing power node, enabling distributed data processing. A subtask data file refers to the subset of data that meets specific conditions extracted from the data file to be processed. Data processing metrics are processing result standards used to measure, evaluate, and describe the performance of data business during data processing, providing clear goals and directions for data processing. At this point, the data processing metrics correspond to the processing requirement information of the data processing subtasks. The processing requirement information of all data processing subtasks is consistent with the processing requirement information of the computing power scheduling request, only the data volume differs. For example, the data file to be processed in the computing power scheduling request includes insurance order data from multiple provinces, and the processing requirement is to calculate the average amount of insurance orders in each province. At this point, the task scenario field is "province", the data processing indicator is "average insurance order amount", the processing result of one data processing subtask is the average insurance order amount of one province, and the processing results of all data processing subtasks are summarized to generate the processing result of the computing power scheduling request.

[0035] This embodiment breaks down the computing power scheduling request into multiple data processing subtasks based on task scenario fields and data processing metrics. This helps to distribute each data processing subtask to computing power nodes with different architectures for parallel processing, thereby greatly shortening the total processing time of the request and improving data processing efficiency.

[0036] S40. Based on the node load parameters of the candidate computing power nodes, allocate a subtask target node for each data processing subtask from all the candidate computing power nodes, so that the subtask target node completes the allocated data processing subtask.

[0037] Understandably, each candidate computing node corresponds to its own node load parameters, which are parameters used to characterize the amount of unused resources on heterogeneous computing nodes. Each data processing subtask corresponds to its own subtask target node, which refers to the specific candidate computing node assigned to each data processing subtask according to specific rules and responsible for executing that subtask. Different data processing subtasks may correspond to the same or different subtask target nodes.

[0038] This embodiment, based on unified scheduling of heterogeneous computing resources, shields the complexity of resource scheduling for heterogeneous computing power, matches the optimal heterogeneous computing resources to the requested resource requirements, and improves data processing efficiency. In the scenario of domestic IT innovation upgrade, this embodiment can realize unified scheduling and management of three types of heterogeneous computing resources: domestic IT innovation environment (Hygon / Kunpeng processor) and non-domestic IT innovation environment (Intel processor), supporting gradual gray-scale domestic IT innovation transformation and upgrade.

[0039] In one embodiment, step S40, namely, allocating a subtask target node for each data processing subtask from all the candidate computing power nodes based on the node load parameters of the candidate computing power nodes, includes: S401. For each of the data processing subtasks, determine the subtask resource requirements of the data processing subtask based on the subtask data file of the data processing subtask. S402. Match the node load parameters of all candidate computing power nodes one by one according to the resource requirements of the subtask, and determine the matched candidate computing power nodes as the target nodes of the subtask.

[0040] Understandably, the server can determine the resource requirements of each data processing subtask based on its subtask data file. The subtask resource requirement refers to the amount of computing resources required to process the subtask data file. The server can determine the remaining resources of each candidate computing node based on its node load parameters. For each data processing subtask, the server compares its resource requirements with the remaining resources of all candidate computing nodes, filters out candidate computing nodes whose remaining resources are greater than the subtask resource requirement, sorts them, and selects the candidate computing node with the smallest remaining resources as the target node for the subtask.

[0041] This embodiment matches suitable candidate computing nodes according to the specific needs of each data processing subtask, ensuring that the allocated computing node resources match the task requirements, making resource allocation more scientific and reasonable, and improving the overall resource utilization rate.

[0042] In one embodiment, step S40, that is, after allocating a subtask target node for each data processing subtask from all the candidate computing power nodes according to the node load parameters of the candidate computing power nodes, includes: S403. The first data processing tool is invoked through the subtask target node to execute the data processing subtask, and the result of the first subtask is obtained. The second data processing tool is invoked to execute the data processing subtask, and the result of the second subtask is obtained. The first data processing tool is a Hive tool, and the second data processing tool is a Spark tool. S404. Determine the tool accuracy deviation based on the results of the first subtask and the second subtask. S405. When the tool accuracy deviation is greater than a preset deviation threshold, a tool selection request is sent to the client.

[0043] Understandably, for each subtask target node, when executing the data processing subtask, the server can choose to call either a first data processing tool or a second data processing tool. The first data processing tool is Hive, and the second data processing tool is Spark. The server calls both the first and second data processing tools to process the subtask data files within the data processing subtask, obtaining the first and second subtask results. The first subtask result is the result processed using Hive, and the second subtask result is the result processed using Spark. The server determines the tool precision deviation based on the first and second subtask results. Tool precision deviation refers to the difference in results obtained by processing the same data using different data processing tools. A preset deviation threshold is a pre-defined critical value used to evaluate the degree of influence of different data processing tools on the results. When the tool precision deviation is less than or equal to the preset deviation threshold, the server determines the final subtask result based on the first and second subtask results (e.g., randomly selecting a subtask result). When the tool precision deviation is greater than the preset deviation threshold, the server sends a tool selection request to the client. A tool selection request is used to ask the client user to actively select a specific data processing tool as needed. When the server receives the tool selection result corresponding to the tool selection request from the client, it determines the data processing tool specified in the tool selection result as the actual data processing tool and uses the selected data processing tool to generate the final subtask result.

[0044] This embodiment executes the same data processing subtask by calling different data processing tools and analyzes the accuracy deviations based on the results to cross-verify the accuracy of the results. Furthermore, client users can flexibly select suitable data processing tools based on their actual business needs and the accuracy deviations of the tools.

[0045] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0046] In one embodiment, a computing resource scheduling device is provided, which corresponds one-to-one with the computing resource scheduling method in the above embodiments. For example... Figure 4 As shown, the computing resource scheduling device includes a request acquisition module 10, an application type matching module 20, a subtask determination module 30, and a computing node allocation module 40. Detailed descriptions of each functional module are as follows: The request acquisition module 10 is used to acquire computing power scheduling requests and determine the request application type, data file to be processed, and processing requirement information based on the computing power scheduling requests. Application type matching module 20 is used to determine multiple candidate computing power nodes that match the requested application type from the heterogeneous computing power node pool; The subtask determination module 30 is used to determine at least one data processing subtask based on the data file to be processed and the processing requirement information; The computing power node allocation module 40 is used to allocate a subtask target node from all the candidate computing power nodes for each data processing subtask according to the node load parameters of the candidate computing power nodes, so that the subtask target node completes the allocated data processing subtask.

[0047] In one embodiment, the application type matching module 20 includes: The computing power node partitioning unit is used to perform multi-instance partitioning of computing power resources of heterogeneous application system architecture to obtain multiple heterogeneous computing power nodes; the heterogeneous application system architecture includes at least one x86 application system architecture and at least one ARM application system architecture. The node pool generation unit is used to determine the set of all the heterogeneous computing power nodes as a heterogeneous computing power node pool.

[0048] In one embodiment, the application type matching module 20 further includes: The heterogeneous compatibility identifier acquisition unit is used to acquire the heterogeneous compatibility identifier of the computing power scheduling request when the application type of the request is a first application type corresponding to the x86 application system architecture. The first candidate node determination unit is used to determine all idle first heterogeneous computing power nodes in the heterogeneous computing power node pool as candidate computing power nodes that match the requested application type if the heterogeneous compatibility identifier fails the verification. The second candidate node determination unit is used to determine all idle second heterogeneous computing nodes in the heterogeneous computing node pool as candidate computing nodes that match the requested application type if the heterogeneous compatibility identifier is verified as passed.

[0049] In one embodiment, the application type matching module 20 further includes: The instruction data packet analysis unit is used to obtain the instruction data packet in the data file to be processed and to determine whether there is heterogeneous compatible compiled code in the instruction data packet; The verification failure processing unit is used to determine the heterogeneity compatibility identifier of the computing power scheduling request as verification failure if there is no heterogeneous compatible compiled code in the instruction data packet; The verification processing unit is used to determine the heterogeneity compatibility identifier of the computing power scheduling request as verified if heterogeneous compatible compiled code exists in the instruction data packet.

[0050] In one embodiment, the subtask determination module 30 includes: The requirement information parsing unit is used to determine the task scenario fields and data processing indicators based on the processing requirement information. The file splitting processing unit is used to split the data file to be processed according to the task scenario field to obtain multiple sub-task data files; The subtask generation unit is used to determine the data processing subtask corresponding to the subtask data file based on the data processing indicators and each subtask data file.

[0051] In one embodiment, the computing node allocation module 40 includes: The subtask requirement determination unit is used to determine the subtask resource requirement of each data processing subtask based on the subtask data file of the data processing subtask. The target node determination unit is used to match the node load parameters of all the candidate computing power nodes one by one according to the resource requirements of the subtask, and determine the matched candidate computing power nodes as the target nodes of the subtask.

[0052] In one embodiment, the computing node allocation module 40 further includes: The tool invocation unit is used to invoke a first data processing tool to execute the data processing subtask through a subtask target node, and obtain the result of the first subtask; and to invoke a second data processing tool to execute the data processing subtask, and obtain the result of the second subtask; the first data processing tool is a Hive tool, and the second data processing tool is a Spark tool; The accuracy deviation determination unit is used to determine the tool accuracy deviation based on the results of the first subtask and the results of the second subtask. The selection request sending unit is used to send a tool selection request to the client when the tool accuracy deviation is greater than a preset deviation threshold.

[0053] Specific limitations regarding the computing resource scheduling device can be found in the limitations of the computing resource scheduling method above, and will not be repeated here. Each module in the aforementioned computing resource scheduling device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0054] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a readable storage medium and internal memory. The readable storage medium stores an operating system, computer-readable instructions, and a database. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The database stores data related to the computing resource scheduling method. The network interface communicates with external terminals via a network connection. When the computer-readable instructions are executed by the processor, a computing resource scheduling method is implemented. The readable storage medium provided in this embodiment includes both non-volatile and volatile readable storage media.

[0055] In one embodiment, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein the processor performs the following steps when executing the computer-readable instructions: Obtain a computing power scheduling request, and determine the requesting application type, the data file to be processed, and the processing requirements based on the computing power scheduling request; Multiple candidate computing nodes matching the requested application type are identified from the heterogeneous computing node pool; Based on the data file to be processed and the processing requirements information, at least one data processing subtask is determined; Based on the node load parameters of the candidate computing power nodes, a subtask target node is allocated from all the candidate computing power nodes for each data processing subtask, so that the subtask target node completes the allocated data processing subtask.

[0056] In one embodiment, one or more computer-readable storage media storing computer-readable instructions are provided. The readable storage media provided in this embodiment include non-volatile readable storage media and volatile readable storage media. The readable storage media stores computer-readable instructions, which, when executed by one or more processors, perform the following steps: Obtain a computing power scheduling request, and determine the requesting application type, the data file to be processed, and the processing requirements based on the computing power scheduling request; Multiple candidate computing nodes matching the requested application type are identified from the heterogeneous computing node pool; Based on the data file to be processed and the processing requirements information, at least one data processing subtask is determined; Based on the node load parameters of the candidate computing power nodes, a subtask target node is allocated from all the candidate computing power nodes for each data processing subtask, so that the subtask target node completes the allocated data processing subtask.

[0057] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When executed, these computer-readable instructions can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0058] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0059] The software tools or components not belonging to this company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use. The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A computing power resource scheduling method, characterized in that, include: Obtain a computing power scheduling request, and determine the requesting application type, the data file to be processed, and the processing requirements based on the computing power scheduling request; Multiple candidate computing nodes matching the requested application type are identified from the heterogeneous computing node pool; Based on the data file to be processed and the processing requirements information, at least one data processing subtask is determined; Based on the node load parameters of the candidate computing power nodes, a subtask target node is allocated from all the candidate computing power nodes for each data processing subtask, so that the subtask target node completes the allocated data processing subtask.

2. The computing resource scheduling method as described in claim 1, characterized in that, Before determining multiple candidate computing nodes matching the requested application type from the heterogeneous computing node pool, the process includes: The computing resources of the heterogeneous application system architecture are split into multiple instances to obtain multiple heterogeneous computing nodes; the heterogeneous application system architecture includes at least one x86 application system architecture and at least one ARM application system architecture. The set of all the heterogeneous computing power nodes is defined as the heterogeneous computing power node pool.

3. The computing resource scheduling method as described in claim 2, characterized in that, The heterogeneous computing power nodes include multiple first heterogeneous computing power nodes corresponding to the x86 application system architecture, and multiple second heterogeneous computing power nodes corresponding to the ARM application system architecture; The step of identifying multiple candidate computing nodes from the heterogeneous computing node pool that match the requested application type includes: When the requested application type is the first application type corresponding to the x86 application system architecture, obtain the heterogeneous compatibility identifier of the computing power scheduling request; If the heterogeneity compatibility flag fails verification, then all idle first heterogeneous computing nodes in the heterogeneous computing node pool are identified as candidate computing nodes that match the requested application type. If the heterogeneity compatibility flag passes verification, then all idle second heterogeneous computing nodes in the heterogeneous computing node pool are identified as candidate computing nodes that match the requested application type.

4. The computing resource scheduling method as described in claim 3, characterized in that, Before obtaining the heterogeneous compatibility identifier of the computing power scheduling request, the following steps are included: Obtain the instruction data packet from the data file to be processed, and determine whether there is heterogeneous compatible compiled code in the instruction data packet; If the instruction data packet does not contain heterogeneous compatible compiled code, the heterogeneous compatibility flag of the computing power scheduling request will be determined as failing verification; If the instruction data packet contains heterogeneous compatible compiled code, then the heterogeneous compatibility identifier of the computing power scheduling request is determined as verified.

5. The computing resource scheduling method as described in claim 1, characterized in that, The step of determining at least one data processing sub-task based on the data file to be processed and the processing requirement information includes: Based on the processing requirement information, determine the task scenario fields and data processing indicators; The data file to be processed is split according to the task scenario field to obtain multiple sub-task data files; Based on the data processing metrics and each of the subtask data files, determine the data processing subtask corresponding to that subtask data file.

6. The computing resource scheduling method as described in claim 1, characterized in that, The step of allocating a target node for each data processing subtask from all the candidate computing power nodes based on the node load parameters of the candidate computing power nodes includes: For each of the data processing subtasks, the resource requirements of the subtask are determined based on the subtask data file of the data processing subtask. Based on the resource requirements of the subtask, the node load parameters of all the candidate computing power nodes are matched one by one, and the matched candidate computing power nodes are determined as the target nodes of the subtask.

7. The computing resource scheduling method as described in claim 6, characterized in that, After allocating a target node for each data processing subtask from all the candidate computing power nodes based on the node load parameters of the candidate computing power nodes, the process includes: The data processing subtask is executed by calling a first data processing tool through the subtask target node to obtain the first subtask result, and by calling a second data processing tool to execute the data processing subtask to obtain the second subtask result; the first data processing tool is a Hive tool, and the second data processing tool is a Spark tool; Determine the tool accuracy deviation based on the results of the first subtask and the second subtask; When the tool's accuracy deviation exceeds a preset deviation threshold, a tool selection request is sent to the client.

8. A computing resource scheduling device, characterized in that, include: The request acquisition module is used to acquire computing power scheduling requests and determine the request application type, data file to be processed, and processing requirement information based on the computing power scheduling requests. The application type matching module is used to determine multiple candidate computing nodes that match the requested application type from the heterogeneous computing power node pool; The subtask determination module is used to determine at least one data processing subtask based on the data file to be processed and the processing requirement information; The computing power node allocation module is used to allocate a subtask target node from all the candidate computing power nodes for each data processing subtask according to the node load parameters of the candidate computing power nodes, so that the subtask target node completes the allocated data processing subtask.

9. A computer device comprising a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, characterized in that, When the processor executes the computer-readable instructions, it implements the computing resource scheduling method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing computer-readable instructions, characterized in that, When the computer-readable instructions are executed by one or more processors, the one or more processors cause the computing resource scheduling method as described in any one of claims 1 to 7 to be performed.