An OpenSAI intelligent network system and network algorithm matching method and system

By using the OpenSAI intelligent network system, resources are dynamically configured and scheduled, solving the long-term task scheduling problem on the edge computing platform, realizing efficient resource utilization and real-time and reliability of tasks, and improving system performance and response speed.

CN119576522BActive Publication Date: 2025-12-12FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411424236.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-12
Publication Date
2025-12-12
Estimated Expiration
2044-10-12

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively schedule long-term edge tasks on edge computing platforms, and existing algorithms are not suitable for edge platforms, leading to resource waste and performance degradation.

Method used

It provides the OpenSAI intelligent network system, which includes an application task module, a cluster management module, a distributed computing module, an adaptation and optimization module, a resource monitoring module, and a resource control module. By dynamically configuring and scheduling resources, it optimizes the matching of tasks and nodes, ensuring efficient resource utilization and task execution under appropriate conditions.

Benefits of technology

It improves the overall performance and response speed of the system, is suitable for scheduling long-term edge tasks, and ensures efficient use of resources and the real-time performance and reliability of tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119576522B_ABST
    Figure CN119576522B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of network system services, and particularly discloses an OpenSAI intelligent network system and a computing network matching method and system. The application monitors the states of network nodes in a computing node cluster to which a plurality of application tasks are offloaded through a resource monitoring module, obtains a plurality of node performance monitoring values, and sequentially judges whether the plurality of node performance monitoring values meet preset monitoring values. If the preset monitoring values are not met, a first resource scheduling instruction of corresponding demand resources of the nodes that do not meet the preset monitoring values is generated. After receiving the first resource scheduling instruction, a resource control module dynamically schedules computing storage network resources required by the corresponding nodes according to the first resource scheduling instruction. This method can ensure that resources are efficiently utilized, can also ensure that tasks are executed under suitable conditions, thereby improving the overall performance and response speed of the system, and is suitable for scheduling of long-term edge tasks.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of network system services, and in particular to an OpenSAI intelligent network system and an algorithm network matching method and system. BACKGROUND

[0002] OpenSAI is a system emphasizing openness, intelligence and networking characteristics, and specifically is an open platform or framework for intelligent devices or network systems, aiming to provide a series of standardized interfaces and services to facilitate the interconnection and intercommunication between different devices and services. At present, artificial intelligence applications are mainly deployed in the cloud. In the process of artificial intelligence application, it is usually necessary to deploy lightweight machines on the edge computing platform, and to more intelligently utilize real-time data of the Internet of Things under the application of lightweight machines. At the same time, with the development of open source container orchestration engines and cloud native technologies, the resources required by the edge computing platform are more, and the edge computing platform more needs cluster resource management and application lightweight deployment capabilities.

[0003] With the large consumption of edge cluster resources by big data real-time computing and machine learning tasks, effective scheduling is needed. However, existing researches mostly design scheduling based on deep learning and heuristic algorithms, but the number of training times is too large, which is not suitable for edge platforms, and the algorithms are mostly offline, which is not suitable for long-term edge task scheduling. Therefore, an OpenSAI intelligent network system and an algorithm network matching method and system are needed to solve the problem of unsuitability for long-term edge task scheduling. SUMMARY

[0004] The purpose of the present application is to provide an OpenSAI intelligent network system and an algorithm network matching method and system to solve the technical problems raised in the background art.

[0005] To achieve the above-mentioned purpose, the present application provides the following technical solutions:

[0006] An OpenSAI intelligent network system, comprising: an application task module, a cluster management module, a computing distributed module, an adaptation optimization module, a resource monitoring module and a resource control module;

[0007] The application task module is used for obtaining a plurality of application tasks, sorting the plurality of application tasks to obtain a task sorting list, and sequentially configuring interface parameter resources for the plurality of application tasks according to the task sorting list to obtain a plurality of application task configuration resources, and sending the plurality of application task configuration resources to the computing cluster management module;

[0008] The cluster management module is used for obtaining a current node state, and grouping a computing node cluster according to the current node state and the plurality of application task configuration resources;

[0009] The computing distribution module is configured to offload a plurality of application tasks to network nodes in a computing node cluster for execution, to obtain a plurality of execution tasks;

[0010] The adaptive optimization module is configured to dynamically configure the plurality of execution tasks based on resources provided by an elastic base, wherein the dynamic configuration comprises application task configuration, network node configuration, and network resource scheduling configuration;

[0011] The resource monitoring module is configured to monitor the state of the network nodes in the computing node cluster to which the plurality of application tasks are offloaded, to obtain a plurality of node performance monitoring values, and to sequentially determine whether the plurality of node performance monitoring values meet preset monitoring values;

[0012] If not, a first resource scheduling instruction corresponding to the required resources of the node that does not meet the preset monitoring values is generated.

[0013] The resource control module is configured to receive the first resource scheduling instruction and dynamically schedule the computing, storage, and network resources required by the corresponding node according to the first resource scheduling instruction.

[0014] Preferably, the adaptive optimization module is further configured to simultaneously optimize the application task configuration and the network bandwidth and network computing power resources, to obtain task-network bidirectional synchronization optimization information, and to pair the application tasks and the network nodes based on the task-network bidirectional synchronization optimization information.

[0015] The application further provides an algorithm-network matching method for executing the OpenSAI intelligent network system according to any one of claims 1-2, comprising:

[0016] Obtaining a plurality of application tasks, sorting the plurality of application tasks to obtain a task sorting list, and sequentially configuring interface parameter resources for the plurality of application tasks according to the task sorting list to obtain a plurality of application task configuration resources;

[0017] Obtaining a current node state and grouping a computing node cluster according to the current node state and the plurality of application task configuration resources;

[0018] Offloading the plurality of application tasks to network nodes in the computing node cluster for execution, to obtain a plurality of execution tasks;

[0019] Dynamically configuring the plurality of execution tasks based on resources provided by an elastic base, wherein the dynamic configuration comprises application task configuration, network node configuration, and network resource scheduling configuration;

[0020] Monitoring the state of the network nodes in the computing node cluster to which the plurality of application tasks are offloaded, to obtain a plurality of node performance monitoring values, and sequentially determining whether the plurality of node performance monitoring values meet preset monitoring values;

[0021] If not, the first resource scheduling instruction of the corresponding demand resource is generated for the node that does not meet the preset monitoring value;

[0022] According to the first resource scheduling instruction, the computing storage network resources required by the corresponding node are dynamically scheduled.

[0023] As a preferred, the step of configuring interface parameter resources for a plurality of application tasks according to the task scheduling list to obtain a plurality of application task configuration resources further comprises:

[0024] According to the application task configuration resource, the first CPU occupation resource amount, the first RAM occupation resource amount, the first GPU occupation resource amount and the first bandwidth occupation resource amount of the application task are obtained;

[0025] According to the first CPU occupation resource amount and the preset CPU occupation resource amount, the first CPU occupation ratio is calculated;

[0026] According to the first RAM occupation resource amount and the preset RAM occupation resource amount, the first RAM occupation ratio is calculated;

[0027] According to the first GPU occupation resource amount and the preset GPU occupation resource amount, the first GPU occupation ratio is calculated;

[0028] According to the first bandwidth occupation resource amount and the preset bandwidth occupation resource amount, the first bandwidth occupation ratio is calculated;

[0029] According to the first CPU occupation ratio, the first RAM occupation ratio, the first GPU occupation ratio and the first bandwidth occupation ratio, the application task resource occupation score is calculated, and the calculation formula is:

[0030] ;

[0031] Wherein, application task resource occupation score, first CPU occupation ratio, first RAM occupation ratio, first GPU occupation ratio, first bandwidth occupation ratio;

[0032] According to the above steps, a plurality of application task resource occupation scores corresponding to a plurality of application task configuration resources are calculated.

[0033] As a preferred, the step of monitoring the state of the network node in the computing node cluster according to the plurality of application tasks further comprises:

[0034] According to the node performance monitoring value, a second CPU resource occupation amount, a second RAM resource occupation amount, a second GPU resource occupation amount and a second bandwidth resource occupation amount of the node are obtained;

[0035] According to the second CPU resource occupation amount and a preset CPU resource occupation amount, a second CPU occupation ratio is calculated;

[0036] According to the second RAM resource occupation amount and a preset RAM resource occupation amount, a second RAM occupation ratio is calculated;

[0037] According to the second GPU resource occupation amount and a preset GPU resource occupation amount, a second GPU occupation ratio is calculated;

[0038] According to the second bandwidth resource occupation amount and a preset bandwidth resource occupation amount, a second bandwidth occupation ratio is calculated;

[0039] According to the second CPU occupation ratio, the second RAM occupation ratio and the second bandwidth occupation ratio, a node resource occupation amount score is calculated, and the calculation formula is:

[0040] ;

[0041] Wherein, represents the node resource occupation amount score, represents the second CPU occupation ratio, represents the second RAM occupation ratio, represents the first GPU occupation ratio, represents the second bandwidth occupation ratio;

[0042] According to the above calculation steps, a plurality of node resource occupation amount scores corresponding to a plurality of node performance monitoring values are sequentially calculated.

[0043] Preferably, the step of dynamically scheduling the computing and storage network resources required by the corresponding node according to the first resource scheduling instruction further comprises:

[0044] The application task resource occupation amount score is iteratively matched with a plurality of node resource occupation amount scores based on a similarity model to obtain a resource similarity value, wherein the function of the similarity model is:

[0045] ;

[0046] Wherein, represents the resource similarity value, represents the application task resource occupation amount score, represents the node resource occupation amount score, and n represents the serial number of the node resource occupation amount score, n = 1, 2, 3... n.

[0047] The application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method when executing the computer program.

[0048] The application further provides a computer readable storage medium, which stores a computer program, and the computer program implements the steps of the method when executed by a processor.

[0049] The application has the following beneficial effects: the application first acquires multiple application tasks by applying a task module, and prepares the priority and configuration parameters of the application tasks, then a cluster management module is used to acquire the current node state, and a computing node cluster is formed, meanwhile, multiple application tasks are unloaded to the network nodes in the computing node cluster for execution to obtain multiple execution tasks, then an adaptive optimization module dynamically configures the multiple execution tasks based on the resources provided by the elastic base, finally, a resource monitoring module is used to monitor the state of the multiple application tasks unloaded to the network nodes in the computing node cluster to obtain multiple node performance monitoring values, and whether the multiple node performance monitoring values meet preset monitoring values is sequentially judged, if not, a first resource scheduling instruction of the corresponding required resource is generated for the node that does not meet the preset monitoring value, and a resource control module dynamically schedules the computing network resources required by the corresponding node according to the first resource scheduling instruction after receiving the first resource scheduling instruction, so that the system can dynamically match the application task configuration resources and the node resources for matching, and then the optimal resource allocation decision is made according to the real-time data, this method can ensure efficient use of resources, and also ensure that the tasks are executed under suitable conditions, thereby improving the overall performance and response speed of the system, and is suitable for scheduling of long-term edge tasks. BRIEF DESCRIPTION OF DRAWINGS

[0050] Figure 1 The method flowchart of an embodiment of the application.

[0051] Figure 2 The system structure schematic diagram of an embodiment of the application.

[0052] Figure 3 The internal structure schematic diagram of a computer device of an embodiment of the application.

[0053] Figure 4 The example explanation schematic diagram of an embodiment of the application.

[0054] The purpose implementation, functional features and advantages of the application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0055] It should be understood that the specific embodiments described herein are merely intended to explain the application, and are not intended to limit the application.

[0056] As Figures 1-3 shown, the present application provides an OpenSAI intelligent network system, comprising: an application task module, a cluster management module, a computing distributed module, an adaptive optimization module, a resource monitoring module and a resource control module;

[0057] The application task module is used for obtaining a plurality of application tasks, sorting the plurality of application tasks to obtain a task sorting list, and sequentially configuring interface parameter resources for the plurality of application tasks according to the task sorting list to obtain a plurality of application task configuration resources, and sending the plurality of application task configuration resources to the computing cluster management module;

[0058] The cluster management module is used for obtaining a current node state, and grouping a computing node cluster according to the current node state and the plurality of application task configuration resources;

[0059] The computing distributed module is used for unloading the plurality of application tasks to network nodes in the computing node cluster for execution to obtain a plurality of execution tasks;

[0060] The adaptive optimization module is used for dynamically configuring the plurality of execution tasks based on resources provided by an elastic base, wherein the dynamic configuration includes application task configuration, application task and network node configuration, and network resource scheduling configuration;

[0061] The resource monitoring module is used for monitoring states of the plurality of application tasks unloaded to the network nodes in the computing node cluster to obtain a plurality of node performance monitoring values, and sequentially judging whether the plurality of node performance monitoring values meet preset monitoring values;

[0062] If not, a first resource scheduling instruction of a corresponding demand resource is generated for a node that does not meet the preset monitoring value;

[0063] The resource control module is used for receiving the first resource scheduling instruction, and dynamically scheduling computing storage network resources required by the corresponding node according to the first resource scheduling instruction.

[0064] The application obtains multiple application tasks by applying a task module, sorts the multiple application tasks to obtain a task sorting list, and sequentially configures interface parameter resources for the multiple application tasks according to the task sorting list to obtain multiple application task configuration resources, and sends the multiple application task configuration resources to a computing cluster management module, so that corresponding task instances can be quickly created according to the creation of each application task, and the priority of the task and the configuration interface and other parameters are set to provide a basis for subsequent application task calculation, and then the current node state is obtained by the cluster management module, and a computing node cluster is formed according to the current node state and the multiple application task configuration resources, so that the state information (such as CPU usage, memory occupation, network condition, etc.) of all available nodes is collected by the cluster management module, and the computing resource availability is evaluated according to the node state to prepare for task allocation, and then the multiple application tasks are unloaded to the network nodes in the computing node cluster for execution by the computing distributed module to obtain multiple execution tasks, and each node cooperates with each other through the communication module, wherein the computing distributed module is also used to analyze the task complexity, determine whether distributed computing is needed, and allocate specific parameters and resource requirements of the execution task to each node, and then the adaptive optimization module dynamically configures the multiple execution tasks based on the resources provided by the elastic base, wherein the dynamic configuration includes application task configuration, application task and network node configuration, and network resource scheduling configuration, so that the adaptive optimization module can optimize the application and the network at the same time, and the real-time performance and reliability of task execution can be ensured, and finally, the state of the multiple application tasks unloaded to the network nodes in the computing node cluster is monitored by the resource monitoring module to obtain multiple node performance monitoring values, and whether the multiple node performance monitoring values meet the preset monitoring value is sequentially judged, if not, a first resource scheduling instruction of the corresponding required resource is generated for the node that does not meet the preset monitoring value, and the resource control module is used to receive the first resource scheduling instruction and dynamically schedule the required computing, storage and network resources of the corresponding node, so that the system can dynamically match the application task configuration resources and the node resources for matching, and then the optimal resource allocation decision is made according to the real-time data, this method can ensure efficient use of resources, and also ensure that the task is executed under suitable conditions, thereby improving the overall performance and response speed of the system, and is suitable for long-term edge task scheduling.

[0065] In one embodiment, the adaptive optimization module is also used to simultaneously optimize the application task configuration and the network bandwidth and network computing power resources to obtain task-network bidirectional synchronization optimization information, and pair the application task and the network node based on the task-network bidirectional synchronization optimization information.

[0066] The application obtains task-network bidirectional synchronization optimization information by simultaneously optimizing application task configuration and network bandwidth and network computing power resources through the adaptive optimization module, which can optimize the execution environment, thereby ensuring the real-time performance and reliability of task execution, and traverses and pairs application resource configuration and multiple node performance monitoring values based on the task-network bidirectional synchronization optimization information, so that the application can be calculated and run on the most suitable node after the traversal and pairing of application resource configuration and multiple node performance monitoring values.

[0067] The application further provides an algorithm-network matching method for executing the OpenSAI intelligent network system according to any one of claims 1-2, comprising:

[0068] S1, obtaining multiple application tasks, sorting the multiple application tasks to obtain a task sorting list, and sequentially configuring interface parameter resources for the multiple application tasks according to the task sorting list to obtain multiple application task configuration resources;

[0069] S2, obtaining a current node state, and constructing a computing node cluster according to the current node state and the multiple application task configuration resources;

[0070] S3, unloading the multiple application tasks to network nodes in the computing node cluster for execution to obtain multiple execution tasks;

[0071] S4, dynamically configuring the multiple execution tasks based on resources provided by an elastic base, wherein the dynamic configuration includes application task configuration, application task and network node configuration, and network resource scheduling configuration;

[0072] S5, monitoring the state of the multiple application tasks unloaded to the network nodes in the computing node cluster to obtain multiple node performance monitoring values, and sequentially judging whether the multiple node performance monitoring values meet preset monitoring values;

[0073] If not, a first resource scheduling instruction of corresponding demand resources is generated for the node that does not meet the preset monitoring values;

[0074] S6, dynamically scheduling algorithm storage network resources required by the corresponding node according to the first resource scheduling instruction.

[0075] As described in steps S1-S6 above, the application first acquires a plurality of application tasks, sorts the plurality of application tasks to obtain a task sorting list, and sequentially configures interface parameter resources for the plurality of application tasks according to the task sorting list to obtain a plurality of application task configuration resources. Then, the current node state is acquired, and a computing node cluster is constructed according to the current node state and the plurality of application task configuration resources. Then, the plurality of application tasks are unloaded to the network nodes in the computing node cluster for execution to obtain a plurality of execution tasks. Then, the plurality of execution tasks are dynamically configured based on the resources provided by the elastic base, wherein the dynamic configuration includes application task configuration, application task and network node configuration, and network resource scheduling configuration. Finally, the state of the plurality of application tasks unloaded to the network nodes in the computing node cluster is monitored to obtain a plurality of node performance monitoring values, and it is sequentially determined whether the plurality of node performance monitoring values meet a preset monitoring value. In this way, the system can dynamically match the application task configuration resources and the node resources for matching, and then make the optimal resource allocation decision according to the real-time data, thereby improving the overall performance and response speed of the system, and being suitable for long-term edge task scheduling, for example, the actual application of the method (as shown in Figure 4 The above Figure 4 The application is only used for illustration and is not the only reference.

[0076] In one embodiment, the step S1 of sequentially configuring interface parameter resources for the plurality of application tasks according to the task sorting list to obtain a plurality of application task configuration resources further includes:

[0077] S101, acquiring a first CPU occupation resource amount, a first RAM occupation resource amount, a first GPU occupation resource amount, and a first bandwidth occupation resource amount of the application task according to the application task configuration resources;

[0078] S102, calculating a first CPU occupation ratio according to the first CPU occupation resource amount and a preset CPU occupation resource amount;

[0079] S103, calculating a first RAM occupation ratio according to the first RAM occupation resource amount and a preset RAM occupation resource amount;

[0080] S104, calculating a first GPU occupation ratio according to the first GPU occupation resource amount and a preset GPU occupation resource amount;

[0081] S105, calculating a first bandwidth occupation ratio according to the first bandwidth occupation resource amount and a preset bandwidth occupation resource amount;

[0082] S106, calculating an application task resource occupation amount score according to the first CPU occupation ratio, the first RAM occupation ratio, the first GPU occupation ratio, and the first bandwidth occupation ratio, wherein the calculation formula is:

[0083] ;

[0084] wherein, represents an application task resource occupation score, represents a first CPU occupation ratio, represents a first RAM occupation ratio, represents a first GPU occupation ratio, represents a first bandwidth occupation ratio;

[0085] S107, according to the above steps, a plurality of application task resource occupation scores corresponding to a plurality of application task configuration resources are calculated in turn.

[0086] As described above in steps S102-S107, the application first obtains the first CPU occupation resource amount, the first RAM occupation resource amount, the first GPU occupation resource amount and the first bandwidth occupation resource amount of the application task according to the application task configuration resource, then calculates the first CPU occupation ratio according to the first CPU occupation resource amount and the preset CPU occupation resource amount, that is, the proportion of the actually used CPU resource to the preset CPU resource, so as to understand the CPU resource usage of the device, so as to further analyze and adjust the resource allocation strategy, then calculates the first RAM occupation ratio according to the first RAM occupation resource amount and the preset RAM occupation resource amount, that is, the RAM resource usage of the device, so as to further analyze and adjust the resource allocation strategy, then calculates the first bandwidth occupation ratio according to the first bandwidth occupation resource amount and the preset bandwidth occupation resource amount, and finally, the application task resource occupation score is calculated according to the first CPU occupation ratio, the first RAM occupation ratio and the first bandwidth occupation ratio, and then a plurality of application task resource occupation scores corresponding to a plurality of application task configuration resources are calculated in turn according to the above steps, for example, the first CPU occupation resource amount is 1000, the preset CPU occupation resource amount is 2000, the first GPU occupation resource amount is 1000, the preset GPU occupation resource amount is 2000, the first RAM occupation resource amount is 512, the preset RAM occupation resource amount is 1024, the first bandwidth occupation ratio is 100, and the preset bandwidth occupation resource amount is 200, the first CPU occupation ratio = 1000 / 2000 = 0.5, the first RAM occupation ratio = 512 / 1024 = 0.5, the first GPU occupation ratio = 1000 / 2000 = 0.5, the first bandwidth occupation ratio = 100 / 200 = 0.5, and the application task resource occupation score = (0.5+0.5+0.5+0.5) / 4 = 0.5. In this way, the resource usage of the device is evaluated by the application task resource occupation score, which is convenient for managers to optimize and make decisions, and provides an important basis for resource scheduling. The above example is only for illustration and does not constitute a unique reference.

[0087] In one embodiment, the step S5 of monitoring the state of the plurality of application tasks offloaded to the network nodes in the computing node cluster to obtain a plurality of node performance monitoring values further comprises:

[0088] S501, obtaining a second CPU occupied resource amount, a second RAM occupied resource amount, a second GPU occupied resource amount and a second bandwidth occupied resource amount of the node according to the node performance monitoring value;

[0089] S502, calculating a second CPU occupied ratio according to the second CPU occupied resource amount and a preset CPU occupied resource amount;

[0090] S503, calculating a second RAM occupied ratio according to the second RAM occupied resource amount and a preset RAM occupied resource amount;

[0091] S504, calculating a second GPU occupied ratio according to the second GPU occupied resource amount and a preset GPU occupied resource amount;

[0092] S505, calculating a second bandwidth occupied ratio according to the second bandwidth occupied resource amount and a preset bandwidth occupied resource amount;

[0093] S506, calculating a node resource occupation score according to the second CPU occupied ratio, the second RAM occupied ratio and the second bandwidth occupied ratio, wherein the calculation formula is:

[0094]

[0095] wherein, represents the node resource occupation score, represents the second CPU occupied ratio, represents the second RAM occupied ratio, represents the first GPU occupied ratio, represents the second bandwidth occupied ratio;

[0096] S507, sequentially calculating a plurality of node resource occupation scores corresponding to a plurality of node performance monitoring values according to the above calculation steps.

[0097] ​As described in steps S501-S507, the application first obtains the second CPU resource occupation amount, the second RAM resource occupation amount, the second GPU resource occupation amount and the second bandwidth resource occupation amount of the node according to the node performance monitoring value, then calculates the second CPU occupation ratio according to the second CPU resource occupation amount and the preset CPU resource occupation amount, then calculates the second RAM occupation ratio according to the second RAM resource occupation amount and the preset RAM resource occupation amount, then calculates the second GPU occupation ratio according to the second GPU resource occupation amount and the preset GPU resource occupation amount, then calculates the second bandwidth occupation ratio according to the second bandwidth resource occupation amount and the preset bandwidth resource occupation amount, and finally calculates the node resource occupation score according to the second CPU occupation ratio, the second RAM occupation ratio, the second GPU occupation ratio and the second bandwidth occupation ratio. In this way, the node resource occupation score facilitates the system to make resource optimization and decision, and further provides an important basis for subsequent resource scheduling.

[0098] In one embodiment, the step S6 of dynamically scheduling the computing and storage network resources required by the corresponding node according to the first resource scheduling instruction further includes:

[0099] S601, the application task resource occupation score is matched with a plurality of node resource occupation scores based on a similarity model to obtain a resource similarity value, wherein the function of the similarity model is:

[0100] ;

[0101] wherein, the resource similarity value is represented by, the application task resource occupation score is represented by, the node resource occupation score is represented by, and n represents the serial number of the node resource occupation score, n=1, 2, 3...n.

[0102] As described in step S601, the application matches the application task resource occupation score with a plurality of node resource occupation scores based on a similarity model to obtain a resource similarity value. When the resource similarity value tends to 1, it means that the two are more matched. In this way, the system can quickly match the task resources and the node resources, and make the optimal resource allocation decision according to the real-time data. This method can ensure efficient use of resources, and also ensure that the task is executed under suitable conditions, thereby improving the overall performance and response speed of the system, and is suitable for long-term edge task scheduling.

[0103] The application further provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of the OpenSAI intelligent network system and computing network matching method.

[0104] Those of ordinary skill in the art understand that all or part of the processes in the above-mentioned embodiments can be implemented by a computer program instructing relevant hardware, and the computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, the processes of the above-mentioned embodiments can be included. Any reference to memory, storage, database, or other medium provided by the present application and used in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0105] It should be noted that in this document, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusions, so that a process, device, article, or method including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or includes elements inherent to such a process, device, article, or method. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, device, article, or method including the element.

[0106] The above description is only the preferred embodiments of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation using the content of the present application specification and drawings, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.

Claims

1. An OpenSAI intelligent network system, characterized in that, include: Application task module, cluster management module, distributed computing module, adaptation and optimization module, resource monitoring module, and resource control module; The application task module is used to obtain multiple application tasks, sort the multiple application tasks to obtain a task sorting list, configure the interface parameter resources of the multiple application tasks in sequence according to the task sorting list to obtain multiple application task configuration resources, and send the multiple application task configuration resources to the computing cluster management module. The cluster management module is used to obtain the current node status and configure resources to form a computing node cluster based on the current node status and multiple application tasks; The distributed computing module is used to offload multiple application tasks to network nodes in a compute node cluster for execution, resulting in multiple execution tasks. The adaptation and optimization module is used to dynamically configure multiple execution tasks based on the resources provided by the elastic base. The dynamic configuration includes application task configuration, application task and network node configuration, and network resource scheduling configuration. The resource monitoring module is used to monitor the status of network nodes that are offloaded to the computing node cluster for multiple application tasks, obtain multiple node performance monitoring values, and pass the multiple node performance monitoring values ​​to the adaptation and optimization module to determine whether the multiple node performance monitoring values ​​meet the preset monitoring values ​​in turn. If the preset monitoring value is not met, a first resource scheduling instruction for the corresponding required resources will be generated for the node that does not meet the preset monitoring value. The resource control module is used to receive the first resource scheduling instruction and dynamically schedule the computing, storage, and network resources required by the corresponding node according to the first resource scheduling instruction.

2. The OpenSAI intelligent network system according to claim 1, characterized in that, The adaptation and optimization module is also used to simultaneously optimize application task configuration and network bandwidth and network computing resources to obtain task-network bidirectional synchronization optimization information, and to pair application tasks and network nodes based on the task-network bidirectional synchronization optimization information.

3. A network matching method for executing the OpenSAI intelligent network system as described in any one of claims 1-2, characterized in that, include: Multiple application tasks are obtained, the multiple application tasks are sorted to obtain a task sorting list, and the interface parameter resources of the multiple application tasks are configured in turn according to the task sorting list to obtain multiple application task configuration resources. Obtain the current node status, and configure resources to build a computing node cluster based on the current node status and multiple application tasks; Multiple application tasks are offloaded to network nodes in the compute node cluster for execution, resulting in multiple execution tasks. Multiple execution tasks are dynamically configured based on the resources provided by the elastic base, wherein the dynamic configuration includes application task configuration, application task and network node configuration, and network resource scheduling configuration. The status of multiple application tasks unloaded to network nodes in the computing node cluster is monitored to obtain multiple node performance monitoring values, and it is determined in turn whether the multiple node performance monitoring values ​​meet the preset monitoring values. If the preset monitoring value is not met, a first resource scheduling instruction for the corresponding required resources will be generated for the node that does not meet the preset monitoring value. The computing, storage, and network resources required by the corresponding node are dynamically scheduled according to the first resource scheduling instruction.

4. The network matching method according to claim 3, characterized in that, The step of configuring interface parameters and resources for multiple application tasks sequentially according to the task sorting list to obtain configuration resources for multiple application tasks further includes: Based on the application task configuration resources, obtain the first CPU resource usage, first RAM resource usage, first GPU resource usage, and first bandwidth resource usage of the application task; The first CPU utilization ratio is calculated based on the first CPU resource usage and the preset CPU resource usage. The first RAM usage ratio is calculated based on the first RAM usage amount and the preset RAM usage amount; The first GPU utilization ratio is calculated based on the first GPU resource usage and the preset GPU resource usage. The first bandwidth utilization ratio is calculated based on the first bandwidth utilization amount and the preset bandwidth utilization amount; The application task resource usage score is calculated based on the first CPU usage ratio, the first RAM usage ratio, the first GPU usage ratio, and the first bandwidth usage ratio, wherein the calculation formula is: ; in, This indicates the application's resource consumption score. This indicates the first CPU utilization rate. Indicates the first RAM usage ratio. Indicates the first GPU utilization rate. Indicates the first bandwidth utilization ratio; Based on the above steps, the resource usage scores of multiple application tasks corresponding to the configuration resources of multiple application tasks are calculated sequentially.

5. The network matching method according to claim 3, characterized in that, The step of monitoring the status of multiple application tasks unloaded to the network nodes in the computing node cluster to obtain multiple node performance monitoring values ​​further includes: The node's second CPU resource usage, second RAM resource usage, second GPU resource usage, and second bandwidth resource usage are obtained based on the node performance monitoring values. The second CPU utilization ratio is calculated based on the second CPU resource usage and the preset CPU resource usage. The second RAM usage ratio is calculated based on the second RAM usage amount and the preset RAM usage amount; The second GPU utilization ratio is calculated based on the second GPU resource usage and the preset GPU resource usage. The second bandwidth utilization ratio is calculated based on the second bandwidth utilization amount and the preset bandwidth utilization amount. The node resource utilization score is calculated based on the second CPU utilization ratio, the second RAM utilization ratio, and the second bandwidth utilization ratio. The calculation formula is as follows: ; in, This indicates a score representing the node's resource usage. Indicates the second CPU utilization rate. Indicates the second RAM usage ratio. Indicates the first GPU utilization rate. Indicates the second bandwidth utilization ratio; Based on the above calculation steps, the resource usage scores of multiple nodes corresponding to the multiple node performance monitoring values ​​are calculated sequentially.

6. The network matching method according to claim 3, characterized in that, The step of dynamically scheduling the computing, storage, and network resources required by the corresponding node according to the first resource scheduling instruction further includes: The application task resource usage score is matched with the resource usage scores of multiple nodes based on a similarity model to obtain a resource similarity value. The function of the similarity model is: ; in, Indicates resource similarity value, This indicates the application's resource consumption score. This represents the node resource usage score, where n represents the index of the node resource usage score, n = 1, 2, 3...n.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 3 to 6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 3 to 6.

Citation Information

Patent Citations

  • System resource scheduling method and device, machine readable medium and system

    CN111158879A

  • Multidimensional resource scheduling method in kubernetes cluster architecture system

    US20210365290A1