Multi-modal computing power arrangement method and system

By obtaining multimodal service flow data, extracting features and calculating the capacity value, and dynamically adjusting computing power resources based on node utility functions and priority weights, the problem that traditional scheduling methods cannot adapt to multimodal service flow is solved, and efficient and balanced multimodal data processing and resource allocation are achieved.

CN120508394APending Publication Date: 2025-08-19ZHONGSHAN POWER SUPPLY BUREAU OF GUANGDONG POWER GRID
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510683610.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

Traditional computing resource scheduling methods cannot adapt to the dynamic changes in data types and quantities in multimodal service flows, resulting in the inability to meet the needs of multimodal service flows. Especially when multiple data types need to be processed simultaneously and cross-modal information fusion and processing are realized, the limitations of static scheduling and single data type processing cannot meet the needs.

Method used

By obtaining multimodal service flow data, extracting business flow characteristics, calculating information capacity values using the information capacity calculation model, determining the initial computing power value based on the node utility function, and dynamically adjusting the target computing power value according to the priority weight, and using Kubernetes tools to allocate computing power resources to achieve efficient and real-time efficiency of multimodal data processing.

Benefits of technology

It realizes the efficiency of multimodal data processing and the balance of resource allocation, can adapt to the dynamic changes of multimodal service flow, and ensures that the fusion and processing of cross-modal information meet real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508394A_ABST
    Figure CN120508394A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of multi-modal computing power arrangement, and discloses a multi-modal computing power arrangement method and system, and the method comprises the steps: obtaining multi-modal service flow data, extracting the service flow features of the multi-modal service flow data, inputting a preset information capacity calculation model, and carrying out the calculation to obtain an information capacity value needed by each service flow feature; based on a preset node utility function and the information capacity value, determining an initial computing power value required for processing each service; mapping the service flow to a corresponding container according to the initial computing power value; and dynamically adjusting the initial computing power value to generate a target computing power value according to the priority weight of the service flow corresponding to the container, and allocating computing power resources based on the target computing power value. According to the invention, through three core mechanisms of multi-modal feature fusion analysis, information-capacity-driven dynamic computing power prediction and priority-perceived resource scheduling, high efficiency and real-time performance of multi-modal data processing and balance of resource allocation are realized. The technical problem that the requirements of multi-mode service flow cannot be met due to limitation of static scheduling and single data type processing in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal computing power orchestration, and in particular to a multimodal computing power orchestration method and system. Background Art

[0002] The Internet of Things (IoT) connects various physical devices to the internet, enabling real-time data collection, exchange, and analysis. The processing and analysis of this data is driven by computing power. With the surge in the number of IoT devices, the explosive growth in data volumes, and the development of diverse data modalities, the demand for computing power is also increasing. Multimodal data includes images, videos, text, voice, and other data types. These data types are prevalent in IoT applications and require efficient means of perception and comprehensive understanding. With the emergence of new intelligent applications, system latency performance has become a critical consideration. Users expect real-time or near-real-time processing results, which places higher demands on the scheduling and allocation of computing resources.

[0003] However, traditional computing resource scheduling methods allocate resources based on preset rules and policies. In multimodal traffic flows, the data types and quantities change dynamically, and static scheduling methods cannot adapt to these changes. Traditional computing resource scheduling methods often process a single data type, such as images or text. Multimodal traffic flows require simultaneous processing of multiple data types and the fusion and processing of cross-modal information. Due to the limitations of static scheduling and single-data type processing, they cannot meet the needs of multimodal traffic flows. Summary of the Invention

[0004] This invention provides a multimodal computing power orchestration method and system that addresses the technical problem that existing technologies, due to the limitations of static scheduling and single data type processing, cannot meet the needs of multimodal business flows due to the need to simultaneously process multiple data types and achieve cross-modal information fusion and processing.

[0005] A first aspect of the present invention provides a multimodal computing power orchestration method, comprising:

[0006] Acquire multiple modal business flow data, and extract business flow features of each of the modal business flow data;

[0007] Inputting each of the service flow characteristics into a preset credit capacity calculation model for calculation to obtain the credit capacity value required by each of the service flow characteristics;

[0008] Determining the initial computing power required for the business processing corresponding to each of the business flow characteristics based on the preset node utility function and each of the credit values;

[0009] Mapping the business flows corresponding to the business processes to corresponding containers based on the initial computing power required for the business processes;

[0010] According to the priority weight of the business flow corresponding to the container, the initial computing power value required for the business processing corresponding to each business flow is adjusted to generate a target computing power value, and computing power resources are allocated according to the target computing power value.

[0011] Optionally, the acquiring of multiple modal business flow data and extracting business flow features of each of the modal business flow data includes:

[0012] Collecting multiple modal business flow data from the business interfaces of different sensing devices;

[0013] Assigning a priority tag to each of the modal service flow data according to the service type, and storing each of the modal service flow data in a priority queue according to the priority tag;

[0014] Extract the business flow characteristics of each modal business flow data in the priority queue.

[0015] Optionally, inputting each of the service flow characteristics into a preset credit capacity calculation model for calculation to obtain a credit capacity value required by each of the service flow characteristics includes:

[0016] Input the feature data corresponding to each of the business flow features and the business event attention matrix into a preset information quantity equation to obtain the information quantity of each modal event;

[0017] The information volume of each modal event, the maximum value of the credit capacity and the characteristic data dimension are input into a preset credit capacity calculation model to generate the credit capacity value required by each business flow feature.

[0018] Optionally, it also includes:

[0019] The variance of each node's resource utilization is calculated using the node resource utilization and the overall mean of the node resources;

[0020] Determining a node utility function based on the variance of resource utilization of each of the nodes;

[0021] Set constraints based on the capacity of task offloading and resource allocation;

[0022] Based on cooperative game theory, the objective function of multi-resource load balancing scheduling is set.

[0023] Optionally, determining the initial computing power required for the service processing corresponding to each of the service flow characteristics based on a preset node utility function and each of the credit values includes:

[0024] Classify the credit capacity values corresponding to the business flow characteristics and establish a mapping table of basic computing power requirements for each mode;

[0025] Determine the resource allocation correction coefficient based on the preset node utility function;

[0026] Based on the basic computing power requirements and resource allocation correction coefficients corresponding to the basic computing power requirements mapping tables of each modality, the initial computing power value required for the business processing corresponding to each business flow feature is determined.

[0027] Optionally, adjusting the initial computing power required for business processing corresponding to each business flow according to the priority weight of the business flow corresponding to the container to generate a target computing power value, and allocating computing power resources according to the target computing power value includes:

[0028] Calculating the priority weight of the service flow corresponding to each container according to the normalized result of the credit capacity value of each container, the adjustment factor and the service quality weight;

[0029] Determine a normalized error of the actual delay based on a preset delay threshold of the service flow and the actual delay of the service flow;

[0030] Based on a preset load impact factor, the normalized error of the actual delay, the priority weight, and the sensitivity coefficient, an initial computing power value required for service processing corresponding to each of the service flows is adjusted to generate a target computing power value;

[0031] Allocate computing power resources according to the target computing power value.

[0032] Optionally, it also includes:

[0033] Real-time collection of actual node resource utilization and processing delay;

[0034] Calculating a computing power prediction error rate using the actual resource utilization of the node and the processing delay;

[0035] Determining whether the computing power prediction error rate is greater than a preset error rate threshold;

[0036] If so, adjust the business event attention matrix to generate a new business event attention matrix, and jump to execute the step of inputting the feature data corresponding to each of the business flow features and the business event attention matrix into the preset information quantity equation to obtain the information quantity of each modal event.

[0037] A second aspect of the present invention provides a multimodal computing power orchestration system, comprising:

[0038] An acquisition module, configured to acquire multiple modal business flow data and extract business flow features of each of the modal business flow data;

[0039] A credit capacity value module, configured to input each of the service flow characteristics into a preset credit capacity calculation model for calculation to obtain the credit capacity value required by each of the service flow characteristics;

[0040] An initial computing power value module, configured to determine an initial computing power value required for business processing corresponding to each of the business flow characteristics based on a preset node utility function and each of the credit values;

[0041] A container module, configured to map the business flow corresponding to each business process to a corresponding container according to the initial computing power value required for each business process;

[0042] An allocation module is used to adjust the initial computing power value required for the business processing corresponding to each business flow according to the priority weight of the business flow corresponding to the container, generate a target computing power value, and allocate computing power resources according to the target computing power value.

[0043] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed, implements the multimodal computing power orchestration method as described in any one of the above items.

[0044] A fourth aspect of the present invention provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions, wherein when the program instructions are executed by a computer, the computer executes the multimodal computing power orchestration method as described in any one of the above items.

[0045] It can be seen from the above technical solutions that the present invention has the following advantages:

[0046] The present invention first acquires multimodal traffic flow data and extracts its traffic flow characteristics (including basic characteristics such as data type, flow rate, and urgency, as well as modality-specific characteristics). The traffic flow characteristics are then input into a preset capacity calculation model to calculate the capacity value (the information provision capability per unit of data) required for each traffic flow characteristic. Next, based on a preset node utility function (which measures the balance of node resources) and the capacity value, the initial computing power required for each service processing is determined. Based on the initial computing power value, the traffic flow is mapped to the corresponding container (e.g., matching CPU / GPU containers based on computing power requirements). Finally, the initial computing power value is dynamically adjusted according to the priority weight of the container-specific traffic flow to generate a target computing power value, and computing power resources are allocated based on this value, achieving dynamic perception of multimodal traffic flows and intelligent computing power scheduling. Through three core mechanisms: multimodal feature fusion analysis, capacity-driven dynamic computing power prediction, and priority-aware resource scheduling, the present invention systematically addresses the technical bottlenecks of traditional static scheduling, which cannot adapt to the dynamic changes of multimodal traffic flows and cannot achieve cross-modal fusion in single-modality processing. This achieves efficient and real-time multimodal data processing and balanced resource allocation. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0048] Figure 1 A flowchart of the steps of a multi-modal computing power orchestration method provided in Example 1 of the present invention;

[0049] Figure 2 This is a flowchart of a multi-modal computing power orchestration method provided in Example 1 of the present invention;

[0050] Figure 3 This is a structural block diagram of a multimodal computing power orchestration system provided in Example 2 of the present invention. DETAILED DESCRIPTION

[0051] Embodiments of the present invention provide a multimodal computing power orchestration method and system to address the technical problem that existing technologies, due to the limitations of static scheduling and single data type processing, cannot meet the requirements of multimodal business flows.

[0052] In order to make the purpose, features, and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0053] See also Figures 1 to 2 , Figure 1 This is a flowchart of the steps of a multimodal computing power orchestration method provided in Example 1 of the present invention.

[0054] The present invention provides a multimodal computing power orchestration method, comprising:

[0055] Step 101: Acquire multiple modal business flow data and extract business flow features of each modal business flow data.

[0056] In an embodiment of the present invention, multiple modal business flow data refers to the business interface of each sensing device receiving multiple different modal business flows from various sensors and data acquisition terminals through multiple communication ports, such as text, audio, video and other multimodal data.

[0057] Business flow features refer to multi-dimensional features such as basic features, modality-specific features, and cross-modal correlation features.

[0058] It is worth mentioning that the business interface of each sensing device receives multimodal data such as text, audio, and video from various sensors and data acquisition terminals through multiple communication ports, and extracts the features of each modal data to facilitate the subsequent prediction of the computing power required for each modal business.

[0059] Furthermore, step 101 includes the following sub-steps:

[0060] S11. Collect multiple modal business flow data from the business interfaces of different sensing devices.

[0061] In an embodiment of the present invention, the business interface of each sensing device receives multiple different modal business streams from various sensors and data acquisition terminals through multiple communication ports, such as text, audio, video and other multimodal data, and adds a globally unique timestamp (such as UTC time) to each data to ensure cross-modal data synchronization.

[0062] S12. Assign priority tags to the service flow data of each modality according to the service type, and store the service flow data of each modality into priority queues according to the priority tags.

[0063] In the embodiment of the present invention, the service type refers to real-time monitoring or non-real-time logging.

[0064] Priority labels refer to labels that classify multimodal business flow data into three levels: high, medium, and low.

[0065] A priority queue is a data structure whose core characteristic is that each element has a priority, and elements in the queue are sorted and processed according to their priority. In computing power orchestration scenarios that are aware of multimodal business flows, priority queues are primarily used to manage multimodal data awaiting processing, ensuring that high-priority business flows are prioritized, thereby optimizing the system's real-time performance and responsiveness.

[0066] Assign priority tags (high / medium / low) to multimodal business flow data based on business type (such as real-time monitoring, non-real-time log), and store them in the priority queue of the business cache module (high-priority data is processed first).

[0067] S13. Extracting the service flow features of each modal service flow data in the priority queue.

[0068] In the embodiment of the present invention, multi-dimensional features such as basic features, modality-specific features, and cross-modality correlation features of each modality service flow data in the priority queue are extracted respectively. Among them, basic features include: data type, traffic volume, urgency, and timestamp;

[0069] Modality-specific features: Image / video: resolution (e.g., 1920×1080), frame rate (FPS), object detection density (using the YOLO model to count the number of detection boxes);

[0070] Speech: sampling rate (e.g., 44.1kHz), language (identified by NLP classification models), noise decibel level (calculated by audio preprocessing algorithms);

[0071] Text: text length (number of words), semantic complexity (the variance of text embedding vectors calculated by the BERT model);

[0072] Cross-modal correlation features: Calculate the timestamp difference of multimodal data (e.g., the time difference between a video frame and the corresponding audio segment is ≤50ms) and mark modal pairs (e.g., "video-speech" pairs).

[0073] Step 102: Input each service flow feature into a preset credit capacity calculation model for calculation to obtain the credit capacity value required by each service flow feature.

[0074] In the embodiment of the present invention, information capacity refers to the information provision capability of a unit of data.

[0075] The service flow characteristics (such as resolution, frame rate, and semantic complexity) are input into the preset capacity calculation model for calculation to obtain the capacity value required for each service flow characteristic.

[0076] Furthermore, step 102 includes the following sub-steps:

[0077] S21. Input the feature data corresponding to each business flow feature and the business event attention matrix into the preset information quantity equation to obtain the information quantity of each modal event.

[0078] In the embodiment of the present invention, the business event attention matrix refers to the core parameter of the credit capacity calculation model, which is used to quantify the degree of attention of the model to multimodal events in different time and space ranges.

[0079] Modal event information volume is a fundamental parameter of the information capacity computing model, used to measure the total amount of effective information contained in events of a specific modality in a specific space and time. Its core function is to quantify the information value of multimodal data and provide a direct basis for the allocation of computing resources.

[0080] In multimodal cognitive computing, the trust capacity computing model attempts to explain the process by which machines extract information from data. Assuming that the event space X∈Rm×s×t is a tensor of perceptual modality (m), space (s), and time (t), inspired by the above phenomenon, this paper defines the amount of information obtained by an individual from each modal event in the event space as:

[0081] ; (1)

[0082] Where, The matrix is the matrix composed of the information of all events of the i-th mode, which is calculated as (For example, the higher the image resolution, The larger the value); ⊙ represents the matrix multiplication operation, is the business event attention matrix, for any event , whose different modal sub-events share attention Considering that humans have limited ability to process data, for a single individual, it is assumed that the sum of their attention in a certain time and space range is 1.

[0083] It is worth mentioning that the information equation refers to Formula 1.

[0084] S22: Use the information volume of each modal event, the maximum value of the credit capacity, and the characteristic data dimension to input a preset credit capacity calculation model to generate the credit capacity value required by each business flow feature.

[0085] In the embodiment of the present invention, the maximum value of the signal capacity , refers to the maximum value of the information providing capacity of unit data.

[0086] The characteristic data dimension, D, is a key parameter of the capacity computing model, used to describe the complexity and structural characteristics of multimodal business flow data. It reflects the data's combined complexity across multiple dimensions, including modality, space, and time, and is one of the core criteria for quantifying data processing difficulty and computing power requirements.

[0087] The information capacity computing model refers to the core algorithm module used to quantify the information value of multimodal business flow data and provide a foundation for the intelligent scheduling of computing resources. Its design goal is to simulate the human perception and processing mechanism of multimodal information and accurately measure data complexity and information density through mathematical modeling.

[0088] The credibility value C refers to the value of the information provision capability of unit data.

[0089] When the received event information is known, individuals will continuously adjust their limited attention in the process of perceiving the environment to maximize the amount of information. D is the dimension of the feature data, such as the three-dimensional tensor dimension of the video. It is the maximum value of the trust capacity, which indicates the ability to obtain the maximum amount of information from unit data. The calculation formula of the maximum value of the trust capacity is:

[0090] (2)

[0091] The model of credit capacity calculation can be expressed as:

[0092] (3)

[0093] According to the above formula, the required signal capacity value C of each service flow characteristic can be obtained.

[0094] Furthermore, the method further comprises the following sub-steps:

[0095] S31. Calculate the variance of each node resource utilization rate using the node resource utilization rate and the overall mean of the node resources.

[0096] In this embodiment of the present invention, node resource utilization refers to a core metric that measures the resource efficiency of computing nodes (such as servers and edge devices). It is used to assess the actual usage of resources such as CPU, memory, network bandwidth, and disk I / O. It is a fundamental input parameter for capacity calculation and computing power prediction models, directly influencing node utility function calculation and resource scheduling strategies.

[0097] The overall mean of node resources is one of the core metrics for measuring the balance of compute node resource utilization. It reflects the average utilization of various resources, including CPU, memory, network bandwidth, and disk I / O. It is a fundamental parameter for computing resource balance (variance) and node utility functions, helping the system assess the overall state of node load.

[0098] Node resource utilization variance is a core metric used to measure the balanced utilization of various resources (CPU, memory, network, and disk I / O) within a compute node. It quantifies the degree to which different resource utilization rates deviate from the overall mean, reflecting the degree of resource fragmentation. It is a key input parameter for trust capacity calculation and computing power prediction models.

[0099] It is worth mentioning that resource balance is a physical quantity that measures the degree of difference in resource utilization of CPU, memory, network, and disk IO. The smaller the resource balance, the smaller the degree of resource fragmentation. Its calculation formula is as follows:

[0100] (4)

[0101] Where, They represent the resource utilization of CPU, memory, network bandwidth and disk IO in node i after the task is scheduled to node i. is the overall mean of the resource;

[0102] Calculate the variance of resource utilization of each node:

[0103] (5)

[0104] The variance of resource utilization of each node is calculated in the above manner to facilitate the subsequent determination of the node utility function.

[0105] S32. Determine the node utility function based on the variance of resource utilization of each node.

[0106] In this embodiment of the present invention, the node utility function refers to a core mathematical model used to quantify the resource utilization balance and load status of computing nodes (such as servers and edge devices). Its design goal is to provide an intuitive priority basis for computing resource scheduling by comprehensively evaluating the utilization differences of various resources (CPU, memory, network, and disk I / O) within a node, thereby ensuring the efficient and stable processing of multimodal business flows.

[0107] In the multi-resource load balancing scheduling algorithm based on cooperative game theory, the utility function of the node can be defined as:

[0108] (6)

[0109] Where, represents the utility function value of node i, represents the resource utilization variance of node i, represents the overall mean value of the resources of node i.

[0110] S33. Set constraints according to the capacity of task offloading and resource allocation.

[0111] In the embodiment of the present invention, according to the capacity of task offloading and resource allocation, resource allocation needs to meet the following constraints:

[0112] (7)

[0113] Where, Indicates the decision parameter for whether task k is offloaded to edge server m, represents the computing resources allocated to task k on server m, represents the maximum resource requirement of task k.

[0114] S34. Based on cooperative game theory, set the objective function of multi-resource load balancing scheduling.

[0115] In an embodiment of the present invention, the objective function of the multi-resource load balancing scheduling algorithm based on cooperative game theory can be defined as:

[0116] (8)

[0117] Where, represents the objective function of the algorithm, represents the utility function value of node i, T represents the task set, and N represents the node set.

[0118] Step 103: Based on the preset node utility function and each credit value, determine the initial computing power value required for the business processing corresponding to each business flow feature.

[0119] In this embodiment of the present invention, the initial computing power value refers to the basic computing power requirement value generated by computing power prediction models based on the capacity value of multimodal business flows and the resource status of computing nodes. It serves as the starting point for computing power resource allocation, providing a quantitative benchmark for subsequent dynamic adjustments and ensuring a preliminary match between business flow processing and resource supply.

[0120] According to the preset node utility function value The initial computing power required for business processing corresponding to each business flow feature is predicted using each credit value C.

[0121] Furthermore, step 103 includes the following sub-steps:

[0122] S41. Classify the credit capacity values corresponding to the characteristics of each business flow and establish a mapping table of the basic computing power requirements of each mode.

[0123] In the embodiment of the present invention, the basic computing power requirement mapping table for each modality refers to a table that classifies the traffic flow characteristics according to the corresponding credit value and establishes a mapping table of basic computing power requirements for different modalities based on the classification results. The specific table is as follows:

[0124]

[0125] In a specific embodiment, the credit capacity value C of a certain video service is 1.5, and the basic computing power requirement is calculated according to the formula:

[0126] F base =(4+2×1.5,2+1.5,8+5×1.5)=(7,3.5,15.5) (9)

[0127] That means 7 CPU cores, 3.5 GPUs, and 15.5GB of memory.

[0128] S42. Determine a resource allocation correction coefficient based on a preset node utility function.

[0129] In the embodiment of the present invention, the resource allocation correction coefficient refers to a key parameter used to adjust the initial computing power requirement. Its core function is to dynamically correct the initial computing power value generated based on the credit value according to the real-time status of the computing node (such as resource balance, load pressure) or the real-time requirements of the business flow, to ensure that resource allocation meets both data processing requirements and adapts to changes in the system operating environment.

[0130] According to the node utility function value Calculate the resource allocation correction coefficient to prioritize high-capacity tasks and assign them to efficient nodes:

[0131] (10)

[0132] Physical meaning:

[0133] Efficient Node ( ≥8) can obtain 1.2 times the basic computing power, accelerating the processing of high-capacity tasks;

[0134] Inefficient nodes ( <5) Only 0.8 times the base computing power is allocated to avoid overload.

[0135] S43. Based on the basic computing power requirements and resource allocation correction coefficients corresponding to the basic computing power requirements mapping table of each modality, determine the initial computing power value required for the business processing corresponding to each business flow feature.

[0136] In an embodiment of the present invention, the basic computing power requirement corresponding to the modal basic computing power requirement mapping table and the resource allocation correction coefficient are combined to obtain the computing power requirement formula:

[0137] (11)

[0138] Where a is the urgency weight (e.g., 0.3), and urgency is converted from the priority label in the feature extraction stage (high = 1, medium = 0.5, low = 0); It is a three-dimensional vector, representing the number of CPU cores, GPUs, and GB of memory.

[0139] For multimodal fusion tasks (such as "video + voice" joint analysis), it is necessary to add the computing power requirements of each modality and increase the coordination overhead:

[0140] (12)

[0141] Where, Indicates the individual computing power requirements of each mode, is the synergy coefficient (e.g. 0.2), which indicates the additional computing power required for cross-modal synchronization.

[0142] Since container resources are usually integers, the calculation results need to be discretized:

[0143] , , (13)

[0144] For example:

[0145] A fusion task calculates , after discretization it is (7,3,15).

[0146] Specifically, enter:

[0147] Video service flow characteristics: 1080p resolution, 30fps frame rate, high urgency

[0148] Credibility value C=1.0

[0149] Target node utility =8.5 (efficient node)

[0150] Calculation process:

[0151] 1. Basic computing power requirement: F base =(2+1.0,1,2+4×1.0)=(3,1,6)

[0152] 2. Node correction factor: =8.5→Resource Allocation Correction Factor = 1.2

[0153] 3. Urgency correction: Urgency = 1 → F k =(3,1,6)×1.2×(1+0.3×1)=(4.68,1.56,9.36)

[0154] 4. Discrete processing (initial computing power value): F k =(5,2,10)

[0155] Output:

[0156] This video service requires an initial computing power of 5 CPU cores, 2 GPUs, and 10GB of memory.

[0157] Step 104: Map the business flows corresponding to each business process to corresponding containers based on the initial computing power required for each business process.

[0158] In the embodiment of the present invention, the container refers to a business container or a business application (APP software).

[0159] Use Kubernetes to dynamically adjust computing resource mapping, including CPU, memory, storage, and network resources, and map various business flows to different business containers and their business apps.

[0160] Specifically, the identified business flows are mapped to corresponding application programs (APPs). Each APP represents a type of business processing requirement, such as video processing, image recognition, or natural language processing.

[0161] According to F k The computing power type (such as whether GPU requirements are included) in the map will map the business flow to the corresponding APP:

[0162] >0→Video processing app (calling GPU container);

[0163] =0 and >1→Text processing app (calling a multi-core CPU container).

[0164] Step 105: According to the priority weight of the business flow corresponding to the container, the initial computing power value required for the business processing corresponding to each business flow is adjusted to generate a target computing power value, and computing power resources are allocated according to the target computing power value.

[0165] In this embodiment of the present invention, the target computing power value refers to the final computing power required to process the service flow after dynamic multi-dimensional adjustments. It is the result of adjusting the initial computing power value (based on trust value and node utility) in combination with factors such as priority weight, real-time load, and latency requirements. It is the direct basis for computing power resource allocation.

[0166] Computing resources refer to the hardware or virtualized resources available in compute nodes for processing business flows, including physical and virtual resources. Physical resources include: CPU (computing cores, such as Intel Xeon cores); GPU (graphics processing unit, such as NVIDIA A100 graphics card); memory (random access memory, such as DDR4); and storage and network (disk I / O bandwidth and network throughput, such as 10Gbps network cards). Virtual resources include the virtual CPUs and memory allocated to containers (such as Docker containers) or virtual machines (VMs); and serverless function computing resources in cloud-native environments.

[0167] For each app, priority and weight are assigned based on business needs and Quality of Service (QoS) requirements.

[0168] Analyze the real-time requirements of each business flow, such as processing latency and data freshness. Prioritize business flows that require rapid processing.

[0169] The weights are dynamically modified based on the real-time nature of the business flow and the current system load.

[0170] Using deep learning and graph neural network (GNN) technology, we predict the resource requirements of business flows, adjust the initial computing power required for business processing corresponding to each business flow, generate a target computing power value, and map it to the optimal computing resources. Based on the predicted target computing power value and mapping results, we use the Kubernetes container orchestration tool to schedule resources.

[0171] Furthermore, step 105 includes the following sub-steps:

[0172] S51. Calculate the priority weight of the service flow corresponding to each container based on the normalized result of the credit value of each container, the adjustment factor, and the service quality weight.

[0173] In the embodiment of the present invention, the credit value normalization result refers to mapping the credit value C to the interval [0, 1] through mathematical transformation, so as to facilitate linear combination or comparison with other parameters (such as QoS weight).

[0174] The adjustment factor refers to a manually set weight coefficient (0≤λ≤1) used to balance the proportion of credit capacity value and QoS weight in priority calculation.

[0175] The service quality weight refers to the measure of the service quality requirements of the business flow. It is usually quantified based on indicators such as delay threshold and data freshness, and its value range is [0, 1].

[0176] Priority weight refers to the final scheduling priority indicator that takes into account factors such as credit capacity, QoS requirements, and real-time correction. It is used to determine the processing order and resource allocation weight of service flows, and its value range is [0, 1].

[0177] The calculation formula of the initialized priority weight formula is:

[0178] (14)

[0179] Where, is the adjustment factor, such as 0.7; is the normalized result of the confidence value, ranging from [0,1]; is the service quality weight;

[0180] S52: Determine a normalized error of the actual delay based on a preset service flow delay threshold and the actual service flow delay.

[0181] In this embodiment of the present invention, the delay threshold refers to the maximum allowable delay in service flow processing, which is predefined by service requirements or Quality of Service (QoS) protocols and is typically measured in milliseconds. It serves as a baseline for determining whether service flow processing meets standards.

[0182] The actual service flow latency refers to the end-to-end time from when the service flow enters the system (service interface submodule) to when it is processed (output results), measured in milliseconds (ms).

[0183] The normalized error of the actual delay refers to the deviation between the actual delay and the threshold normalized to the range of [-1, 1], which is used to quantify the degree of delay exceeding the standard.

[0184] Dynamic correction needs to combine real-time indicators (such as processing delay) with the current system load. Make adjustments, specific steps:

[0185] Real-time error calculation:

[0186] Defining delay thresholds for service flows (For example, real-time business is 100ms), calculate the actual delay The normalized error is:

[0187] (15)

[0188] The actual delay is calculated according to the above formula The normalized error.

[0189] S53. Based on the preset load impact factor, the normalized error of the actual delay, the priority weight, and the sensitivity coefficient, the initial computing power value required for the business processing corresponding to each business flow is adjusted to generate a target computing power value.

[0190] In this embodiment of the present invention, the load impact factor refers to a regulatory parameter that reflects the current load status of a computing node. It is used to suppress the allocation of computing power to heavily loaded nodes to avoid overload. It is calculated based on the node's real-time utilization (such as CPU and memory utilization) and typically ranges from [0 to 1].

[0191] The sensitivity coefficient is a manually set adjustment parameter used to control the scheduling strategy's response to factors such as real-time errors and priority weights. It reflects the system's sensitivity to a specific indicator and typically ranges from [0 to 1].

[0192] Load balancing factor calculation:

[0193] Extract the current CPU utilization of the node , converted into load impact factor through sigmoid function:

[0194] (16)

[0195] Among them, when the load is greater than 70%, It approaches 1 quickly, suppressing the growth of weight.

[0196] Dynamic weight formula:

[0197] (17)

[0198] Where, is the sensitivity coefficient, such as 0.5; when the delay exceeds the threshold, >0, the weight will automatically increase; when the load is too high, Suppress weight increase.

[0199] Example:

[0200] =0.8 (high priority), =120ms (20% over threshold), =80%

[0201] =0.2, =0.88

[0202] =0.8×(1+0.5×0.2×0.88)=0.8704.

[0203] In the GNN model, As input features, and the computing power prediction value F k Combined, the formula is expressed as:

[0204] (18)

[0205] Key Roles:

[0206] High-weight tasks (such as real-time video) Will be forced to upgrade to ensure priority allocation of resources;

[0207] Low-weight tasks (such as non-real-time logs) It can be compressed appropriately to release resources for high-priority services.

[0208] Constraints for computing power allocation: Resource prediction must satisfy the priority-computing power binding relationship, for example:

[0209] ≥ × (19)

[0210] Where, The basic computing power requirements for the task, such as at least 1 CPU core for text processing.

[0211] Example:

[0212] A text task =1 core, =0.6

[0213] Minimum allocated computing power: 1×0.6=0.6 cores, rounded up to 1 core (meets basic requirements).

[0214] Combining real-time requirements (such as latency thresholds) with current node load, the GNN model is used to modify computing power requirements:

[0215]

[0216] Among them, ensure , retaining the basic computing power requirements.

[0217] S54. Allocate computing power resources according to the target computing power value.

[0218] In an embodiment of the present invention, the Kubernetes container orchestration tool is used to perform resource scheduling, that is, allocate computing power resources, based on the predicted target computing power value and the mapping result.

[0219] According to the target computing power Configure container resources, that is, allocate computing resources:

[0220]

[0221] If the node resources are insufficient, task splitting is triggered:

[0222] (If the task requires 4 GPU computing power, but a single node only has 2 GPUs, it will be split into 2 parallel processing stages).

[0223] Furthermore, the method further comprises the following sub-steps:

[0224] S61. Collect the actual resource utilization and processing delay of the node in real time.

[0225] In this embodiment of the present invention, actual node resource utilization refers to the ratio of the actual resource usage of a computing node (such as a server or edge device) to the total available resources at a given moment or time period. This reflects the real-time resource load and is a fundamental indicator for assessing node busyness.

[0226] Processing latency, the time difference between a business flow entering a computing node and completing processing, is a key metric for measuring service real-time performance. Depending on the business scenario, latency can be categorized as end-to-end latency and node processing latency.

[0227] It is worth mentioning that the actual resource utilization and processing delay of the real-time acquisition node need to be realized in combination with hardware sensors and operating system tools.

[0228] S62. Calculate the computing power prediction error rate using the actual resource utilization and processing delay of the node.

[0229] In the embodiment of the present invention, the computing power prediction error rate refers to the degree of deviation between the computing power prediction value and the actual computing power used, and is used to evaluate the accuracy of the computing power prediction model.

[0230] Collects actual resource utilization of nodes and processing delays , calculate the computing power prediction error:

[0231]

[0232] The computing power prediction error is calculated according to the above formula.

[0233] S63: Determine whether the computing power prediction error rate is greater than a preset error rate threshold.

[0234] In the embodiment of the present invention, the preset error rate threshold is 15%.

[0235] Determine whether the error rate is greater than 15%.

[0236] S64. If yes, adjust the business event attention matrix, generate a new business event attention matrix, and jump to the step of inputting the feature data corresponding to each business flow feature and the business event attention matrix into the preset information quantity equation to obtain the information quantity of each modal event.

[0237] In an embodiment of the present invention, if the error rate is greater than 15%, the attention matrix A in the credit calculation model or the computing power prediction model parameters (such as the weight coefficient 100 in Formula 6) are automatically adjusted, and the corresponding steps are jumped to form a "perception-prediction-scheduling-optimization" closed loop.

[0238] The present invention can process data in multiple modes, such as text, images, voice, etc., to meet the needs of different business scenarios. This multimodal support capability makes the system more flexible and versatile. The present invention constructs a capacity calculation and computing power prediction model for multimodal business, and allocates a computing center for multimodal business flow perception. The present invention also proposes a method for computing power mapping, mapping various types of business to corresponding containers, and allocating computing power resources to them according to business priority. And through computing power orchestration, customized services can be provided according to the needs of different businesses, more computing power resources can be allocated to specific businesses, or specific algorithm models can be configured for them.

[0239] See also Figure 3 , Figure 3 This is a structural block diagram of a multimodal computing power orchestration system provided in Example 2 of the present invention.

[0240] The present invention provides a multimodal computing power orchestration system, comprising:

[0241] The acquisition module 301 is used to acquire multiple modal business flow data and extract business flow features of each modal business flow data;

[0242] The credit capacity value module 302 is used to input each service flow feature into a preset credit capacity calculation model for calculation to obtain the credit capacity value required by each service flow feature;

[0243] The initial computing power value module 303 is used to determine the initial computing power value required for the service processing corresponding to each service flow feature based on the preset node utility function and each credit value;

[0244] The container module 304 is used to map the business flow corresponding to each business process to the corresponding container according to the initial computing power value required by each business process;

[0245] The allocation module 305 is used to adjust the initial computing power value required for the business processing corresponding to each business flow according to the priority weight of the business flow corresponding to the container, generate a target computing power value, and allocate computing power resources according to the target computing power value.

[0246] Furthermore, the acquisition module 301 includes:

[0247] The acquisition submodule is used to collect multiple modal business flow data from the business interfaces of different sensing devices;

[0248] The allocation submodule is used to assign priority tags to the business flow data of each modality according to the business type, and store the business flow data of each modality into the priority queue according to the priority tag;

[0249] The extraction submodule is used to extract the business flow features of each modal business flow data in the priority queue.

[0250] Furthermore, the credibility value module 302 includes:

[0251] The input submodule is used to input the feature data corresponding to each business flow feature and the business event attention matrix into the preset information quantity equation to obtain the information quantity of each modal event;

[0252] The credit capacity value submodule is used to use the information volume of each modal event, the maximum credit capacity value and the characteristic data dimension to input the preset credit capacity calculation model to generate the credit capacity value required by each business flow feature.

[0253] Furthermore, the system also includes:

[0254] The calculation submodule is used to calculate the variance of each node resource utilization rate by using the node resource utilization rate and the overall mean of the node resources;

[0255] The node utility function submodule is used to determine the node utility function based on the variance of resource utilization of each node;

[0256] Setting submodules to set constraints based on the capacity of task offloading and resource allocation;

[0257] The objective function submodule is used to set the objective function of multi-resource load balancing scheduling based on cooperative game theory.

[0258] Furthermore, the initial computing power value module 303 includes:

[0259] Establish a submodule to classify the credit capacity values corresponding to each business flow feature and establish a mapping table of basic computing power requirements for each modality;

[0260] A resource allocation correction coefficient submodule is used to determine a resource allocation correction coefficient based on a preset node utility function;

[0261] The initial computing power value submodule is used to determine the initial computing power value required for business processing corresponding to each business flow feature based on the basic computing power requirements and resource allocation correction coefficient corresponding to the basic computing power requirements mapping table of each modality.

[0262] Furthermore, the allocation module 305 includes:

[0263] The priority weight submodule is used to calculate the priority weight of the service flow corresponding to each container based on the normalized result of the capacity value of each container, the adjustment factor and the service quality weight;

[0264] A normalized error submodule, configured to determine a normalized error of an actual delay based on a preset service flow delay threshold and an actual service flow delay;

[0265] The target computing power submodule is used to adjust the initial computing power required for the service processing of each service flow based on the preset load impact factor, the normalized error of the actual delay, the priority weight, and the sensitivity coefficient to generate the target computing power value;

[0266] The computing power resource submodule is used to allocate computing power resources according to the target computing power value.

[0267] Furthermore, the system also includes:

[0268] Real-time acquisition submodule, used to collect the actual resource utilization and processing delay of nodes in real time;

[0269] The error rate submodule is used to calculate the computing power prediction error rate using the actual resource utilization and processing delay of the node;

[0270] The judgment submodule is used to determine whether the computing power prediction error rate is greater than a preset error rate threshold;

[0271] The adjustment submodule is used to adjust the business event attention matrix, generate a new business event attention matrix, jump to execute the step of inputting the feature data corresponding to each business flow feature and the business event attention matrix into the preset information amount equation to obtain the information amount of each modal event.

[0272] A third embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed, a multimodal computing power orchestration method as described in any embodiment of the present invention is implemented.

[0273] A fourth embodiment of the present invention provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes a multimodal computing power orchestration method as described in any embodiment of the present invention.

[0274] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0275] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0276] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0277] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0278] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0279] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multimodal computing power orchestration method, characterized in that: include: Acquire multiple modal business flow data, and extract business flow features of each of the modal business flow data; Inputting each of the service flow characteristics into a preset credit capacity calculation model for calculation to obtain the credit capacity value required by each of the service flow characteristics; Determining the initial computing power required for the business processing corresponding to each of the business flow characteristics based on the preset node utility function and each of the credit values; Mapping the business flows corresponding to the business processes to corresponding containers based on the initial computing power required for the business processes; According to the priority weight of the business flow corresponding to the container, the initial computing power value required for the business processing corresponding to each business flow is adjusted to generate a target computing power value, and computing power resources are allocated according to the target computing power value.

2. The multimodal computing power orchestration method according to claim 1, characterized in that: The acquiring of multiple modal business flow data and extracting business flow features of each of the modal business flow data includes: Collecting multiple modal business flow data from the business interfaces of different sensing devices; Assigning a priority tag to each of the modal service flow data according to the service type, and storing each of the modal service flow data in a priority queue according to the priority tag; Extract the business flow characteristics of each modal business flow data in the priority queue.

3. The multimodal computing power orchestration method according to claim 1, characterized in that: The inputting each of the service flow characteristics into a preset credit capacity calculation model for calculation to obtain the credit capacity value required by each of the service flow characteristics includes: Input the feature data corresponding to each of the business flow features and the business event attention matrix into a preset information quantity equation to obtain the information quantity of each modal event; The information volume of each modal event, the maximum value of the credit capacity and the characteristic data dimension are input into a preset credit capacity calculation model to generate the credit capacity value required by each business flow feature.

4. The multimodal computing power orchestration method according to claim 1, characterized in that: Also includes: The variance of each node's resource utilization is calculated using the node resource utilization and the overall mean of the node resources; Determining a node utility function based on the variance of resource utilization of each of the nodes; Set constraints based on the capacity of task offloading and resource allocation; Based on cooperative game theory, the objective function of multi-resource load balancing scheduling is set.

5. The multimodal computing power arrangement method according to claim 1, characterized in that: The determining, based on the preset node utility function and each of the credit values, the initial computing power required for the service processing corresponding to each of the service flow characteristics includes: Classify the credit capacity values corresponding to the business flow characteristics and establish a mapping table of basic computing power requirements for each mode; Determine the resource allocation correction coefficient based on the preset node utility function; Based on the basic computing power requirements and resource allocation correction coefficients corresponding to the basic computing power requirements mapping tables of each modality, the initial computing power value required for the business processing corresponding to each business flow feature is determined.

6. The multimodal computing power orchestration method according to claim 1, characterized in that: The step of adjusting the initial computing power required for the business processing corresponding to each business flow according to the priority weight of the business flow corresponding to the container, generating a target computing power value, and allocating computing power resources according to the target computing power value includes: Calculating the priority weight of the service flow corresponding to each container according to the normalized result of the credit capacity value of each container, the adjustment factor and the service quality weight; Determine a normalized error of the actual delay based on a preset delay threshold of the service flow and the actual delay of the service flow; Based on a preset load impact factor, the normalized error of the actual delay, the priority weight, and the sensitivity coefficient, an initial computing power value required for service processing corresponding to each of the service flows is adjusted to generate a target computing power value; Allocate computing power resources according to the target computing power value.

7. The multimodal computing power arrangement method according to claim 3, characterized in that: Also includes: Real-time collection of actual node resource utilization and processing delay; Calculating a computing power prediction error rate using the actual resource utilization of the node and the processing delay; Determining whether the computing power prediction error rate is greater than a preset error rate threshold; If so, adjust the business event attention matrix to generate a new business event attention matrix, and jump to execute the step of inputting the feature data corresponding to each of the business flow features and the business event attention matrix into the preset information quantity equation to obtain the information quantity of each modal event.

8. A multimodal computing power orchestration system, characterized in that: include: An acquisition module, configured to acquire multiple modal business flow data and extract business flow features of each of the modal business flow data; A credit capacity value module, configured to input each of the service flow characteristics into a preset credit capacity calculation model for calculation to obtain the credit capacity value required by each of the service flow characteristics; An initial computing power value module, configured to determine an initial computing power value required for business processing corresponding to each of the business flow characteristics based on a preset node utility function and each of the credit values; A container module, configured to map the business flow corresponding to each business process to a corresponding container according to the initial computing power value required for each business process; An allocation module is used to adjust the initial computing power value required for the business processing corresponding to each business flow according to the priority weight of the business flow corresponding to the container, generate a target computing power value, and allocate computing power resources according to the target computing power value.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the multimodal computing power orchestration method according to any one of claims 1 to 7 is implemented.

10. A computer program product, characterized in that The computer program product includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions, wherein when the program instructions are executed by a computer, the computer executes the multimodal computing power orchestration method as described in any one of claims 1 to 7.