Edge Model Scheduling Method and System Integrating Graph Neural Network and Mixture-of-Experts Model

Through the edge model scheduling method of fusion graph neural and hybrid expert model, the resource requirements and compatibility problems of deep learning model in edge environments are solved, and the efficient deployment and stable operation of the model on edge devices is achieved.

CN119883660BActive Publication Date: 2025-06-10SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510376494.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-10
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

In edge environments, the high resource requirements of deep learning models are difficult to meet on weaker edge computing resources, and at the same time, large models are deployed into unseen application environments, resulting in unstable operation or performance degradation.

Method used

The edge model scheduling method of fusion graph neural and hybrid expert models is adopted to preprocess and feature extraction of multimodal data, and a general feature data set is generated, and the edge device relationship diagram is constructed using heuristic algorithms and graph neural networks to optimize the placement and operation of the model in heterogeneous edge devices.

Benefits of technology

It effectively solves the problem of placement of large language models in edge environments, improves the running speed and load balancing of the model in resource-constrained environments, and ensures the successful deployment and stable operation of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119883660B_ABST
    Figure CN119883660B_ABST
Patent Text Reader

Abstract

The present invention belongs to the fields of computer artificial intelligence and wireless communication networks, and provides an edge model scheduling method and system that integrates graph neural networks and a mixture of experts model. By collecting multi-modal, hardware parameter, and model performance data, and based on the resource status of edge devices and the computing requirements of tasks, a graph neural network structure is constructed to capture the topological relationships and conditional dependencies among data, hardware, and models. By performing feature processing locally where the data is generated, the time delay of data transmission to the cloud or central server is significantly reduced, thereby improving the resource utilization and model processing efficiency of the overall system. The graph neural network prediction model constructed based on general features and a new message passing mechanism can effectively address the zero-shot prediction problem in unfamiliar hardware device environments, without the need for specialized training and optimization for each hardware device and can generate diverse prediction results, effectively reducing the risk of overfitting of the prediction results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer artificial intelligence and wireless communication network, and in particular relates to an edge model scheduling method and system that integrates graph neural network and hybrid expert model. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] Large models have made remarkable progress in the practice of multiple application fields such as computer vision, natural language processing, and big data analysis. With the rapid development of emerging technologies such as the Internet of Things (IoT) and 5G communications, the deployment and reasoning needs of large models in cloud-edge environments are becoming increasingly urgent. However, the application of large model technology in edge scenarios such as the Industrial Internet is limited by complex and diverse application scenarios and equipment conditions. There are still many challenges in implementing the deployment and operation of large models in edge environments. In order to meet the computing resource requirements of large models, a common practice is to transmit data with large-scale servers far away from the data source through network communication, so as to assist in processing data and training models with the help of cloud computing resources. However, this method has problems such as transmission delay and privacy security. Among them, one of the main challenges is to meet the high resource requirements of deep learning on weaker edge computing resources. These devices often have relatively weak CPU and GPU performance and limited storage capacity, which cannot be compared with high-performance servers in the cloud. Another is how to deploy a trained large model to an actual zero-shot application environment, which means that the model needs to recognize and classify parameters and scenarios that have never been seen before. The hardware platforms in different actual application environments have different parameters and structures, and the models may face compatibility issues, resulting in failure to operate normally or performance degradation. How to efficiently manage and allocate edge resources and ensure the success rate of large models after the first deployment, especially how to predict and optimize the placement of large models in complex edge environments, ensure the efficiency and quality of completing tasks to ensure good end-to-end applications, has also become a difficult challenge in the field of edge computing. Summary of the invention

[0004] In order to solve the above problems, the present invention proposes an edge model scheduling method and system that integrates graph neural networks and hybrid expert models. The present invention can effectively solve the placement problem of large language models in edge environments in the prior art, so that the expert models in the multimodal hybrid expert model can be reasonably and efficiently distributed to different edge devices, thereby ensuring the successful deployment and stable operation of large language models in heterogeneous edge devices, and improving the running speed and load balance of artificial intelligence models in resource-constrained environments.

[0005] According to some embodiments, the first solution of the present invention provides an edge model scheduling method that integrates graph neural networks and a mixture of experts model, adopting the following technical solutions:

[0006] The edge model scheduling method that integrates graph neural networks and a mixture of experts model includes:

[0007] Preprocess the multimodal data and generate a multimodal mixture of experts model based on the preprocessed multimodal data;

[0008] Obtain the edge device hardware parameters and splice the multimodal data features to generate a general feature dataset;

[0009] Based on the general feature dataset, use a heuristic algorithm to enumerate and generate multiple candidate placement schemes for the mixture of experts model, and filter out multiple feasible candidate placement schemes according to the edge device hardware constraints;

[0010] Based on the general feature dataset, use the feature embedding method and the message passing mechanism to set the expert models and the corresponding edge devices as independent nodes with feature attributes, and establish edge connections according to the feasible candidate placement schemes to construct an edge device relational graph neural network;

[0011] Establish a cost function according to the resource status of the edge devices and the computational requirements of the models, and iteratively train the candidate placement schemes for the edge device relational graph neural network with the goal of minimizing the cost function value to obtain the optimal placement scheme for the mixture of experts model.

[0012] Furthermore, the processing process of the mixture of experts model is specifically as follows:

[0013] Obtain the multimodal data for normalization processing and feature extraction, and use feature embedding to map the multimodal data to a unified feature space;

[0014] Generate modality tokens corresponding to the multimodality from the unified feature space to obtain multiple modality tokens;

[0015] Use gated routing to sort the experts models in descending order of weights, and assign the multiple modality tokens to the corresponding number of experts models with high weights and not fully occupied for processing;

[0016] Fuse the multiplication results of the output results of the multiple experts models and their weights with the original modality tokens to obtain a fusion result, and then normalize and decode the fusion result as the final output result of the mixture of experts model.

[0017] Furthermore, the step of using a heuristic algorithm to enumerate and generate multiple candidate placement schemes for the mixture of experts model based on the general feature dataset and filtering out multiple feasible candidate placement schemes according to the edge device hardware constraints is specifically as follows:

[0018] Initialize the edge device relational graph neural network and assign values according to the types and features of each node to generate multiple model nodes and hardware nodes, where the model nodes include data nodes, routing nodes, expert nodes, and output nodes;

[0019] Set the constraints of the hardware nodes based on the feature values in each node;

[0020] After sorting the expert nodes in a random and linear sequence, use the heuristic enumeration method to combine and allocate the model nodes and hardware nodes to generate multiple candidate placement schemes for the mixture of experts model;

[0021] Judge the validity of the candidate placement schemes based on the hardware constraints, filter out the infeasible candidate placement schemes, and screen out multiple feasible candidate placement schemes.

[0022] Furthermore, the placement process of the hardware nodes is combined and allocated with the model nodes according to the following three rules, specifically:

[0023] Each hardware node receives at least one model node;

[0024] The hardware nodes are arranged from weak to strong in terms of resources and performance;

[0025] The model nodes are placed into the hardware nodes in the order of a single line of data flow during the placement process.

[0026] Furthermore, the construction process of the edge device relational graph neural network is specifically as follows:

[0027] Use the feature embedding method to map the general feature data set to a continuous vector space to obtain a set of feature vectors;

[0028] According to the candidate placement schemes, sequentially transfer the corresponding feature vectors from the expert nodes, data nodes, routing nodes, and output nodes to the hardware nodes assigned to them to transmit feature information;

[0029] Then, perform the reverse transfer of feature information from the hardware nodes to the associated nodes;

[0030] Perform the transfer of feature information according to the data flow order of processing modal tokens in each run of the expert model;

[0031] Based on this, construct edges according to the relationships between each node to obtain the edge device relational graph neural network.

[0032] Furthermore, the cost function includes the amount of data processed by the expert model per unit time, the load balance of the edge nodes, and the communication time for the expert model to process modal tokens.

[0033] According to some embodiments, the second solution of the present invention provides an edge model scheduling system integrating graph neural network and mixture of experts model, adopting the following technical solutions:

[0034] An edge model scheduling system integrating graph neural network and mixture of experts model, comprising:

[0035] A mixture of experts model generation module, configured to preprocess multimodal data and generate a multimodal mixture of experts model based on the preprocessed multimodal data;

[0036] A feature splicing module, configured to obtain the edge device hardware parameters and splice the multimodal data features to generate a general feature dataset;

[0037] A feasible candidate placement scheme determination module, configured to enumerate and generate multiple candidate placement schemes of the mixture of experts model based on the general feature dataset by using a heuristic algorithm, and screen out multiple feasible candidate placement schemes according to the edge device hardware constraints;

[0038] A graph neural network construction module, configured to set the expert model and the corresponding edge device as independent nodes with feature attributes based on the general feature dataset by using a feature embedding method and a message passing mechanism, and establish edge connections according to the feasible candidate placement schemes to construct an edge device relationship graph neural network;

[0039] An optimal placement scheme determination module, configured to establish a cost function according to the resource status of the edge device and the calculation requirements of the model, and iteratively train the candidate placement schemes of the edge device relationship graph neural network with the goal of minimizing the cost function value to obtain the optimal placement scheme of the mixture of experts model.

[0040] According to some embodiments, the third solution of the present invention provides a computer-readable storage medium.

[0041] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in the edge model scheduling method integrating graph neural network and mixture of experts model as described in the first aspect above are implemented.

[0042] According to some embodiments, the fourth solution of the present invention provides a computer device.

[0043] A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps in the edge model scheduling method integrating graph neural network and mixture of experts model as described in the first aspect above are implemented.

[0044] According to some embodiments, the fifth aspect of the present invention provides a computer program product or a computer program.

[0045] The present invention provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the edge model scheduling method that integrates a graph neural network and a mixture-of-experts model as described in the first aspect above.

[0046] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0047] Facing the application scenario of timely and efficient processing of multi-modal data in the industrial Internet environment, the present invention can make full use of the resources of edge nodes with limited local computing power, directly perform feature processing locally where the data is generated, reduce the time delay of transmitting data to the cloud or the central server, thereby improving the computing speed and operation efficiency of the multi-modal mixture-of-experts model. At the same time, it can also make full use of the computing resources of edge nodes, improve the resource utilization of the overall system and the model processing efficiency;

[0048] The present invention can handle the zero-shot prediction problem in hardware device environments that have not been seen, that is, beyond the scope of the training data set. Under zero-shot conditions, it is not necessary to perform specialized training and optimization for each type of hardware device, thereby improving the applicable range and flexibility of the expert model, increasing the success rate and response speed of the first placement of the multi-modal mixture-of-experts model in an unfamiliar edge environment, and reducing the complexity and cost of the preliminary preparation work for model deployment;

[0049] The GNN zero-shot prediction model of the present invention based on a new message passing mechanism uses a heuristic enumeration algorithm to iteratively obtain an allocation scheme defined based on multiple cost parameters, generating diverse prediction results for the same model execution cycle, effectively reducing the risk of overfitting of the prediction results. Each node of the model can better synthesize the complex relationships and feature information in the overall graph neural network, which helps to effectively search and identify the most cost-effective allocation strategy in a complex decision space. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0051] Figure 1 is a flowchart of an edge model scheduling method that integrates a graph neural network and a mixture-of-experts model in an embodiment of the present invention;

[0052] Figure 2 is the heuristic enumeration process in an embodiment of the present invention;

[0053] Figure 3 It is a flowchart of the feature embedding method in an embodiment of the present invention;

[0054] Figure 4 It is a flowchart of the new message passing mechanism in an embodiment of the present invention;

[0055] Figure 5 It is a structural diagram of the multi-modal mixture of experts model (MMoE) in an embodiment of the present invention;

[0056] Figure 6 It is a flowchart of deployment and implementation in an embodiment of the present invention. Detailed implementation manners

[0057] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0058] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0059] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless otherwise clearly specified in the context, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0060] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0061] Embodiment 1

[0062] This embodiment provides an edge model scheduling method that integrates graph neural networks and a mixture of experts model. This embodiment takes the application of this method to a server as an example. It can be understood that this method can also be applied to a terminal, or to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, web servers, cloud communications, middleware services, domain name services, security services CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not limit this. In this embodiment, the method includes the following steps:

[0063] Preprocess the multimodal data, and generate a multimodal mixture of experts model based on the preprocessed multimodal data;

[0064] Obtain the edge device hardware parameters and concatenate the multimodal data features to generate a general feature dataset;

[0065] Based on the general feature dataset, use a heuristic algorithm to enumerate and generate multiple candidate placement schemes for the mixture of experts model, and screen out multiple feasible candidate placement schemes according to the edge device hardware constraints;

[0066] Based on the general feature dataset, use the feature embedding method and the message passing mechanism to set the expert models and the corresponding edge devices as independent nodes with feature attributes, and establish edge connections according to the feasible candidate placement schemes to construct an edge device relational graph neural network;

[0067] Establish a cost function according to the resource status of the edge devices and the computational requirements of the models, and iteratively train the candidate placement schemes for the edge device relational graph neural network with the goal of minimizing the cost function value to obtain the optimal placement scheme for the mixture of experts model.

[0068] As Figure 1 shown, an edge model placement prediction algorithm based on a graph neural network and a multimodal mixture of experts model includes the following steps:

[0069] Step (1): Generate a dataset composed of general features according to the edge hardware parameters and multimodal data, such as the number of CPUs, transmission bandwidth, etc.;

[0070] Step (2): Generate multiple candidate placement schemes based on heuristic enumeration and judge the feasibility of the schemes according to the hardware constraints;

[0071] Step (3); Using the overall graph structure constructed from the feature dataset and the candidate placement scheme, with the help of the feature embedding method and the new message passing mechanism, set the expert model and the corresponding edge devices as independent nodes with feature attributes, and establish edge connections according to the placement scheme;

[0072] Step (4); Based on the generated graph neural network, establish a training dataset and a loss function based on parameters such as the running success rate and IO latency, and iteratively train the GNN model to generate prediction values to select the best placement scheme.

[0073] The specific process is as follows:

[0074] In the said step (1), the meaning and content of the dataset composed of general features include:

[0075] In order to better utilize the model to predict the running cost and the advantages and disadvantages of the placement scheme, the general features must best reflect the attributes of the running cost, such as hardware parameters, data structure and other parameters.

[0076] On the one hand, general features mean that they have a wider adaptability in different edge devices and multi-modal data. This enables the same model or algorithm to run on multiple devices without the need for specialized optimization and adjustment for each device, improving the running success rate when placing the model. Moreover, general features can map the features of data from different modalities (such as image data, text data, audio data, etc.) into the same feature space, that is, the multi-modal feature space, to achieve the fusion and unification of multi-modal data, helping the model to better understand and process complex information, and thus enabling migration and application in different tasks.

[0077] On the other hand, when deploying the trained model to the actual application environment, it is necessary to consider how to deal with data content different from the training data, that is, the generalization ability of the model. Training the model with general features can have better generalization performance because general features often can better reflect the core patterns and laws, and can also simplify the overall complexity to a certain extent, having better robustness to noise interference in the real scenario and the migration ability to handle diverse tasks, thus being able to help the model cross different scenarios and data to adapt to diverse features.

[0078] The following Table 1 shows the general features used in the present invention to initialize the parameters of each node in the graph neural network:

[0079] Table 1 Graph Node Features

[0080]

[0081] In step (2), for generating multiple candidate placement schemes based on heuristic enumeration and judging the feasibility of the schemes according to hardware constraints, the specific steps are as follows, and the overall process is as Figure 2 shown:

[0082] Initialize the graph neural network and assign values according to the node types and features in Table 1 to generate multiple model nodes and hardware nodes; among them, the model nodes include data nodes, routing nodes, expert nodes, and output nodes

[0083] Set the constraints of the hardware nodes based on the feature values in each node. For example, the resource parameters of the hardware nodes can meet the storage and calculation requirements of the model node parameters, and whether the transmission speed between the hardware nodes does not exceed the preset threshold, etc.;

[0084] After sorting the expert nodes in a random and linear sequence, use the heuristic enumeration method to combine and allocate the model nodes and hardware nodes, and at the same time ensure that the placement process of the hardware nodes conforms to the following three rules to obtain multiple candidate placement schemes for the hybrid expert model:

[0085] Each hardware node receives at least one model node, that is, multiple model nodes can be assigned to the same hardware node to ensure the independence of each expert model;

[0086] The hardware nodes are arranged from weak to strong in terms of resources and performance to ensure the load balance of multiple nodes;

[0087] The model nodes are placed into the hardware nodes in the order of a single line of data flow during the placement process. Specifically, the model nodes are placed into the hardware nodes in the order of data flow from "data node" to "routing node" to "expert node" to "output node" during the placement process, and there is no data loop, thereby reducing the complexity of data flow and the overhead of network communication.

[0088] Based on the hardware constraints, initially judge the effectiveness of the candidate placement schemes, filter out the infeasible candidate placement schemes, and screen out multiple feasible candidate placement schemes.

[0089] In step (3), based on the above steps, multiple feasible candidate placement schemes are obtained. The present invention designs and adopts a feature embedding method (Feature Embedding) and a new message passing mechanism (Massage Passing) to help the nodes obtain the feature information of adjacent nodes and the global network. The process is as Figure 3 and Figure 4 shown:

[0090] Using the Feature Embedding method, high-dimensional sparse features are mapped to a continuous vector space, thereby preserving the core semantic information of the data, better reflecting the implicit relationships and distribution laws between features, and improving the generalization ability of the model;

[0091] The present invention designs a new Message Passing mechanism to iteratively update the special information of each independent node multiple times. By aggregating parameters, the information in adjacent nodes is used to update the parameters in the central node, so that the node contains the feature information of adjacent nodes and the global network, and relevant feature information is obtained according to the data flow situation. The specific steps are as follows:

[0092] According to the feasible candidate placement scheme in step (2), feature information is transmitted from the Expert Node, Dataset Node, Router Node, and End Node to the hardware nodes assigned to them;

[0093] Through the reverse direction of the transmission in the previous step, feature information is transmitted from the hardware nodes to the associated nodes;

[0094] According to the data flow order of processing tokens in each run of the MMoE model, feature information is transmitted;

[0095] After the above updates are completed, the feature information of all nodes is integrated in the Final MLP Layer, and the resulting result will be used to predict the running cost of the MMoE model using this candidate placement scheme for optimal placement scheme prediction.

[0096] In step (4), based on the graph neural network generated in the above steps, a loss function based on parameters such as running success rate and latency is established, and the GNN model is iteratively trained using the training dataset to generate the MMoE model cost prediction value to select the optimal placement scheme. The process is as Figure 2 shown:

[0097] In order to accurately reflect the effect and performance of the model placement scheme, the present invention defines a model cost function representing the graph neural network:

[0098] (1);

[0099] Throughput ( Throughout , ( ) refers to the amount of data that the MMoE model can process per unit time. For the execution of a given query, the throughput is defined as the number of tokens arriving at the output node per unit time, expressed in tokens / s (the number of tokens processed per second).

[0100] Load balancing ( Balance , ) is used to select a solution that can make edge hardware nodes obtain balanced data processing opportunities as much as possible during the model application process, thereby avoiding over - invocation of some edge hardware devices or neglect of some devices caused by significant data category imbalance, and helping the model better integrate multi - modal data information, improving the accuracy and robustness of video understanding. Among them, is the number of experts, is the probability that the -th expert is selected, specifically:

[0101] (2);

[0102] Processing time ( Processing Time , ) represents the time required for a token to be passed into the hardware node where an expert model is located until the processed data is sent out from the node.

[0103] Communication time ( Communication Time , ) represents the sum of all the time required for a token to be transmitted between nodes , including transmission time ( Transfer Time , ) and queuing time ( Queue Time , ), is the number of experts:

[0104] (3);

[0105] Placement success ( Placement Success , ) represents the success situation of the MMoE model running:

[0106] (4);

[0107] Train and iterate the GNN model according to multiple cost function parameters. The loss function used here is the mean squared logarithmic error (MSLE), which contains the actual value and the predicted value , and the formula is as follows:

[0108] (5);

[0109] The GNN model generated through multiple iterations produces prediction results for multiple candidate placement schemes of the MMoE models and conducts comprehensive analysis, and finally selects the optimal MMoE model placement scheme.

[0110] For the iterative training of the GNN model, iterative processing is performed on different feasible candidate placement schemes. A trained MMoE model is required first. The GNN is divided into two processes: training and running.

[0111] Training uses the loss function to generate multiple candidate schemes based on the training data including hardware nodes and MMoE model parameters, and iteratively optimizes the model until a good prediction result is formed.

[0112] Running is to generate an optimal placement scheme result using the trained GNN model according to the hardware nodes and MMoE model parameters to be predicted.

[0113] Such as Figure 5 As shown, a complete execution process of the multi-modal mixture of experts model is as follows:

[0114] Multi-modal data collection, including image data, text data, and audio data;

[0115] With the help of an encoder for standardization processing and feature extraction, the multi-modal data is feature-mapped to a unified multi-modal feature space through data embedding and multiple tokens are generated. Each execution process processes the content of one token.

[0116] Using the gating router to sort in descending order according to the weights given by the expert model, the data is allocated to k experts with high weights and not fully occupied, that is, these k experts perform better on the current token than other experts. Among them, the structure of the expert model (MoE) is uniformly a feed-forward neural network (FFN), and the main difference is the different parameters learned during the training process;

[0117] Using the method of residual connection and normalization, the multiplication result of the output results of multiple expert models and their weights is combined with the original token. Assuming the input token is , the th expert weight calculated by the gating router layer is , the output result of the th expert model is , then the final output result can be expressed as:

[0118] (6);

[0119] Fuse the multi-modal data information, perform normalization, and decode it into the final output result;

[0120] Set tags and parameter values according to the above content to facilitate the construction and prediction of the graph neural network.

[0121] As Figure 6 shown, this embodiment aims to solve the placement problem of the multi-modal mixture of experts model in the edge environment and improve the running speed and load balance of the model in resource-constrained environments. The following is the specific implementation of the present invention for deployment and operation on edge devices:

[0122] Data collection and preprocessing. First, determine the multi-modal task objective, and use sensors and input devices to obtain and store multi-modal data from edge devices, including data of three modalities: raw text data, raw image data, and raw audio data; obtain hardware parameter data according to the Internet of Things hardware devices, and obtain corresponding model performance and structure data according to the training and performance optimization of the multi-modal mixture of experts model.

[0123] Preprocess the above data, such as data cleaning operations, to ensure the quality and consistency of the data. For example, natural language processing techniques can be used to perform operations such as word segmentation and word vector representation on text data, that is, split the text into words so that the model can process each word, remove irrelevant characters, punctuation marks, numbers, etc. and retain meaningful text information, and then convert the text into a sequence form that the model can process; perform operations such as grayscale conversion, size adjustment, and channel order adjustment on image data. For example, the picture size can be adjusted to a fixed size and the channel order can be adjusted to the format required by the model to facilitate the model to quickly receive data; for the adjustment of audio data, operations such as converting the audio sampling rate to the sampling rate required by the model and splitting long audio into short segments of fixed length can be performed to ensure that the data can be effectively understood and processed by the model.

[0124] Construction and training of the multi-modal mixture of experts model. A complete multi-modal mixture of experts model needs to be constructed in advance before deploying the model. The model contains multiple expert sub-models, and each expert sub-model is responsible for processing one or more modalities of data, such as feasible models like DeepSeek-V3, Mixtral 8x7B, Qwen2.5-Max, etc. One expert sub-model can specifically process text data, and another expert sub-model can process image data. Currently, feasible mixture of experts models can improve the flexibility and adaptability of the model by introducing the mixture of experts mechanism.

[0125] Model and Hardware Node Parameter Settings. According to Table 1, summarize the labels and parameters of each part of the trained multi-modal mixture-of-experts model, and at the same time collect the relevant labels and parameters of edge hardware resources as the basis for the next model placement prediction; initialize the edge device relationship graph neural network and assign values according to each node type and feature to generate multiple model nodes and hardware nodes, where the model nodes include data nodes, routing nodes, expert nodes, and output nodes; set the constraints of the hardware nodes based on the feature values in each node; after sorting the expert nodes in a random and linear sequence, use the heuristic enumeration method to combine and allocate the model nodes and hardware nodes to generate multiple candidate placement schemes for the mixture-of-experts model; judge the effectiveness of the candidate placement schemes based on the hardware constraints, filter out the infeasible candidate placement schemes, and select multiple feasible candidate placement schemes.

[0126] Construct a Graph Neural Network and Edge Model Placement Prediction Algorithm. Construct a graph neural network. According to the resource status of edge devices and the computational requirements of the model, use the edge model placement prediction algorithm in the present invention. Take edge devices as nodes in the graph and the communication links between devices as edges to capture the topological relationships and data dependencies between edge devices to obtain the optimal prediction scheme, and reasonably allocate different layers and expert models in the multi-modal mixture-of-experts model to different edge devices. The algorithm will consider factors such as the computing power, storage capacity, and network bandwidth of the devices and make decisions through the node features of the graph neural network.

[0127] Use the feature embedding method to map the general feature data set to a continuous vector space to obtain a set of feature vectors;

[0128] Based on the new message passing mechanism, according to the candidate placement scheme, transfer the corresponding feature vectors from the expert nodes, data nodes, routing nodes, and output nodes to the hardware nodes assigned to them to transfer feature information;

[0129] Then perform the reverse transfer of feature information from the hardware nodes to the associated nodes;

[0130] Perform the transfer of feature information in the data flow order of processing modal tokens during each run of the expert model;

[0131] Based on this, construct edges according to the relationships between each node to obtain the edge device relationship graph neural network.

[0132] Establish a cost function according to the resource status of edge devices and the computational requirements of the model, and iteratively train the candidate placement schemes for the edge device relationship graph neural network with the goal of minimizing the cost function value to obtain the optimal placement scheme for the mixture-of-experts model.

[0133] In summary, through the above specific embodiments, the present invention can effectively solve the placement problem of the multi-modal mixture of experts model in the edge environment, improve the running speed and load balance of the model in resource-constrained environments, and ensure the successful deployment and stable operation of the model.

[0134] In this embodiment, a graph neural network (GNN) is applied to solve the problem of how to place a multi-modal mixture of experts model into multiple weak-resource edge devices under zero-shot conditions. A graph neural network is a deep learning model specifically designed for processing graph-structured data, which consists of nodes and edges. Each node or edge contains the feature data of the entity it represents. Multiple nodes and edges can represent the relationships and interactions between entities. By performing information passing and aggregation on the graph-structured data, the hidden information of the nodes and edges in the graph structure is automatically passed and aggregated along the connected edges, so as to represent the mutual dependence relationships by combining the connections of all adjacent nodes and edges of the nodes, and update the global feature information of the nodes in the graph structure. Specifically, it includes:

[0135] General feature extraction and fusion: Considering the edge hardware parameters and multi-modal data comprehensively, a data set containing general features (such as the number of CPUs, transmission bandwidth, etc.) is generated to help the model select the best placement scheme under zero-shot conditions, providing a basis for the subsequent deployment and operation of the model.

[0136] Heuristic scheme generation: Multiple candidate placement schemes are generated through heuristic enumeration, and the feasibility of the schemes is judged based on hardware constraints, which broadens the scope of scheme selection and ensures that the schemes are implementable in the actual hardware environment, effectively improving the success rate of model placement and the level of performance optimization, enabling the model to better adapt to the edge computing environment.

[0137] New message passing mechanism: Using the feature data set and candidate placement schemes to construct an overall graph structure, an innovative new message passing mechanism is adopted. The expert model and the corresponding edge devices are set as independent nodes with feature attributes, and feature information is passed according to the candidate schemes and data execution processes, enabling the model to better capture the complex relationships between nodes and providing strong support for accurate prediction.

[0138] Embodiment 2

[0139] This embodiment provides an edge model scheduling system integrating a graph neural network and a mixture of experts model, including:

[0140] A mixture of experts model generation module, configured to preprocess multi-modal data and generate a multi-modal mixture of experts model based on the preprocessed multi-modal data;

[0141] A feature splicing module, configured to obtain the edge device hardware parameters and splice the multi-modal data features to generate a general feature data set;

[0142] A feasible candidate placement scheme determination module, configured to enumerate and generate multiple candidate placement schemes of the mixture of experts model based on a general feature dataset by using a heuristic algorithm, and screen out multiple feasible candidate placement schemes according to the edge device hardware constraints;

[0143] A graph neural network construction module, configured to set the expert model and the corresponding edge device as independent nodes with feature attributes by using a feature embedding method and a message passing mechanism based on a general feature dataset, and establish edge connections according to the feasible candidate placement schemes to construct an edge device relationship graph neural network;

[0144] An optimal placement scheme determination module, configured to establish a cost function according to the resource status of the edge device and the calculation requirements of the model, and iteratively train the candidate placement schemes of the edge device relationship graph neural network with the goal of minimizing the cost function value to obtain the optimal placement scheme of the mixture of experts model.

[0145] The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the first embodiment above. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer executable instructions.

[0146] In the above embodiments, the descriptions of each embodiment have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0147] The proposed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the above module division is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0148] Embodiment Three

[0149] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps in the edge model scheduling method for fusing a graph neural network and a mixture of experts model as described in the first embodiment above are implemented.

[0150] Embodiment Four

[0151] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps in the edge model scheduling method for fusing a graph neural network and a mixture of experts model as described in the first embodiment above are implemented.

[0152] Embodiment Five

[0153] This embodiment provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the edge model scheduling method of the fusion graph neural and mixture of experts model described in the first embodiment above.

[0154] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.

[0155] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0156] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0157] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0158] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0159] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that based on the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.

Claims

1. The edge model scheduling method integrating graph neural network and hybrid expert model is characterized by: include: Preprocess the multimodal data, and generate a multimodal hybrid expert model based on the preprocessed multimodal data; the processing process of the hybrid expert model is specifically as follows: Obtain multimodal data for standardization and feature extraction, and use feature embedding to map multimodal data into a unified feature space; Generate modal tokens corresponding to multiple modalities from the unified feature space to obtain multiple modal tokens; Using gated routing, the expert models are sorted in descending order of their weights, and multiple modal tokens are assigned to a corresponding number of expert models with high weight values ​​and not fully occupied for processing; The output results of multiple expert models are fused with the multiplication results of their weights and the original modal tokens to obtain a fusion result, and then the fusion result is normalized and decoded as the final output result of the hybrid expert model; Obtain edge device hardware parameters and multimodal data feature splicing to generate a general feature data set; Based on the common feature dataset, a heuristic algorithm is used to enumerate and generate multiple candidate placement schemes of the hybrid expert model, and multiple feasible candidate placement schemes are screened out according to the hardware constraints of the edge device; Based on the common feature dataset, the expert model and the corresponding edge devices are set as independent nodes with feature attributes by using feature embedding method and message passing mechanism, and edge connections are established according to feasible candidate placement schemes to construct the edge device relationship graph neural network. A cost function is established according to the resource status of the edge device and the computing requirements of the model. The edge device relationship graph neural network is iteratively trained to find candidate placement solutions with the goal of minimizing the cost function value, and the optimal placement solution of the hybrid expert model is obtained.

2. The edge model scheduling method of fusion graph neural and hybrid expert model as claimed in claim 1, characterized in that: Based on the general feature data set, a heuristic algorithm is used to enumerate and generate multiple candidate placement schemes of the hybrid expert model, and multiple feasible candidate placement schemes are screened out according to the hardware constraints of the edge device, specifically: Initialize the edge device relationship graph neural network and assign values ​​according to the type and characteristics of each node to generate multiple model nodes and hardware nodes, where the model nodes include data nodes, routing nodes, expert nodes, and output nodes; Setting constraints on hardware nodes based on the eigenvalues ​​in each node; After sorting the expert nodes in random and linear sequence, the model nodes and hardware nodes are combined and allocated using a heuristic enumeration method to generate multiple candidate placement schemes for the hybrid expert model; The validity of candidate placement solutions is judged based on hardware constraints, infeasible candidate placement solutions are filtered out, and multiple feasible candidate placement solutions are screened out.

3. The edge model scheduling method of fusion graph neural and hybrid expert model as claimed in claim 2, characterized in that: The placement process of hardware nodes is combined and allocated with model nodes according to the following three rules: Each hardware node receives at least one model node; Hardware nodes are arranged from weak to strong based on resources and performance; During the placement process, model nodes are placed into hardware nodes in the order of data flow in a line.

4. The edge model scheduling method of fusion graph neural and hybrid expert model as claimed in claim 1, characterized in that: The construction process of the edge device relationship graph neural network is specifically as follows: The feature embedding method is used to map the general feature data set into a continuous vector space to obtain a feature vector set; According to the candidate placement scheme, the corresponding feature vector is transmitted from the expert node, the data node, the routing node and the output node to the hardware node assigned to it; Then reversely transmit the characteristic information from the hardware node to the associated node; The feature information is passed in the order of the data flow in which the modal tokens are processed in each expert model run; In this way, edges are constructed according to the relationship between each node to obtain the edge device relationship graph neural network.

5. The edge model scheduling method of fusion graph neural and hybrid expert model as claimed in claim 1, characterized in that: The cost function includes the amount of data processed by the expert model per unit time, the load balance of the edge nodes, and the communication time of the expert model in processing the modality token.

6. The edge model scheduling system integrating graph neural network and hybrid expert model is characterized by: include: The hybrid expert model generation module is configured to preprocess the multimodal data and generate a multimodal hybrid expert model based on the preprocessed multimodal data; the processing process of the hybrid expert model is specifically as follows: Obtain multimodal data for standardization and feature extraction, and use feature embedding to map multimodal data into a unified feature space; Generate modal tokens corresponding to multiple modalities from the unified feature space to obtain multiple modal tokens; Using gated routing, the expert models are sorted in descending order of their weights, and multiple modal tokens are assigned to a corresponding number of expert models with high weight values ​​and not fully occupied for processing; The output results of multiple expert models are fused with the multiplication results of their weights and the original modal tokens to obtain a fusion result, and then the fusion result is normalized and decoded as the final output result of the hybrid expert model; A feature stitching module is configured to obtain edge device hardware parameters and multimodal data feature stitching to generate a general feature data set; A feasible candidate placement scheme determination module is configured to enumerate and generate multiple candidate placement schemes of a hybrid expert model based on a common feature data set using a heuristic algorithm, and screen out multiple feasible candidate placement schemes according to hardware constraints of edge devices; The graph neural network building module is configured to use a feature embedding method and a message passing mechanism to set the expert model and the corresponding edge device as independent nodes with feature attributes based on a common feature data set, and to establish edge connections according to feasible candidate placement solutions to build a graph neural network for edge device relationships; The optimal placement plan determination module is configured to establish a cost function based on the resource status of the edge device and the computing requirements of the model, iteratively train the edge device relationship graph neural network to obtain the candidate placement plan with the goal of minimizing the cost function value, and obtain the optimal placement plan of the hybrid expert model.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in the edge model scheduling method of the fusion graph neural and hybrid expert model as described in any one of claims 1 to 5 are implemented.

8. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the edge model scheduling method of fusing graph neural and hybrid expert models as described in any one of claims 1-5 are implemented.

9. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps in the edge model scheduling method of fusing graph neural and hybrid expert models as described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Twin network performance prediction method and device based on graph neural network, and storage medium

    CN119299327A

  • Virtual power plant collaborative optimization scheduling method and system based on deep reinforcement learning

    CN119494521A