Quantized graph neural network collaborative reasoning method for cross-modal Internet of Things data

By constructing a cross-modal association graph and dynamically adjusting the quantization strategy, the collaborative reasoning problem of cross-modal IoT data under changing network bandwidth is solved, efficient and robust multimodal data processing is achieved, and the overall performance of the system is improved.

CN120812062APending Publication Date: 2025-10-17高忆楠
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510939394.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies cannot adapt to the dynamic changes in network bandwidth when processing cross-modal IoT data, resulting in a balance between inference accuracy and latency. They also ignore the semantic relevance between modalities, leading to information loss and performance being limited by the trade-off of the Pareto frontier.

Method used

By constructing a cross-modal association graph, monitoring network bandwidth changes in real time, dynamically adjusting the quantization bit width combination and calculating the split factor, and adopting a global optimization objective function and heuristic search algorithm, the quantization strategy is optimized to maximize the overall system performance.

Benefits of technology

It achieves efficient and robust collaborative reasoning in a dynamic network environment, ensures the stability of reasoning accuracy and latency, avoids unnecessary computing overhead, and improves the accuracy of multimodal fusion reasoning and the overall performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120812062A_ABST
    Figure CN120812062A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, in particular to a cross-modal Internet of Things data-oriented quantitative graph neural network collaborative reasoning method, which specifically comprises the following steps of: state perception and association graph construction: constructing a cross-modal association graph according to an association relationship between Internet of Things equipment, and establishing a neural network collaborative reasoning model according to the cross-modal association graph; obtaining a current network bandwidth between the edge device and the cloud server; cooperative strategy solving: responding to the change of the network bandwidth at the current moment, taking maximization of the comprehensive efficiency of the system as a target, and jointly solving to determine an optimal quantization bit width combination and calculate a segmentation factor; the cross-modal input data and the graph neural network model are quantized according to the optimal quantization bit width combination and the calculation segmentation factor, and the calculation task is deployed, so that the system does not fall into one-way sacrifice of precision or time delay when coping with dramatic change of the network environment, and the calculation efficiency is improved. However, the overall improvement of the comprehensive efficiency of the system is realized through the cooperative adjustment of the quantification strategy and the calculation load.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer technology, and in particular to a quantized graph neural network collaborative inference method for cross-modal Internet of Things data. BACKGROUND

[0002] With the popularity of Internet of Things devices, massive, heterogeneous, and cross-modal data are continuously generated at the network edge. In order to process these data in real time, graph neural networks (GNNs) are widely used in Internet of Things scenarios due to their strong graph structure data representation capabilities. Edge-cloud collaborative inference architecture deploys part of the GNN model on edge devices and the other part on cloud servers to balance inference latency and computing resource consumption.

[0003] The prior art solution faces the following fundamental contradictions when processing cross-modal Internet of Things data:

[0004] Limitations of static strategies. Traditional model quantization and model segmentation usually adopt static or quasi-static strategies. These strategies cannot adapt to highly dynamic external variables in the Internet of Things environment, especially the dramatic fluctuations in network bandwidth;

[0005] Splitting of information between modalities. There is often an inherent semantic association between cross-modal data. Existing quantization methods usually process each modality independently, ignoring this association, which may cause key joint information to be disproportionately lost during quantization, thereby severely affecting the final multi-modal inference accuracy;

[0006] Bound by the Pareto frontier. In terms of inference accuracy and end-to-end latency, the existing technology can usually only achieve a trade-off, i.e., improving one indicator will inevitably sacrifice the other, and its performance is limited to a Pareto frontier. When network bandwidth and other conditions deteriorate to the critical point of the system, the system performance will drop sharply, and robust and efficient collaborative inference cannot be achieved.

[0007] Therefore, there is an urgent need for a new method that can dynamically perceive system state, collaboratively optimize modality quantization strategies, and intelligently migrate computing boundaries to solve the above problems. SUMMARY

[0008] The present application aims to provide a quantized graph neural network collaborative inference method for cross-modal Internet of Things data, which solves the problems in the background art.

[0009] To solve the above technical problems, the present application provides a quantized graph neural network collaborative inference method for cross-modal Internet of Things data, the specific steps of which include:

[0010] Step one, state perception and association graph construction step: according to the association relationship between the Internet of Things devices, a cross-modal association graph is constructed, and the current network bandwidth between the edge device and the cloud server is obtained;

[0011] Step two, collaborative strategy solving step: in response to the change of the current network bandwidth, the optimal quantization bit width combination and the calculation split factor are determined by joint solving to maximize the system comprehensive performance;

[0012] Step three, collaborative reasoning execution step: according to the optimal quantization bit width combination and the calculation split factor, the cross-modal input data and the graph neural network model are quantized and the calculation task is deployed to complete the collaborative reasoning

[0013] Preferably, the state perception and association graph construction step specifically includes:

[0014] S11, according to the physical proximity relationship or data flow logical association between Internet of Things devices, the cross-modal association graph is constructed; wherein the node of the cross-modal association graph represents the modal data source, and the edge represents the association strength between the data sources;

[0015] S12, through the network monitor, the current network bandwidth is measured and obtained in real time.

[0016] Preferably, when the change rate of the current network bandwidth exceeds the preset trigger threshold, the collaborative strategy solving step is executed; when the change rate of the current network bandwidth does not exceed the preset trigger threshold, the existing optimal quantization bit width combination and calculation split factor are used.

[0017] Preferably, the collaborative strategy solving step further includes:

[0018] Based on the cross-modal association graph, the cross-modal quantization symbiotic cost of different modalities using a specific quantization bit width combination is calculated; wherein the cross-modal quantization symbiotic cost is used to evaluate the degree of maintaining the association information between modalities by the specific quantization bit width combination.

[0019] Preferably, the collaborative strategy solving step further includes:

[0020] Based on the quantization bit width combination, the intermediate feature tensor size transmitted from the edge device to the cloud server is determined;

[0021] Combined with the current network bandwidth, the cross-modal quantization symbiotic cost, and the intermediate feature tensor size, the calculation split factor is dynamically generated.

[0022] Preferably, the joint solving to determine the optimal quantization bit width combination and the calculation split factor includes:

[0023] construct a global optimization objective function based on prediction inference accuracy and prediction end-to-end latency;

[0024] Solve the global optimization objective function by using a heuristic search algorithm to obtain the optimal quantization bit width combination that optimizes the global optimization objective function.

[0025] Preferably, the prediction inference accuracy is generated by a pre-trained accuracy predictor; the prediction end-to-end latency is generated by a latency predictor, wherein the latency predictor combines the quantization bit width combination, the calculation split factor and the current time network bandwidth.

[0026] Preferably, the collaborative inference execution step comprises:

[0027] S31, the edge device quantizes each modality data according to the optimal quantization bit width combination, and performs partial graph neural network calculation determined by the calculation split factor to generate a quantized intermediate feature tensor;

[0028] S32, the edge device sends the quantized intermediate feature tensor to the cloud server;

[0029] S33, the cloud server receives the intermediate feature tensor and completes the remaining graph neural network calculation to output a final inference result.

[0030] Compared with the prior art, the present application has the following beneficial effects:

[0031] The present application provides a kind of quantization graph neural network collaborative inference method for cross-modal Internet of Things data, with the following beneficial effects:

[0032] (1) For the defect that the prior art static strategy cannot adapt to environmental changes, the present scheme introduces a dynamic triggering and solving mechanism based on current network bandwidth changes; The prior art usually uses a fixed quantization and segmentation strategy, and the performance will drop sharply when the network condition deteriorates; The present scheme monitors the network state in real time, and only starts the collaborative strategy solving step when the bandwidth change rate exceeds the preset threshold, dynamically re-solves the optimal quantization bit width combination and calculation split factor; This design not only realizes accurate self-adaptation to dynamic network environment, ensures stable inference performance under various network conditions, but also avoids unnecessary frequent calculation, and takes into account the responsiveness and operating economy of the system.

[0033] (2) To solve the problem of information loss caused by ignoring inter-modal correlation in the prior art, the scheme proposes a quantization cost evaluation method based on cross-modal correlation graph; the prior art treats each data modality independently during quantization, which easily destroys the joint semantic information contained between modalities; the scheme first constructs a cross-modal correlation graph according to the physical or logical relationship between devices to represent the correlation strength between data sources; on this basis, the cross-modal quantization symbiotic cost is calculated to punish the quantization scheme that destroys the consistency of information of strongly correlated modalities; this mechanism can effectively guide the optimization process and preferentially protect key joint information, thereby fundamentally improving the final accuracy of multi-modal fusion reasoning.

[0034] (3) To solve the problem of the Pareto frontier constraint of the prior art due to the trade-off between precision and latency, the scheme constructs a collaborative optimization framework with a global efficiency target; the prior art optimizes one indicator at the expense of the other; the scheme constructs a global optimization objective function that integrates the prediction reasoning accuracy and the prediction end-to-end latency, and uses a heuristic search algorithm to solve it, which converts multiple conflicting sub-problems into a unified global optimization problem; this enables the system to explore and discover solutions beyond the traditional trade-off boundary, such as significantly reducing latency by collaboratively adjusting quantization and segmentation strategies while maintaining high accuracy, achieving a overall leap in system comprehensive performance. BRIEF DESCRIPTION OF DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, a brief introduction will be given below to the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings;

[0036] Figure 1 is a logic block diagram of the system of the present application;

[0037] Figure 2 is a logic block diagram of state perception and correlation graph construction of the present application;

[0038] Figure 3 is a logic block diagram of collaborative strategy solving of the present application;

[0039] Figure 4 is a logic block diagram of joint solving of the present application. DETAILED DESCRIPTION

[0040] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, and not all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0041] Embodiment one

[0042] Please refer to Figure 1 The present application provides a kind of quantization graph neural network collaborative inference method for cross-modal Internet of Things data, specific steps include:

[0043] Step one, state perception and associated graph construction step: according to the association between Internet of Things devices Cross-modal associated graph is constructed, and the current time network bandwidth between edge device and cloud server is obtained;

[0044] Step two, collaborative strategy solving step: in response to the change of current network bandwidth, the optimal quantization bit width combination and calculation split factor are determined by solving jointly with the goal of maximizing system comprehensive performance;

[0045] Step three, collaborative inference execution step: according to the optimal quantization bit width combination and calculation split factor, cross-modal input data and graph neural network model are quantized, and calculation task is deployed to complete collaborative inference.

[0046] The embodiment provides a kind of quantization graph neural network collaborative inference method for cross-modal Internet of Things data, provides an effective technical solution for the technical challenge that inference precision and time delay are difficult to consider under dynamic network environment;In an Internet of Things application scenario deployed in intelligent factory, the method is via state perception and associated graph construction step, not only the internal connection between each sensor is clear, the current communication link quality is also mastered;On this basis, the core collaborative strategy solving step is activated, this step abandons static configuration, instead, according to the fluctuation of real-time network bandwidth, a set of collaborative strategy containing optimal quantization bit width combination and calculation split factor is dynamically generated with the goal of global system performance optimization;Finally, in collaborative inference execution step, the system strictly executes data quantization and edge, cloud deployment of calculation task according to this strategy, completes an efficient and accurate collaborative inference;This complete process constructs an adaptive closed-loop control system, and its fundamental effect is that when the network environment changes dramatically, the system is no longer trapped in the one-way sacrifice of precision or time delay, but through the collaborative adjustment of quantization strategy and calculation load, the overall improvement of system comprehensive performance is realized.

[0047] Embodiment two

[0048] Please refer toFigure 2 The state awareness and association graph construction step specifically includes:

[0049] S11, constructing a cross-modal association graph according to the physical proximity relationship or data flow logical association between Internet of Things devices; wherein the nodes of the cross-modal association graph represent modal data sources, and the edges represent the association strength between data sources, and the edge representing the connection between the two is assigned a high association strength weight. Specifically, the quantification of the association strength can adopt a discrete scoring method based on expert knowledge, for example, the association strength is divided into four levels: very strong, strong, general, and none, and is respectively mapped to dimensionless weight values w ij ∈{1.0,0.7,0.3,0}. Alternatively, the contribution of different modal data to a specific task can also be analyzed by analyzing historical data in the offline stage, and normalized as a weight.

[0050] S12, measuring and acquiring the current network bandwidth in real time through a network monitor.

[0051] In this embodiment, the state awareness and association graph construction step provides key structured input and environmental state parameters for subsequent collaborative optimization; in an autonomous vehicle system, the construction process of the cross-modal association graph defines multiple modal data sources such as front-view cameras, lidar, and inertial measurement units as nodes of the graph; given the strong data flow logical association between the front-view camera and the lidar in the task of perceiving obstacles in front, the edge representing the connection between the two is assigned a high association strength weight; at the same time, a network monitor as a standard component uses conventional network probing techniques to continuously measure and update the 5G / V2X communication bandwidth between the vehicle and the roadside unit or the cloud server, thereby acquiring the current network bandwidth B t ; through this step, the method not only establishes a static representation of the internal relationship of multi-modal data, but also realizes real-time capture of key external dynamic variables, providing an accurate data basis for dynamic decision-making and ensuring the pertinence and effectiveness of subsequent strategies.

[0052] Embodiment Three

[0053] When the change rate of the current network bandwidth exceeds the preset trigger threshold, the collaborative strategy solving step is executed; when the change rate of the current network bandwidth does not exceed the preset trigger threshold, the existing optimal quantization bit width combination and calculation segmentation factor are used.

[0054] The trigger mechanism introduced in this embodiment realizes on-demand allocation of computing resources to avoid unnecessary strategy recalculation overhead; the execution of the collaborative strategy solving step is not continuous, but is controlled by an event trigger based on the relative change of the network bandwidth; the system sets a preset relative change trigger threshold Δth , for example 10%. When the network monitor reports the current network bandwidth B t Compared to the bandwidth B when the last strategy was determined t-1 The relative rate of change Exceeding this threshold Δ th When the rate of change exceeds the threshold, the system determines that the network has changed significantly, which may cause the current strategy to fail, and activates the collaborative strategy solution step; on the contrary, if the rate of change does not exceed the threshold, it indicates that the network state is relatively stable and the current strategy is still valid. The system will continue to use the existing optimal quantization bit width combination and calculate the split factor; this design greatly improves the system's operating efficiency, allowing the system to adjust its strategy only at critical moments and maintain low-overhead operation in a stable state, thus achieving a delicate balance between responsiveness and economy.

[0055] To systematically determine the threshold, we can introduce the recalculation cost C recomp (the time or energy required to perform a collaborative strategy solution step) and strategy mismatch loss (the performance degradation caused by continuing to use the old strategy due to network changes). Optimization is triggered only when the expected loss exceeds the recalculation cost. Δ th The following principles can be followed in determining

[0056] Evaluate the sensitivity of performance to bandwidth and evaluate the global performance loss L in offline state global For network bandwidth B t The partial derivative of This value reflects the impact of bandwidth fluctuation on system performance.

[0057] Set the trigger condition, when the expected performance loss is greater than the recalculation cost, that is, , triggers policy recalculation.

[0058] Derived threshold Δ th , it can be deduced from the above conditions that the trigger threshold of bandwidth change rate is:

[0059]

[0060] When the system is deployed, the calculated C recomp and Combined with a typical bandwidth value B t , calculate a reasonable Δ th as the default value.

[0061] Example 4

[0062] Please refer to Figure 3, which aims to abstractly demonstrate the interrelation between the computational elements in this step (e.g. symbiotic cost, tensor size, splitting factor). The collaborative strategy solving step further comprises:

[0063] Based on the cross-modal association graph, a cross-modal quantization symbiotic cost is calculated when different modalities adopt a specific combination of quantization bit widths; wherein the cross-modal quantization symbiotic cost is used to evaluate the degree of preservation of the inter-modal association information by a specific combination of quantization bit widths.

[0064] The collaborative strategy solving step further comprises:

[0065] Based on the combination of quantization bit widths, the intermediate feature tensor size transmitted from the edge device to the cloud server is determined;

[0066] And combined with the current network bandwidth, the cross-modal quantization symbiotic cost, and the intermediate feature tensor size, a dynamic calculation splitting factor is generated.

[0067] The core of the embodiment is the coupling decision mechanism introduced in the collaborative strategy solving step, which deeply binds the quantization and splitting decision processes by calculating the cross-modal quantization symbiotic cost and dynamically generating the adaptive calculation splitting factor;

[0068] To evaluate a set of quantization decisions for the preservation of inter-modal association information, the method calculates the cross-modal quantization symbiotic cost C co based on the cross-modal association graph; the construction of this cost function draws on the potential energy model in physics to describe the strength of interaction, and its technical motivation is to provide a computable and optimized mathematical penalty term for the need to protect the information consistency of strongly associated modal data in the quantization process; its calculation formula is:

[0069]

[0070] Wherein, represents the total symbiotic cost when a set of quantization bit widths {q k} is given, which is a dimensionless scalar output;

[0071] is an N-dimensional vector, where the element q i represents the quantization bit width set for the i-th modality, with units of bits, and takes values in a discrete set such as {2, 4, 8, 16};

[0072] E is the edge set from the cross-modal association graph G;

[0073] w ij is the edge weight connecting nodes i and j in graph G, which is a dimensionless value quantifying the association strength between modalities i and j;

[0074] is the mutual information between modalities i and j, also in bits, measuring their statistical correlation. The mutual information is obtained as follows: in the offline phase, a feature extractor is first pre-trained for each modality separately, mapping the high-dimensional raw data into a unified dimensional feature vector f i and f j . Subsequently, a large number of pairs of feature vectors are sampled, and their joint probability distribution p(f i ,f j ) and marginal probability distributions p(f i ) and p(f j ) are estimated by a non-parametric entropy estimation method (such as k-nearest neighbor method), and finally calculated according to the mutual information standard definition . The calculation result is stored as static prior knowledge for online inference;

[0075] is a small positive number set to ensure numerical stability, with the same dimension as , to prevent the denominator from being zero;

[0076] In the strategy search, the optimization algorithm tends to choose the combination that minimizes C co , which achieves a nonlinear redistribution of information fidelity and maximizes the preservation of important joint semantic information across modal data;

[0077] The cross-modal quantified symbiotic cost formula is the starting point of the entire collaborative mechanism; its practical significance lies in that it first converts the abstract expert knowledge of maintaining the consistency of associated modal information into a calculable and optimizable quantitative indicator; the calculation result C co is a key internal state variable, which is input into the subsequent calculation segmentation decision, forming the first link of the decision chain.

[0078] On this basis, the method dynamically generates an adaptive calculation segmentation factor π t combined with the symbiotic cost and other system states; the generation function of this factor draws on the feedback regulation idea in control theory, aiming to create a unified decision function that integrates network bandwidth, quantitative strategy and harmony, and transmission data volume, these three core influencing factors; its calculation formula is

[0079]

[0080] where π tis a dimensionless scalar between (0, 1) that determines the partitioning ratio of the GNN model computation, and σ(·) is the Sigmoid activation function.

[0081] B t is the current real-time bandwidth, B ref is the reference bandwidth, both of which have the same dimension, so the ratio is dimensionless.

[0082] is the optimized minimum cross-modal quantization symbiotic cost, which is dimensionless.

[0083] is the optimal quantization combination estimated intermediate feature tensor size in bits.

[0084] The value can be obtained by training methods such as reinforcement learning before deployment. Among them, the reference bandwidth B ref can be set according to the network conditions of the target application scenario, for example, set to the average bandwidth or the expected minimum guaranteed bandwidth in the network environment. The reference tensor size S ref can be set to the size of the intermediate feature tensor generated when using the highest precision quantization (such as FP32) as a normalization benchmark.

[0085] To ensure dimensional consistency, a reference tensor size S ref is introduced for normalization, so that this term is also dimensionless; α, β, γ are three dimensionless positive weights used to adjust the sensitivity of the system to each factor, and their values can be obtained by training methods such as reinforcement learning before deployment; This multi-factor collaborative decision-making mechanism makes the computation partitioning strategy no longer a simple threshold switch, but a result of deep coupling and collaborative evolution with the quantization strategy, significantly enhancing the system's ability to adapt to complex dynamic environments and robustness.

[0086] It should be noted that the formula of the adaptive computation partitioning factor π t is different in the search process of the collaborative strategy solving step and the final application of the collaborative reasoning execution step:

[0087] In the iterative process of the heuristic search algorithm: when the algorithm evaluates a set of candidate quantization combinations {q k,i}, the symbiotic cost C co,i and the tensor size S model ({q k,i}) are used to calculate the corresponding partitioning factor π t,i . The formula does not contain an asterisk (*) at this time, indicating that it uses intermediate candidate values rather than the final optimal values:

[0088]

[0089] This π t,i With {q k,i} were sent to L global The function is evaluated and the evaluation result is used to guide the next search step of the algorithm.

[0090] When determining the optimal solution and performing reasoning: When the search process converges, the optimal quantization combination {q k} * and its corresponding optimal symbiotic cost After that, the system calculates the final calculation split factor to be deployed and executed Only then use the formula with an asterisk:

[0091]

[0092] By distinguishing these two stages, it is clear that π t The calculation logic solves the circular dependency problem in the iterative process and ensures the logical consistency of the technical solution.

[0093] Example 5

[0094] See also Figure 4 , jointly solve and determine the optimal quantization bit width combination and calculate the split factor, including:

[0095] Construct a global optimization objective function based on the predicted inference accuracy and predicted end-to-end latency;

[0096] A heuristic search algorithm is used to solve the global optimization objective function to obtain the optimal quantization bit width combination that makes the global optimization objective function optimal.

[0097] The predicted inference accuracy is generated by a pre-trained accuracy predictor; the predicted end-to-end delay is generated by a delay predictor, where the delay predictor combines the quantization bit width, calculates the split factor and the current network bandwidth for calculation.

[0098] The specific construction method of the pre-trained accuracy predictor is as follows:

[0099] Model architecture. This predictor can use a lightweight neural network, such as a multi-layer perceptron with 2 to 3 hidden layers. Its input is a feature vector concatenated by the combination of quantization bit width and the calculated split factor. To handle discrete quantization bits, each element in can be one-hot encoded.

[0100] Training data generation, in the offline phase, generate training samples by reasoning on a large amount of representative graph data. Specifically, randomly sample hundreds of different pairs of quantization combinations and partition factors, perform complete collaborative reasoning in the target edge cloud hardware environment, and record the actual reasoning accuracy. These pairs constitute the training data set of the accuracy predictor;

[0101] Training process, use standard supervised learning methods for training, for example, use mean square error as loss function, update the network weights of MLP through back propagation algorithm, until the prediction accuracy converges.

[0102] This embodiment illustrates the core mechanism of jointly solving the optimal strategy, that is, by constructing and optimizing a global objective function, the end-to-end optimization of the system final performance is realized; this method constructs a global optimization objective function L global The technical motivation is to unify the reasoning accuracy and end-to-end delay, two mutually conflicting core performance indicators, into a single framework that can be minimized; the minimization problem is expressed as

[0103]

[0104] Where L global is the global performance loss to be minimized;

[0105] {q k} and π t are variables to be solved;

[0106] Acc pred is a pre-trained reasoning accuracy predictor, whose input is a set of candidate quantization combinations {q k} and partition factor π t , and the output is the predicted value of the final reasoning accuracy;

[0107] Lat pred is a delay predictor that estimates the end-to-end delay based on {q k}, π t and the current network bandwidth B t ;

[0108] λ is a hyperparameter for balancing the importance of accuracy and delay;

[0109] The value of λ can be set according to the needs of specific application scenarios. The basic principle of its setting is: the larger the value of λ, the more the optimization target focuses on reducing the delay; the smaller the value of λ, the more the optimization target focuses on maintaining high accuracy;

[0110] For delay-sensitive applications, a larger λ value (e.g., λ>1) should be set to strongly punish the increase of delay;

[0111] For applications with stringent accuracy requirements, a smaller λ value (for example, 0<λ<1) can be set to prioritize accuracy.

[0112] For applications where accuracy and latency are equally important, λ≈1 can be set;

[0113] By providing clear lambda setting ranges and principles for different application scenarios, the debugging work required during deployment is reduced.

[0114] Since {q k The selection space of L is discrete and high-dimensional. This method uses heuristic search algorithms such as genetic algorithms to search for L global Solve;

[0115] The solution process shows the logical connection and deduction between the decision variables:

[0116] The algorithm generates a set of candidate {q k}, based on which C is calculated co , and combined with the current bandwidth B t Calculate the corresponding π t , this strategy is effective for ({q k},π t ) is then fed into the accuracy predictor Acc pred and delay predictor Lat pred To evaluate its L global ;

[0117] This iterative process continues until convergence, and the final output is L global Achieve the minimum strategy combination The fundamental technical effect of this process is that it places all interrelated decision variables under a unified goal for co-evolution, thereby being able to discover and determine the globally optimal solution under the current network environment, making the overall system performance exceed the limits of traditional trade-offs.

[0118] The minimization problem is formulated as:

[0119]

[0120] Among them, L global is the global performance loss that needs to be minimized;

[0121] {q k} is the quantization bit width combination to be solved, which is the only independent variable of the optimization problem;

[0122] π t ({q k}) clearly indicates that the calculation of the split factor is based on the independent variable {q k} and other system states (such as current bandwidth B t ) is the dependent variable calculated using the formula in .

[0123] Example 6

[0124] The collaborative reasoning execution steps include:

[0125] S31. The edge device quantizes each modal data according to the optimal quantization bit width combination, and performs a partial graph neural network calculation determined by the calculation split factor to generate a quantized intermediate feature tensor.

[0126] S32, the edge device sends the quantized intermediate feature tensor to the cloud server;

[0127] S33. The cloud server receives the intermediate feature tensor and completes the remaining graph neural network calculations to output the final inference results.

[0128] In this embodiment, the collaborative reasoning execution step is a specific stage of implementing the optimal strategy solved in the previous step; after receiving the optimal quantization bit width combination output by the collaborative strategy solving step, and calculate the split factor After that, the system enters the execution phase; the computing program deployed on the edge device strictly follows The input data of each modality is quantized to a specified bit width; after that, the edge device calculates the split factor The edge device sends this compact intermediate feature tensor to the cloud server via the communication link. After receiving the tensor, the cloud server directly executes the remaining computation layers of the graph neural network model without any additional coordination. The latter part of the computing layer is determined and ultimately outputs the reasoning results of the entire system. The execution process is clearly divided and the responsibilities are clear. Its technical effect is to accurately convert the optimized theoretical strategy into actual computing operations. Through edge-side preprocessing and data compression, it effectively reduces the occupancy of network bandwidth, while utilizing the powerful computing power of the cloud to complete complex computing tasks, ultimately ensuring the low latency and high precision of the entire edge-cloud collaborative reasoning process.

[0129] The calculated boundary is divided by the split factor Specifically, assuming that the entire GNN model contains L computing layers, the computational cost of each layer l is C l The total computational cost of the model is

[0130] The split point, i.e., the number of layers k executed on the edge device, is determined by the following rule: find the largest number of layers k such that the cumulative computation cost from layer 1 to layer k does not exceed the cost quota defined by C quota. Mathematically, this is expressed as:

[0131]

[0132] Thus, the edge device will execute layers 1 to k of the model, while the cloud server executes the remaining layers k+1 to L. The computation cost C l of each layer can be obtained once and for all in the offline phase by a model analysis tool.

[0133] The above merely describes the preferred embodiments of the present application, but not intended to limit the present application in other forms. Any person skilled in the art can make changes or modifications to the above disclosed technical contents to apply to other fields with equivalent embodiments of equivalent changes. However, any simple modification, equivalent change and modification made according to the technical essence of the present application to the above embodiments, without departing from the technical solution of the present application, still falls within the protection scope of the present application.

Claims

1. A quantitative graph neural network collaborative reasoning method for cross-modal IoT data, characterized by: The specific steps include: Step 1: State perception and association graph construction: Build a cross-modal association graph based on the association relationships between IoT devices, and obtain the current network bandwidth between the edge device and the cloud server; Step 2: Collaborative strategy solving step: In response to the change of the current network bandwidth, with the goal of maximizing the overall system performance, jointly solving and determining the optimal quantization bit width combination and calculating the splitting factor; Step 3: Collaborative reasoning execution step: Based on the optimal quantization bit width combination and the computational splitting factor, the cross-modal input data and the graph neural network model are quantized, and computing tasks are deployed to complete collaborative reasoning.

2. The quantized graph neural network collaborative reasoning method for cross-modal IoT data according to claim 1 is characterized by: The state perception and association graph construction steps specifically include: S11. Constructing the cross-modal association graph based on the physical proximity relationship or data flow logical association between IoT devices; wherein the nodes of the cross-modal association graph represent modal data sources, and the edges represent the association strength between the data sources; S12: Measure and obtain the current network bandwidth in real time through a network monitor.

3. The quantized graph neural network collaborative reasoning method for cross-modal IoT data according to claim 1 is characterized by: When the change rate of the network bandwidth at the current moment exceeds the preset trigger threshold, the collaborative strategy solving step is executed; when the change rate of the network bandwidth at the current moment does not exceed the preset trigger threshold, the existing optimal quantization bit width combination and the calculation split factor are used.

4. The quantized graph neural network collaborative reasoning method for cross-modal IoT data according to claim 3 is characterized by: The collaborative strategy solving step further includes: Based on the cross-modal association graph, the cross-modal quantization symbiosis cost is calculated when different modalities adopt a specific quantization bit width combination; wherein the cross-modal quantization symbiosis cost is used to evaluate the degree to which the specific quantization bit width combination preserves the inter-modal association information.

5. The quantized graph neural network collaborative reasoning method for cross-modal IoT data according to claim 4 is characterized by: The collaborative strategy solving step further includes: determining, based on the quantization bit width combination, a size of an intermediate feature tensor transmitted from the edge device to the cloud server; The computational splitting factor is dynamically generated by combining the current network bandwidth, the cross-modal quantization symbiosis cost, and the size of the intermediate feature tensor.

6. The quantized graph neural network collaborative reasoning method for cross-modal IoT data according to claim 5 is characterized by: The joint solution for determining the optimal quantization bit width combination and the calculation of the splitting factor includes: Construct a global optimization objective function based on the predicted inference accuracy and predicted end-to-end latency; A heuristic search algorithm is used to solve the global optimization objective function to obtain the optimal quantization bit width combination that makes the global optimization objective function optimal.

7. The quantized graph neural network collaborative reasoning method for cross-modal IoT data according to claim 6 is characterized by: The predicted inference accuracy is generated by a pre-trained accuracy predictor; the predicted end-to-end delay is generated by a delay predictor, wherein the delay predictor is combined with the quantization bit width, and the calculated split factor is calculated with the current network bandwidth.

8. The quantized graph neural network collaborative reasoning method for cross-modal IoT data according to claim 1 is characterized by: The collaborative reasoning execution step includes: S31, the edge device quantizes each modal data according to the optimal quantization bit width combination, and performs a partial graph neural network calculation determined by the calculation segmentation factor to generate a quantized intermediate feature tensor; S32. The edge device sends the quantized intermediate feature tensor to the cloud server; S33. The cloud server receives the intermediate feature tensor and completes the remaining graph neural network calculations to output the final inference result.