Elastic cooperative reasoning method in dynamic Internet of Vehicles environment

By introducing a flexible collaborative inference method in the dynamic Internet of Vehicles environment, combining the dual-path parallel architecture of task offloading and local computing, the low latency and high reliability requirements of DNN inference services in the dynamic Internet of Vehicles environment are solved, and efficient and reliable inference services are achieved.

CN120018210AActive Publication Date: 2025-05-16CHONGQING UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510044562.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-12
Publication Date
2025-05-16
Estimated Expiration
2045-01-12

AI Technical Summary

Technical Problem

In dynamic vehicle networking environments, the dynamic nature of the vehicle environment and the reliability of edge nodes make it difficult for the prior art to provide efficient and reliable deep neural network (DNN) inference services, especially in the case of low latency and high reliability requirements.

Method used

A flexible collaborative inference method in dynamic vehicle network environment is proposed. Through adaptive control of load segmentation, division and unloading, combining task offloading and local computing, the real-time and reliability of DNN inference tasks are optimized.

Benefits of technology

It realizes the provision of low-latency and high-reliability DNN inference services in a dynamic Internet of Vehicles environment, which can quickly recover in the event of uninstallation failure, ensure the reliability of the system, and make full use of idle resources of edge devices to reduce overall inference delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120018210A_ABST
    Figure CN120018210A_ABST
Patent Text Reader

Abstract

The invention provides an elastic cooperative reasoning method in a dynamic Internet of Vehicles environment, which is based on a reasoning scene of an edge acceleration deep neural network in the dynamic Internet of Vehicles environment, and combines a cooperative reasoning time delay model established by a direction propagation characteristic, dynamic Internet of Vehicles environment characteristics and edge node reliability. Solving an edge node optimal distribution strategy and an optimal distribution scheme of reasoning block segmentation and load division; then, the client vehicle decomposes the model into reasoning blocks, and the workload of each reasoning block is refined into smaller independent subtasks; the subtasks are unloaded to the selected edge nodes for parallel processing; meanwhile, the client vehicle calculates the same reasoning block locally; and the client vehicle iteratively executes the process on each reasoning block until reasoning of the whole model is completed. By adopting the method provided by the invention, load segmentation, division and unloading in the dynamic vehicle-mounted network can be adaptively controlled, and low-delay and high-reliability service is provided for an intelligent network-connected vehicle reasoning task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet of Vehicles edge computing, and in particular to a flexible collaborative reasoning method in a dynamic Internet of Vehicles environment. Background Art

[0002] The rapid development of Industrial Cyber-Physical Systems (ICPSs) is reshaping multiple fields such as smart factories, intelligent transportation systems, and intelligent connected vehicles. Among them, the Internet of Vehicles, as a typical application of ICPSs, is gradually becoming a focus. In the Internet of Vehicles, vehicles equipped with a variety of sensors, communication modules, and computing units interact with edge infrastructure and cloud centers through V2X (Vehicle-to-Everything) communication to achieve information sharing and collaborative processing.

[0003] At the same time, the development of deep neural networks (DNN) has demonstrated excellent performance in high-level feature extraction and content processing. Embedding DNN models in connected vehicle devices can support many intelligent applications, such as augmented reality navigation and autonomous driving. However, with the continuous growth of network model scale and the continuous increase in computing requirements, the efficient deployment of DNN on resource-constrained vehicle devices faces huge challenges, especially when the demand for low latency and high reliability becomes increasingly prominent. Traditional methods can no longer meet these requirements.

[0004] To improve the efficiency of DNN task execution, cooperative inference (CI) is developing rapidly. Its main collaborative modes can be divided into two modes: model collaboration and data collaboration. Model collaboration optimizes resource utilization efficiency by splitting the DNN model into layers or blocks, and unloading the split parts to edge nodes (such as intelligent connected vehicles or roadside equipment) for step-by-step execution. In contrast, data collaboration divides the reasoning task into smaller loads based on the spatial independence of convolution operations, and unloads them to multiple peer devices for parallel execution, thereby significantly accelerating the reasoning process.

[0005] Although the above technologies have substantially improved task execution efficiency, they still face many challenges in deployment in actual IoVs: the dynamic nature of the vehicle environment, such as high vehicle mobility, heterogeneous computing resources, and occasional edge node downtime, will lead to variable service performance, which potentially damages reasoning efficiency and functional safety; moreover, most existing technologies are oriented towards stable and reliable edge computing environments, resulting in a one-time fixed decision-making process that is not suitable for dynamic IoV environments. At the same time, existing technologies ignore offload fault protection strategies and cannot quickly restore services in the event of an offload failure. Summary of the invention

[0006] Based on this, the present invention provides a flexible collaborative reasoning method in a dynamic vehicle networking environment to adaptively control load segmentation, division and unloading in the dynamic vehicle network, ensure the real-time and reliability of DNN reasoning tasks, and provide low-latency and high-reliability services for intelligent connected vehicle reasoning tasks.

[0007] In order to achieve the above object, the present invention provides a flexible collaborative reasoning method in a dynamic Internet of Vehicles environment, comprising the following steps:

[0008] S100. Based on the edge accelerated DNN inference scenario in a dynamic Internet of Vehicles environment, a collaborative inference delay model is established considering the forward propagation characteristics of DNN inference. Combined with the characteristics of the dynamic Internet of Vehicles environment and the reliability of edge nodes, the optimal allocation strategy of edge nodes and the optimal allocation scheme of inference block segmentation and load division are solved.

[0009] S200. Based on the assigned edge nodes and offloading strategy, the client vehicle decomposes the model into inference blocks and further refines the workload of each inference block into smaller independent subtasks. These subtasks are offloaded to the selected edge nodes for parallel processing. At the same time, the client vehicle calculates the same inference block locally to quickly restore service when offloading fails. The specific process is as follows:

[0010] Task offloading path: The client vehicle partitions the input data or intermediate features and sends the required input for each inference block to the selected edge node. The edge node completes the calculation of each subtask in this block and returns the output to the client vehicle, and the same operation is performed for each inference block.

[0011] Local calculation path: The client vehicle calculates the same reasoning block locally at the same time. If the output of the offload path is successfully received, the local calculation is terminated and the next reasoning block is processed; otherwise, the local calculation will complete the subsequent reasoning block.

[0012] S300. The client vehicle iteratively executes step 2 for each inference block until the inference of the entire model is completed.

[0013] Furthermore, the following minimization problem is used to minimize the average collaborative reasoning delay in the overall optimization task:

[0014]

[0015] Among them, 1, 2, …, T is the inference task sequence generated by the client vehicle. is the set of allocated edge nodes, is the reasoning block partitioning strategy, For offloading strategies (including allocated edge nodes, inference block segmentation and its inference load partitioning and allocation scheme), CIDt is the collaborative reasoning latency of task t, m is the DNN model deployed on the customer vehicle, is the kth inference block of the model starting at layer i and ending at layer j, is the execution delay of the inference block in the task offloading path, is the execution delay of the inference block in the local computation path of the client vehicle v, is the number of output rows of the last layer of the model, A collection of layer subscripts for the model.

[0016] The CID is obtained by a collaborative reasoning delay model.

[0017] Furthermore, the collaborative reasoning delay CID is calculated as follows:

[0018]

[0019] The collaborative inference delay is composed of all block inference delays The sum is calculated.

[0020] Furthermore, the reasoning block Inference latency The calculation is as follows:

[0021]

[0022] The inference latency of the inference block depends on the local computation path and the task offloading path, where I(·)∈{0,1} is used to determine whether the offloading process of the edge node is successful.

[0023] Furthermore, local computing latency The reasoning process is as follows:

[0024]

[0025] where c v is the computing power of the customer vehicle v, f k For the inference block The computational overhead.

[0026] Furthermore, the unloading delay The reasoning process is as follows:

[0027]

[0028] in, The upload latency of the input data for the inference block, Processing latency for edge nodes, Aggregate latency for the output results at the end of the inference block, which are calculated as follows:

[0029] Model the DNN network structure as a directed acyclic graph vertex represents an inference layer of DNN, such as convolution, pooling, etc., and the edge e i =(l i ,l j ) represents the data flow between inference layers.

[0030] For a specific DNN layer l i , define its kernel size, padding and step size as k i ,p i ,s i . Define the input data size of each layer as Represent the height, width and number of channels of the input matrix respectively. Similarly, the output data size is defined as Represent the height, width and number of channels of the output matrix respectively.

[0031] Since each element in the output layer can be calculated by sliding the kernel window on the input layer, according to the corresponding properties, the height of the output matrix is:

[0032]

[0033] The same formula applies to the width of the output matrix.

[0034] The computational cost of the layer f i It can be quantified by floating point operations (FLOPs):

[0035]

[0036] Data size between layers e i for:

[0037]

[0038] in The memory usage of the unit data.

[0039] Because the reasoning of each layer is not an atomic operation, it can be divided into multiple smaller reasoning fragments and assigned to multiple edge nodes for parallel execution. Assuming that the layers are divided along the height, a matrix row in the layer is regarded as the smallest division unit. Define layer l i The inference fragment of the output matrix of is Given an expected output matrix of an inference fragment, according to the sliding window principle of convolution and pooling, the corresponding required input data range is in:

[0040]

[0041] Define v as the customer vehicle, is the set of available edge nodes, and the set of nodes selected to participate in offloading is The computing power of customer vehicle v and edge node n are c v and c n , the V2X transmission bandwidth is B.

[0042] Selection-based node collection Assuming a specific DNN model It can be split into a series of inference blocks containing multiple DNN layers. The set of inference blocks for segmentation is defined as:

[0043]

[0044] in, represents the kth inference block starting at the i-th layer and ending at the j-th layer, Indicates the last level subscript. Because blocks do not intersect, the start and end subscripts should usually meet the condition

[0045] Based on the assigned edge node set The DNN model is divided into Inference blocks, each of which It is further divided into multiple reasoning segments, and the offloading strategy is defined as a dictionary

[0046] Among them, it is represented as edge node n is responsible for executing the inference block The inference load of part , and the output matrix Sent to customer vehicle v.

[0047] Assume that all available edge nodes share a spectrum with a total bandwidth of B for V2X transmission through orthogonal frequency division multiplexing (OFDM), and the client vehicle allocates bandwidth equally to each selected node. The transmission rate between the client vehicle and the edge node is expressed as:

[0048]

[0049] Where P represents the transmission power, σ 2 represents the received noise power, g vn represents the channel gain between vehicle v and edge node n, which can be expressed as is the small-scale Rayleigh fading parameter, is the large-scale fading coefficient.

[0050] Furthermore, the upload delay of input data The calculation method is as follows:

[0051]

[0052] in, Represents the input matrix height of the inference block, r vn is the transmission rate between client vehicle v and edge node n.

[0053] Related research shows that for a specific layer and device, the communication and computation overhead is proportional to the data size. Based on this, the latency of the edge node processing the inference task block is The calculation method is as follows:

[0054]

[0055] Furthermore, the output result aggregation delay at the end of the inference block is The calculation method is as follows:

[0056]

[0057] Furthermore, in the overall optimization task, an online learning scheduling algorithm based on dynamic programming is used to determine the optimal allocation strategy for edge nodes and the best allocation scheme for inference block segmentation and load partitioning.

[0058] Furthermore, the optimal allocation strategy for edge nodes and the best allocation scheme for reasoning block segmentation and load partitioning are determined according to the following strategies.

[0059] For the allocated edge node set The capability of the edge node n is defined as:

[0060]

[0061] Among them, c n is the computing power of edge nodes, is the unloading success rate of edge node n observed by the client vehicle, is a monotonically decreasing function in the range [1,+∞).

[0062] Within each inference block, considering the heterogeneity of the computing power of the selected edge nodes, the output matrices are distributed in the following way to balance the workload and ensure that they return the corresponding outputs at approximately the same time.

[0063]

[0064] Define the success rate of edge node n returning the assigned reasoning fragment within a specific time threshold as Pr(n), then define the reasoning block Uninstall success rate for:

[0065]

[0066] When the kth inference block is successfully executed in the offload path, the delay is the sum of its execution delay in the offload path and the expected delay of subsequent inference blocks; when the kth inference block fails to execute in the offload path, the kth and subsequent inference blocks will be executed locally. The recursive formula for the expected delay is:

[0067]

[0068] Taking into account the scenarios of successful and failed offloading, combined with the offloading success rate binary tree, the general expression of the expected delay is:

[0069]

[0070] Define I(·) as an indicator function to indicate whether the uninstallation process of the edge node is successful, and define the uninstallation success rate of the edge node n after update:

[0071]

[0072] Based on this, follow the steps below to determine the optimal allocation strategy for edge nodes and the best allocation scheme for inference block segmentation and load partitioning:

[0073] S500. Initialization: Since the actual value of the unloading success rate Pr(n) cannot be known in advance, an estimated value is used It represents the unloading success rate of the edge node observed by the client vehicle, and is initially set to

[0074] S600. Dynamic programming: sort all available edge nodes in descending order according to θ(n), and then sort them from 1 to Traverse top_N, iteratively and incrementally select top_N edge nodes according to the sorted queue, then determine the inference block segmentation scheme based on dynamic programming, and determine the load division and distribution scheme within the inference block according to the computing power of the edge nodes. Finally, merge some inference blocks according to the offloading success rate of each edge node, in order to obtain the offloading scheme with the lowest inference latency and the highest offloading success rate.

[0075] S700. Local calculation: For each inference block, the client vehicle executes the task offloading path and the local calculation path simultaneously, and updates the offloading success rate of each edge node. If the result of the offloading path is received before the local calculation is completed, the inference block is considered completed, the local calculation is terminated, and the initial offloading strategy is continued to execute subsequent inference blocks. If the system dynamics cause the offloading path to fail, the local calculation completes the block, and the subsequent inference blocks are rescheduled as new tasks until they are completed.

[0076] Compared with the existing technology, the technical advantages of the elastic collaborative reasoning method in a dynamic Internet of Vehicles environment are at least reflected in:

[0077] First, in a dynamic Internet of Vehicles environment, the high mobility of vehicles, heterogeneous computing resources, and occasional edge node downtime may affect reasoning efficiency and functional safety. The present invention introduces a dual-path parallel architecture of task offloading and local computing to interactively complete reasoning tasks and deliver final results. While improving computing efficiency, it can achieve rapid recovery in the event of offloading failure to ensure system reliability.

[0078] Secondly, in order to make full use of the idle resources of edge devices to reduce end-to-end reasoning latency and maximize service quality, the present invention adaptively splits the DNN model and divides the reasoning load and offloads it to the edge node for processing, thereby minimizing the overall unloading latency. At the same time, some DNN reasoning blocks are adaptively merged to reduce the number of unloading stages, thereby improving the success rate of unloading tasks and optimizing the overall reasoning performance.

[0079] Thirdly, in order to solve the problem of unknown reliability of edge nodes, the present invention integrates online learning methods into the scheduling process and uses a binary tree model to evaluate the offloading success rate of reasoning tasks. This method not only ensures real-time performance, but also takes reliability into account, effectively making up for the shortcomings of existing technologies and providing low-latency and high-reliability services for intelligent connected vehicle reasoning tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] The accompanying drawings, which constitute a part of the specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0081] Figure 1 It is a workflow diagram of a flexible collaborative reasoning method in a dynamic Internet of Vehicles environment;

[0082] Figure 2 It is a schematic diagram of a task scenario of an embodiment of a flexible collaborative reasoning method in a dynamic Internet of Vehicles environment;

[0083] Figure 3It is a schematic diagram of an embodiment of a flexible collaborative reasoning method in a dynamic Internet of Vehicles environment;

[0084] Figure 4 is a schematic diagram of the inference block segmentation and load division mechanism based on dynamic programming adopted in the provided embodiment;

[0085] Figure 5 is an example diagram of potential offloading strategies after the inference blocks are merged in the provided embodiment;

[0086] Figure 6 is a schematic diagram of a binary tree-based expected inference delay evaluation method in the provided embodiment;

[0087] Figure 7 It is a schematic diagram of reasoning rescheduling under unloading failure in the provided embodiment;

[0088] Figure 8 It is a flowchart for determining the optimal allocation strategy for edge nodes and the best allocation scheme for reasoning block segmentation and load partitioning. DETAILED DESCRIPTION

[0089] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. The description of the exemplary embodiments is merely illustrative and is by no means intended to limit the present disclosure and its application or use. The present disclosure can be implemented in many different forms, not limited to the embodiments described herein.

[0090] like Figure 1 As shown, the present invention provides a flexible collaborative reasoning method in a dynamic vehicle networking environment, the steps comprising:

[0091] S100: Based on the inference scenario of edge accelerated deep neural network in a dynamic IoV environment, the collaborative inference delay model established by propagation characteristics, dynamic IoV environment characteristics and edge node reliability are combined to solve the optimal allocation strategy of edge nodes and the optimal allocation scheme of inference block segmentation and load partitioning;

[0092] S200: Based on the assigned edge nodes and offloading strategy, the client vehicle decomposes the model into inference blocks, and further refines the workload of each inference block into smaller independent subtasks; the subtasks are offloaded to the selected edge nodes for parallel processing; at the same time, the client vehicle calculates the same inference block locally to quickly restore service when offloading fails;

[0093] S300: The client vehicle iteratively executes S200 for each inference block until the inference of the entire model is completed.

[0094] The present invention proposes a flexible collaborative reasoning method in a dynamic Internet of Vehicles environment, which is a flexible accelerated reasoning mechanism that takes into account both reasoning efficiency and functional safety. This mechanism completes reasoning tasks and delivers final results in an interactive manner by introducing a dual-path parallel architecture of task offloading and local computing. While improving computing efficiency, it can achieve rapid recovery in the event of offloading failure to ensure system reliability. In addition, the present invention minimizes the overall offloading delay by adaptively splitting the DNN model and dividing the inference load and offloading it to the edge node for processing. At the same time, some DNN reasoning blocks are adaptively merged to reduce the number of offloading stages, thereby improving the success rate of offloading tasks and optimizing the overall reasoning performance.

[0095] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the elastic collaborative reasoning method in a dynamic Internet of Vehicles environment proposed by the present invention will be further described in detail below with reference to the accompanying drawings.

[0096] As Figure 2 The illustrated vehicle networking environment is used to exemplify the present invention. The vehicle that generates the DNN reasoning task such as classification or detection is defined as the client vehicle, and other peer vehicles and infrastructure such as roadside units (RSUs) are defined as edge nodes. In general, the client vehicle is defined as v, and the edge node set is defined as The client vehicle adaptively divides the DNN model and unloads the inference load to the edge node for processing through V2X communication. At the same time, in order to ensure rapid response to the inference service in the event of unloading failure, the present invention integrates an additional execution path for active inference on the local computing platform. Finally, the entire model inference is completed through the parallel execution of the task offloading path and the local computing path.

[0097] This embodiment adopts the following DNN reasoning delay model: During the implementation process, the edge accelerated DNN reasoning mechanism is basically as follows: Figure 3 As shown, the DNN network structure is modeled as a directed acyclic graph vertex represents an inference layer of DNN, such as convolution, pooling, etc., and the edge e i =(l i ,l j ) represents the data flow between the inference layers. Related studies have shown that for any complex DNN structure, the feature extraction stage occupies most of the total computing time. In addition, the client vehicle will process the output of the feature extraction stage locally to obtain the final results of classification, detection or other downstream tasks. Therefore, this embodiment only considers the more general feature extraction stage in model inference.

[0098] For a specific DNN layer l i, define its kernel size, padding and step size as k i ,p i ,s i , the input data size of each layer is Represent the height, width and number of channels of the input matrix respectively. Similarly, the output data size is defined as Represent the height, width and number of channels of the output matrix respectively.

[0099] Since each element in the output layer can be calculated by sliding the kernel window on the input layer, according to the corresponding properties, the height of the output matrix is:

[0100]

[0101] The same formula applies to the width of the output matrix.

[0102] Layer i The computational cost of i It can be quantified by floating point operations (FLOPs):

[0103]

[0104] Data size between layers e i for:

[0105]

[0106] in The memory usage of the unit data.

[0107] In the implementation, the reasoning of each layer is divided into multiple smaller reasoning fragments and assigned to multiple edge nodes for parallel execution. Assuming that the layers are divided along the height, a matrix row in the layer is regarded as the smallest division unit. Define layer l i The matrix output of a partitioned reasoning fragment is The required input data range is According to the sliding window principle of convolution and pooling, it is as follows:

[0108]

[0109] Define the set of edge nodes selected to participate in offloading as The computing power of customer vehicle v and edge node n are c v and c n , the available V2X transmission bandwidth is B.

[0110] In a dynamic IoV environment, the available edge nodes and their operating states may change over time. This example assumes a specific DNN model It can be split into a series of inference blocks containing multiple DNN layers, and the set of segmented inference blocks is defined as

[0111] in represents the kth inference block starting at the i-th layer and ending at the j-th layer, Indicates the last level subscript. Because blocks do not intersect, the start and end subscripts should usually meet the condition

[0112] Based on the assigned edge node set The DNN model is divided into Inference blocks, each of which Can be further divided into multiple reasoning segments, defining the offloading strategy as a dictionary It is represented as edge node n is responsible for executing the inference block The inference load of part , and the output matrix Sent to the client vehicle v. According to the offloading strategy, the input matrix required for all the front layers in a certain inference block can be determined Each inference fragment is treated as a separate subtask and offloaded to an edge node for parallel inference. Specifically, the execution of an inference block on the offload path consists of three stages: the client vehicle sends the input matrix to the edge node (communication), parallel processing of the edge node (computation), and output aggregation from the edge node to the client vehicle (communication).

[0113] Assuming that all available edge nodes share a spectrum with a total bandwidth of B for V2X transmission through orthogonal frequency division multiplexing (OFDM), and the client vehicle allocates bandwidth equally to each selected node, the transmission rate r between the client vehicle and the edge node can be solved by the Shannon formula: vn In this embodiment, the calculation process is as follows but not limited to this:

[0114]

[0115] Where P represents the transmission power, σ 2 represents the received noise power, g vn represents the channel gain between vehicle v and edge node n, which can be expressed as is the small-scale Rayleigh fading parameter, is the large-scale fading coefficient, which reflects the influence of path propagation loss and log-normal shadow effect. In this embodiment, the path loss model of the V2X link is composed of 38.77+18.2log(f c )+16.7log(d vn ) indicates that f c represents the center carrier frequency (GHz), d vn represents the Euclidean distance (m) between the client vehicle and the edge node. Based on the above definition, the upload delay of the input data The calculation method is as follows:

[0116]

[0117] in, Represents the input matrix height of the inference block, r vn is the transmission rate between client vehicle v and edge node n.

[0118] The customer vehicle will slice the input of the inference load that the edge node is responsible for and then transmit it to the edge node; when the edge node receives the complete input data, it will independently execute to generate the specified output matrix. Related research shows that for specific layers and devices, the communication and computing overhead is proportional to the data size. Based on this, the latency of the edge node processing the inference task block is The calculation method is as follows:

[0119]

[0120] After the edge nodes complete the computation, they transmit the corresponding output matrix back to the client vehicle for output aggregation and repartitioning for the next inference block. The latency of transmitting the output of the inference block back to the client vehicle The calculation method is as follows:

[0121]

[0122] Only when A complete output matrix can be concatenated at the end of the inference block only when all edge nodes in the block successfully return results. Based on this, since all edge nodes execute their assigned segments in parallel, the offloading delay of this block is given by The maximum processing time between edge nodes determines the aggregation delay of the output results at the end of the inference block. The calculation method is as follows:

[0123]

[0124] On the other hand, this embodiment performs active computation on the local path to accelerate service recovery under uncertain conditions. When the customer vehicle divides the inference block and offloads it to the edge node for collaborative inference, the same inference block is processed in parallel on the local computing platform. Local computing latency The reasoning process is as follows:

[0125]

[0126] where c v is the computing power of the customer vehicle v, f k For the inference block The computational overhead.

[0127] The inference latency of the inference block depends on the local computation path and the task offloading path, and I(·)∈{0,1} is used to indicate whether the offloading process of the edge node is successful. Inference latency The calculation is as follows:

[0128]

[0129] Collaborative inference delay CID is expressed as the inference delay of all blocks sum:

[0130]

[0131] Based on the above model, this embodiment adopts Figure 3 The method flow shown is used to optimize the allocation scheme of edge accelerated inference tasks.

[0132] Assumptions is the unloading strategy of customer vehicle v, where is the set of allocated edge nodes, is the reasoning block partitioning strategy, Represented as edge node n responsible for executing the inference block And the output matrix Sent to the client vehicle v. The proposed task offloading problem is to minimize the average collaborative reasoning delay:

[0133]

[0134]

[0135] Among the above three calculation formulas, the second calculation formula indicates that the execution delay of the task offloading path should be lower than the execution delay of the local calculation path processing the same block, and the third calculation formula indicates that the number of selected edge nodes should not exceed the number of output rows of the last layer of the model, because a row of the output matrix is ​​regarded as the smallest part of the DNN partition.

[0136] In this implementation, the dual-path parallel architecture of task offloading and local computing is used to ensure both service efficiency and functional safety. First, in order to ensure service efficiency, the customer vehicle adaptively divides the DNN model according to the model characteristics and the heterogeneous resource characteristics of the edge node, and divides the inference load and offloads it to the edge node for processing, reducing the overall offloading delay. At the same time, some DNN inference blocks are adaptively merged to reduce the number of offloading stages, thereby improving the success rate of offloading tasks and optimizing the overall inference performance.

[0137] Considering that any system dynamic changes such as transmission interruption, service preemption, edge node crash, etc. will affect the task offloading process, Pr(n)∈[0,1] is used to represent the success probability of edge node n returning the allocated partition segment within a specific time threshold. It is worth noting that Pr(n) is an independent and identically distributed random variable and has an unknown value for the customer vehicle. Only when A complete output matrix can only be concatenated at the end of the block when all edge nodes in the inference block return successfully. Uninstall success rate for:

[0138]

[0139] Among them, this embodiment uses a binary tree to evaluate the unloading success rate of each inference block edge node, and uses an online learning scheduling algorithm based on dynamic programming to determine the best matching strategy for edge nodes and the best allocation scheme for inference block segmentation and load division, such as Figure 8 As shown, the specific steps are as follows:

[0140] S500. Initialization: Since the actual value of the unloading success rate Pr(n) cannot be known in advance, an estimated value is used It represents the unloading success rate of the edge node observed by the client vehicle, and is initially set to

[0141] S600. Dynamic programming, including:

[0142] First, based on ability Sort all available edge nodes in descending order, from 1 to Traverse top_N, iteratively and incrementally select top_N edge nodes according to the sorted queue, and then perform subsequent operations.

[0143] Then, without considering the success rate of edge node offloading, the inference block partitioning scheme S is determined based on dynamic programming: Figure 4 As shown, from 1 to Traverse j, then traverse i from 1 to j-1, and calculate the execution delay of the offload path of various inference block partitioning schemes Minimum offloading latency based on inference block Add inference blocks in sequence to obtain the inference block partitioning scheme S. Then determine the load partitioning distribution scheme within the inference block according to the computing power of the edge nodes to balance the workload and ensure that they return the corresponding output at approximately the same time.

[0144]

[0145] Then, the expected collaborative reasoning time E[CID(b m,k ,b m,K ], and then merge some reasoning blocks, such as Figure 5 The potential offloading strategy after the inference block is merged is shown. This embodiment uses a binary tree to calculate the offloading success rate of each edge node, such as Figure 6 As shown in Figure 1, when the kth inference block is successfully executed in the offload path, the delay is the sum of its execution delay in the offload path and the expected delay of subsequent inference blocks; when the kth inference block fails to execute in the offload path, the kth and subsequent inference blocks will be executed locally. The recursive formula for the expected delay is:

[0146]

[0147] Taking into account the scenarios of successful and failed offloading, combined with the offloading success rate binary tree, the general expression of the expected delay is:

[0148]

[0149] Finally, the offloading strategy (optimal allocation strategy for edge nodes and optimal allocation scheme for inference block segmentation and load partitioning) is obtained based on dynamic programming, and the optimal number of edge nodes combined is obtained.

[0150] S700. Local computing, including:

[0151] First, for each inference block, the client vehicle executes the task offloading path and the local computation path simultaneously.

[0152] Then, the uninstallation success rate of each edge node is updated. I(·) is defined as an indicator function to indicate whether the uninstallation process of the edge node is successful. Then the uninstallation success rate of the edge node n after update is defined as:

[0153]

[0154] Finally, if Figure 7As shown in Figure 1, if the result of the offload path is received before the local computation is completed, the reasoning block is considered completed, the local computation is terminated, and the initial offload strategy is used to execute subsequent reasoning blocks. If the offload path fails due to system dynamics, the local computation completes the block, and the subsequent reasoning blocks are re-executed as new tasks in steps 2) and 3) and scheduled until completion.

[0155] To sum up, the method in this embodiment takes into account the heterogeneity of edge node computing and communication capabilities, edge node reliability, DNN reasoning forward propagation and other characteristics, takes into account both reliability and real-time performance, effectively makes up for the shortcomings of the existing technology, and can provide low-latency and high-reliability services for intelligent connected vehicle reasoning tasks.

[0156] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A flexible collaborative reasoning method in a dynamic Internet of Vehicles environment, characterized by: S100: Based on the inference scenario of edge accelerated deep neural network in a dynamic IoV environment, the collaborative inference delay model established by propagation characteristics, dynamic IoV environment characteristics and edge node reliability are combined to solve the optimal allocation strategy of edge nodes and the optimal allocation scheme of inference block segmentation and load partitioning; S200: Based on the assigned edge nodes and offloading strategy, the client vehicle decomposes the model into inference blocks, and further refines the workload of each inference block into smaller independent subtasks; the subtasks are offloaded to the selected edge nodes for parallel processing; at the same time, the client vehicle calculates the same inference block locally to quickly restore service when offloading fails; S300: The client vehicle iteratively executes S200 for each inference block until the inference of the entire model is completed.

2. The elastic collaborative reasoning method in a dynamic vehicle networking environment according to claim 1 is characterized by: In the overall optimization task, the following minimization problem representation is used to minimize the average collaborative reasoning delay: Among them, 1, 2, …, T is the inference task sequence generated by the client vehicle. is the set of allocated edge nodes, is the reasoning block partitioning strategy, For uninstallation policy, CID t is the collaborative reasoning delay of task t, m is the deep neural network model deployed on the customer vehicle, is the kth inference block of the model starting at layer i and ending at layer j, is the execution delay of the inference block in the task offloading path, is the execution delay of the inference block in the local computation path of the client vehicle v, is the number of output rows of the last layer of the model, A collection of layer subscripts for the model.

3. The elastic collaborative reasoning method in a dynamic vehicle networking environment according to claim 2 is characterized by: The CID is obtained by the collaborative reasoning delay model, and the collaborative reasoning delay CID is calculated as follows:

4. The elastic collaborative reasoning method in a dynamic Internet of Vehicles environment according to claim 3 is characterized by: Inference Block Inference latency The calculation method is: The inference latency of the inference block depends on the local computation path and the task offloading path, where I(·)∈{0,1} is used to determine whether the offloading process of the edge node is successful.

5. The elastic collaborative reasoning method in a dynamic vehicle networking environment according to claim 4 is characterized by: Local computing latency The reasoning process is: Among them, c v is the computing power of the customer vehicle v, f k Inference Block The computational overhead.

6. The elastic collaborative reasoning method in a dynamic vehicle networking environment according to claim 4 is characterized by: Unloading delay The reasoning process is: in, The upload latency of the input data for the inference block, Processing latency for edge nodes, Aggregate latency for output results at the end of an inference block.

7. The flexible collaborative reasoning method in a dynamic Internet of Vehicles environment according to claim 6 is characterized by: Input data upload delay The calculation method is: in, Represents the input matrix height of the inference block, r vn is the transmission rate between the client vehicle v and the edge node n; The latency of edge nodes processing inference task blocks The calculation method is: Output result aggregation delay at the end of the inference block The calculation method is:

8. The elastic collaborative reasoning method in a dynamic vehicle networking environment according to claim 6 is characterized by: In the overall optimization task, the optimal allocation strategy for edge nodes and the optimal allocation scheme for inference block segmentation and load division are determined according to the following strategies: For the allocated edge node set The capability of the edge node n is defined as: Among them, c n is the computing power of edge nodes, is the unloading success rate of edge node n observed by the client vehicle, It is a monotonically decreasing function in the range of [1, +∞); Within each inference block, considering the heterogeneity of the computing power of the selected edge nodes, the output matrix is ​​distributed in the following way to balance the workload and ensure that they return the corresponding outputs at approximately the same time; Define the success rate of edge node n returning the assigned reasoning fragment within a specific time threshold as Pr(n), then define the reasoning block Uninstall success rate for: When the kth inference block is successfully executed on the offload path, the latency is the sum of its execution latency on the offload path and the expected latency of subsequent inference blocks; when the kth inference block fails to execute on the offload path, the kth and subsequent inference blocks will be executed locally.

9. The elastic collaborative reasoning method in a dynamic vehicle networking environment according to claim 6 is characterized by: The recursive formula for expected delay is: Taking into account the scenarios of successful and failed offloading, combined with the offloading success rate binary tree, the general expression of the expected delay is: Define I(·) as an indicator function to indicate whether the uninstallation process of the edge node is successful, and define the uninstallation success rate of the edge node n after update:

10. The elastic collaborative reasoning method in a dynamic Internet of Vehicles environment according to claim 9 is characterized by: The process of determining the optimal allocation strategy for edge nodes and the best allocation scheme for inference block segmentation and load partitioning includes: S500. Initialization: Since the actual value of the unloading success rate Pr(n) cannot be known in advance, an estimated value is used It represents the unloading success rate of the edge node observed by the client vehicle, and is initially set to S600. Dynamic programming: sort all available edge nodes in descending order according to θ(n), and then sort them from 1 to Traverse top_N, iteratively and incrementally select top_N edge nodes according to the sorting queue, then determine the inference block segmentation scheme based on dynamic programming, and determine the load division and distribution scheme within the inference block according to the computing power of the edge nodes, and finally merge some inference blocks according to the offloading success rate of each edge node, in order to obtain the offloading scheme with the lowest inference latency and the highest offloading success rate; S700. Local calculation: For each inference block, the client vehicle executes the task offloading path and the local calculation path simultaneously, and updates the offloading success rate of each edge node; if the result of the offloading path is received before the local calculation is completed, the inference block is considered completed, the local calculation is terminated, and the initial offloading strategy is continued to be used to execute subsequent inference blocks; if the system dynamics causes the offloading path to fail, the local calculation completes the block, and the subsequent inference blocks are rescheduled as new tasks until completion.

Citation Information

Patent Citations

  • Intelligent edge calculation method in Internet of Vehicles

    CN111124647A

  • Reliable edge accelerated reasoning task allocation method in Internet of Vehicles environment

    CN116360996A

  • Online scheduling method and device for vehicle edge collaborative deep neural network reasoning in Internet of Vehicles

    CN117042149A

  • DNN edge-end collaborative reasoning method for edge intelligence

    CN117808049A

  • Distributed processing for determining network paths

    WO2021035084A1