A fast collaborative inference method for multimodal deep learning models

By dynamically selecting the segmentation points and compression rates of the feature encoder of the multimodal deep learning model between mobile devices and edge servers, and optimizing the collaborative inference strategy using reinforcement learning algorithms, the high latency and high energy consumption problems of the multimodal deep learning model on resource-constrained devices are solved, achieving a more efficient inference process.

CN116911362BActive Publication Date: 2025-08-08XIAMEN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310718827.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-16
Publication Date
2025-08-08
Estimated Expiration
2043-06-16

AI Technical Summary

Technical Problem

The prior art cannot effectively solve the high latency and high energy consumption problems of computing-intensive multimodal deep learning models on resource-constrained mobile devices, especially for multimodal deep neural network models with branched structures.

Method used

The reinforcement learning algorithm is used to dynamically optimize the collaborative inference strategy of multimodal deep learning models. By dynamically selecting the segmentation points and compression rates of the feature encoder between mobile devices and edge servers, the segmentation points and deep learning model scale of the multimodal deep learning model are optimized, and the reinforcement learning algorithm is used to select the best segmentation points and compression rates to reduce delay and energy consumption.

Benefits of technology

Without significantly reducing the inference quality, the delay and overall energy consumption of multimodal inference services are significantly reduced, and the inference speed and energy efficiency are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116911362B_ABST
    Figure CN116911362B_ABST
Patent Text Reader

Abstract

A fast collaborative inference method for multimodal deep learning models relates to multimodal deep learning models. To address the existing problems of high latency and high energy consumption when heterogeneous multimodal deep learning networks for computationally intensive applications are deployed on resource-scarce mobile devices, a fast collaborative inference method for multimodal deep learning models is provided. A reinforcement learning algorithm is used to dynamically optimize the collaborative inference strategy for multimodal deep learning model services for mobile devices in wireless mobile edge networks. This strategy adapts to the characteristics of computationally intensive multimodal deep learning applications with multiple heterogeneous feature encoders and can reduce the latency and overall energy consumption of deep learning-based multimodal inference services without significantly reducing the inference quality. The split points of each feature encoder in the multimodal deep learning model and the scale of the deep learning model are dynamically selected to improve the speed and energy efficiency of multimodal deep learning model inference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multimodal deep learning model, and in particular to a fast collaborative inference method for a multimodal deep learning model. Background Art

[0002] Multimodal deep learning models connect and process data from multiple sensor modalities, such as audio and images, to provide inference results that provide a more comprehensive understanding of the real world and enhance the robustness of inference results. Examples include audio-video joint speech recognition and the emerging multimodal large language model. However, running computationally intensive multimodal deep learning models on resource-constrained mobile devices can lead to extended inference times and high energy consumption. By splitting the chain-like deep neural network model, some resource-intensive deep learning inference tasks can be offloaded from resource-constrained mobile devices to edge devices.

[0003] In order to effectively utilize the computing resources between the cloud and mobile devices and achieve low latency, low energy consumption, and high throughput for inference, the literature [Y. Kang, J. Hauswald, C. Gao, A. Rovinski, T. Mudge, J. Mars, and L. Tang, “Neuros urgeon: Collaborative intelligence between the cloud and mobile edge,” in Proc. ACM Int. Conf. Architectural Support Programming Languages Operating Syst (ASPLOS). Xi'an, China, April 2017, pp. 615–629.] describes a collaborative inference strategy between mobile devices and edge servers based on a regression model. A hierarchical computation partitioning strategy is used to select the optimal partitioning points to reduce mobile energy consumption and overall latency. Chinese patent 202011268445.X proposes a collaborative reasoning method for deep neural networks based on an end-edge-cloud architecture. It divides the neural network into three parts according to the network environment, resource quotas, and usage of the three parties: end-edge-cloud. It adopts an end-edge-cloud collaborative computing model to reduce the latency and energy consumption of model reasoning.

[0004] During the inference process, it is necessary to adjust the structure of the deep learning model on the mobile device to achieve a trade-off between system load and inference service quality. The literature [R.Lee, SIVenieris, L.Dudziak, S.Bhattacharya, and NDLane, “Mob isr: Efficient on-device super-resolution through heterogeneous mobile processors,” in The 25th Annual International Conference on Mobile Computing and Networking, ser. MobiCom'19. New York, NY, USA: Association for Computing Machinery, 2019.] introduces a dynamic model selection mechanism to achieve a balance between image quality and processing speed by switching between small and large networks. The paper [S. Laskaridis, S. IVenieris, M. Almeida, I. Leontiadis, and N. D. Lane, “Spinn: Synergistic progressive inference of neural networks over device and cloud,” in Proceedings of the 26th Annual International Conference on Mobile Computing and Networking, ser. MobiCom’20. New York, NY, USA: Association for Computing Machinery, 2020.] uses collaborative device cloud computing and progressive inference methods to provide fast and robust CNN inference in different settings, adapting to dynamic conditions and meeting user service level requirements, improving throughput and accuracy, and significantly saving energy. Chinese patent CN114662690A proposes a mobile device collaborative inference system for deep learning Transformer-like models. This system partitions the model and distributes multiple slices to each device. The collaborative inference process is controlled through DNS networking to address resource constraints.

[0005] Collaborative inference schemes based on reinforcement learning can better adapt to complex environments. The literature [Y.Xiao, L.Xiao, K.Wan, H.Yang, Y.Zhang, Y.Wu, and Y.Zhang, “Reinforcement learning based energy-efficient collaborative inference for mobile edge computing,” IEEE Trans.Commun., vol.71, no.2, pp.864–876, Feb.2023.] applies reinforcement learning to select deep learning model split points and edge servers, reducing the overall inference latency and energy consumption of mobile devices while maximizing the long-term expected discounted benefits of reinforcement learning. The literature [Wu, Wen and Yang, Peng and Zhang, Weiting and Zhou, Conghaoand Shen, Xuemin, “Accuracy-guaranteed collaborative DNN inference in industrial IoT via deep reinforcement learning,” IEEE Transactions on Industrial Informatics, vol. 17, no. 7, pp. 4988-4998, 2021.] comprehensively considers the adaptation of sampling rate, offloading of inference tasks, and allocation of edge computing resources, and uses Lyapunov optimization technology and deep reinforcement learning to solve the MDP, reducing the average service delay while maintaining high-probability long-term inference accuracy.

[0006] Existing collaborative inference methods based on model segmentation assume that the deep neural network model has a chain structure and are not applicable to multimodal deep neural network models with branching structures. Summary of the Invention

[0007] The purpose of this invention is to address the existing problems of high latency and high energy consumption in heterogeneous multimodal deep learning networks for computationally intensive applications when deployed on resource-constrained mobile devices. This method provides a fast collaborative inference method for multimodal deep learning models. This method can reduce the latency and overall energy consumption of deep learning-based multimodal inference services without significantly reducing inference quality. It also dynamically selects the split points of each feature encoder in the multimodal deep learning model and the scale of the deep learning model to improve the speed and energy efficiency of multimodal deep learning model inference.

[0008] The present invention comprises the following steps:

[0009] Step 1: The multimodal deep learning model contains multiple feature encoders of modal data based on the deep learning model, and a fusion backend; each modality corresponds to a feature encoder, the number of feature encoders is M (M≥1), and the number of deep learning model layers of feature encoder i (1≤i≤M) is L i (L i ≥1), the deep learning model split point is x i (1≤x i ≤L i ); The optional compression rate level is recorded as E (E ≥ 1), and for each feature encoder i, it is compressed into E-1 different levels of compression models, and the compression rate of each model is c i ∈(0,1), c i ∈{c1,c2,...,c E-1}, use the dataset to train all compressed models for fine-tuning, and store the compressed and fine-tuned models together with the original models in the candidate model pool ν i ; The number of inference tasks generated by the mobile device in each time slot is z, where 1≤z≤z max , z max is the maximum number of tasks generated by the device at one time; the prediction vector is y (k) ,in and represents the probability of T predefined classes;

[0010] Step 2: Initialize the number of states R, the number of actions W, and the Q value matrix Q R×W , learning factor η∈[0,1], discount factor γ∈[0,1], exploration factor 0<ε min <ε max <1, annealing step number τ>0, benefit weight parameters w0>0, w1>0, w2>0; initialize channel gain h (0) , total delay t (0) , overall energy consumption (0) and the long-term confidence score α (0) ;

[0011] Step 3: At the kth time slot, the mobile device generates z based on the multimodal sensing data. (k) Inference task, and observe the channel gain h of the previous time slot (k-1) , total delay t (k-1) , overall energy consumption (k-1) And the model long-term confidence score α, construct the current state vector s (k) =[h (k-1) ,z (k) ,t (k-1) ,e (k-1) ,α];

[0012] Step 4: At the kth time slot, according to the Q value matrix, select the compression rate with the maximum Q value in the current state with a probability of 1-ε and deep learning model split points Randomly select other actions with probability ε;

[0013] Step 5: Based on the selected strategy a (k) =[c (k) ,x (k) ], for feature encoder i∈[1,M], the mobile device selects from the candidate model pool ν i Select the compression ratio as The compression model δ i (k) Input the raw data of all inference tasks into δ i (k) Before completion The calculation of the layer gets the intermediate result φ i (k) After all feature encoders have completed their inference tasks, the inference latency on the mobile device is calculated by the internal processor of the mobile device. and inferred energy consumption

[0014] Step 6: Set the compression ratio of each feature encoder Split Point Intermediate result φ i (k) The information is sent to the edge server with a transmission power of P;

[0015] Step 7: After receiving data from the mobile device, the edge server measures the transmission delay and transmission energy consumption And process the data to get the compression ratio Split Point Intermediate result φ i (k) ; For modality i, from the candidate model pool ν i Select the corresponding compression ratio as Model From the split point Then start executing the remaining layer calculation; after all feature encoders have completed the inference task, each The calculated output data is merged and used as the input layer of the fusion backend to calculate the final inference result ζ (k) , and then the internal processor of the edge server calculates the total inference delay of the local end And use Softmax regression to turn the output into a probability distribution in the fusion backend to get the prediction vector y (k), the confidence score ρ is calculated using the following formula:

[0016]

[0017] Step 8: The edge server collects the inference results (k) , calculate the long-term confidence score Forming feedback information Send to mobile device;

[0018] Step 9: After receiving the feedback information, the mobile device calculates the total delay of k time slots and overall energy consumption Calculate the benefit u generated by this inference (k) :

[0019] u (k) =w0α (k) -w1t (k) -w2e (k)

[0020] Step 10: Update Q(s (k) ,a (k) ):

[0021]

[0022] Step 11: Repeat steps 3 to 10 until |Q(s (k+1) ,a (k+1) )-Q(s (k) ,a (k) )|<0.01, which means that the mobile device has learned a stable inference selection strategy.

[0023] In step 1, the compression ratio of the original model of each feature encoder is recorded as 1.0.

[0024] In step 4, the exploration factor ε is calculated by ε in τ time slots. max Uniformly reduced to ε min .

[0025] Compared with the prior art, the present invention has the following outstanding advantages:

[0026] This paper utilizes a reinforcement learning algorithm to dynamically optimize the collaborative inference strategy for multimodal deep learning models serving mobile devices in wireless mobile edge networks. This strategy adapts to the characteristics of computationally intensive multimodal deep learning applications with multiple heterogeneous feature encoders. It reduces the latency and overall energy consumption of deep learning-based multimodal inference services without significantly compromising inference quality. It also dynamically selects the split points for each feature encoder in the multimodal deep learning model and the size of the deep learning model, improving the speed and energy efficiency of multimodal deep learning model inference.

[0027] The present invention solves the problems of high latency and high energy consumption when heterogeneous multimodal deep learning networks for computationally intensive applications are deployed on resource-scarce mobile devices. Based on the channel state, inference latency, inference energy consumption, long-term confidence score of the model and the amount of inference tasks generated in real time between the mobile device and the edge device in the previous time slot, the reinforcement learning algorithm is applied to dynamically optimize the selection of split points and compression rates of the multimodal deep learning model. The present invention effectively reduces the overall latency of the inference process and reduces the energy consumption during the inference process. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is the overall delay of the inference process of the collaborative inference method described in the embodiment of the present invention.

[0029] Figure 2 It is the total energy consumption of the inference process of the collaborative inference method described in the embodiment of the present invention. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the following embodiments will be further described in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0031] The embodiment of the present invention includes the following steps:

[0032] Step 1: The multimodal deep learning model contains multiple feature encoders for modal data based on the deep learning model, and a fusion backend. Each modality corresponds to a feature encoder. Let the number of feature encoders M = 3, and the number of deep learning model layers for feature encoder i (1≤i≤M) be L. i ∈{32,32,256}, the deep learning model split point is x i , where x1∈{8,16,30},x2∈{6,12,18},x3∈{12,32,56,108}; the optional compression rate level is recorded as E=3, for each feature encoder i, it is compressed into two different levels of compression models, and the compression rate of each model c i ∈{0.5,0.75,1.0}, and use the dataset to train all compressed models for fine-tuning, and store the compressed and fine-tuned models together with the original models in the candidate model pool ν i The number of inference tasks generated by the mobile device in each time slot is z, where 1≤z≤z max , z max =4 is the maximum number of tasks generated by the device at one time. Let the prediction vector be y (k) ,in and represents the probability of T predefined classes.

[0033] Step 2: Initialize the number of states R = 2000, the number of actions W = 972, and the Q value matrix Q = 0 R×W , learning factor η = 0.8, discount factor γ = 0.6, exploration factor ε = ε max =1,ε min =0.1, annealing steps τ = 1000, benefit weight parameters w0 = 0.6, w1 = 0.4, w2 = 1.3; initialize channel gain h (0) =0, overall delay t (0) =6, overall energy consumption e (0) = 10 and the long-term confidence score α (0) =0.

[0034] Step 3: At the kth time slot, the mobile device generates z based on the multimodal sensing data. (k) Inference task, and observe the channel gain h of the previous time slot (k-1) , total delay t (k-1) , overall energy consumption (k-1) And the model long-term confidence score α, construct the current state vector s (k) =[h (k-1) ,z (k) ,t (k-1) ,e (k-1) ,α].

[0035] Step 4: At the kth time slot, according to the Q value matrix, select the compression rate with the maximum Q value in the current state with a probability of 1-ε and deep learning model split points The other actions are chosen randomly with probability ε.

[0036] Step 5: Based on the selected strategy a (k) =[c (k) ,x (k) ], for feature encoder i∈[1,M], the mobile device selects from the candidate model pool ν i Select the compression ratio as The compression model δ i (k) , input the raw data of all inference tasks into δ i (k) Before completion The calculation of the layer gets the intermediate result φ i (k) After all feature encoders have completed the inference task, the internal processor of the mobile device calculates the inference latency on the mobile device. and inferred energy consumption

[0037] Step 6: Set the compression ratio of each feature encoder Split Point Intermediate result φ i (k) The information is sent to the edge server with a transmission power of P.

[0038] Step 7: After receiving data from the mobile device, the edge server measures the transmission delay and transmission energy consumption And process the data to get the compression ratio Split Point Intermediate result φ i (k) For modality i, from the candidate model pool ν i Select the corresponding compression ratio as Model From the split point Then start executing the remaining After all feature encoders have completed the inference task, each δ i (k) (1≤i≤M) The output data is merged and used as the input layer of the fusion backend to calculate the final inference result ζ (k) , and then the internal processor of the edge server calculates the total inference delay of the local end And use Softmax regression to turn the output into a probability distribution in the fusion backend to get the prediction vector y (k) , the confidence score ρ is calculated using the following formula:

[0039]

[0040] Step 8: The edge server collects the inference results (k) , calculate the long-term confidence score Forming feedback information Send to mobile device.

[0041] Step 9: After receiving the feedback information, the mobile device calculates the total delay of k time slots and overall energy consumption Calculate the benefit u generated by this inference (k) :

[0042] u (k) =w0α (k) -w1t (k) -w2e (k)

[0043] Step 10: Update Q(s (k) ,a (k) ):

[0044]

[0045] Step 11: Repeat steps 3 to 10 until |Q(s (k+1) ,a (k+1) )-Q(s (k) ,a (k) )|<0.01, which means that the mobile device has learned a stable inference selection strategy.

[0046] Depend on Figures 1-2 It can be seen that the embodiments of the present invention can reduce inference time and save system energy consumption. The present invention addresses the problems of extended inference time and high device energy consumption faced by mobile devices when performing inference tasks involving multiple sensor modal data, such as audio-visual speech recognition and multimedia event detection. The wireless channel state is estimated, the inference task amount is obtained, and the inference delay, energy consumption and confidence level of the inference result are obtained based on the feedback information of the edge device. Reinforcement learning is used to select the segmentation points and compression rates of each deep neural network feature encoder in the multimodal model, thereby reducing the inference delay and energy consumption while ensuring the accuracy of the inference result.

[0047] The above embodiments are only preferred embodiments of the present invention and should not be considered to limit the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent of the present invention.

Claims

1. A fast collaborative inference method for multimodal deep learning models, characterized by The following steps are involved: Step 1: The multimodal deep learning model contains multiple feature encoders of modal data based on the deep learning model and a fusion backend; each modality corresponds to a feature encoder, and the number of feature encoders is M, where M ≥ 1, and the number of deep learning model layers of feature encoder i is L i , where 1≤i≤M, L i ≥1, the deep learning model split point is x i , where 1≤x i ≤L i ; The optional compression rate level is recorded as E, where E≥1. For each feature encoder i, it is compressed into E-1 different levels of compression models, and the compression rate of each model is c i ∈(0,1), c i ∈{c1,c2,...,c E-1 }, use the dataset to train all compressed models for fine-tuning, and store the compressed and fine-tuned models together with the original models in the candidate model pool ν i ; The number of inference tasks generated by the mobile device in each time slot is z, where 1≤z≤z max , z max is the maximum number of tasks generated by the device at one time; record the prediction vector and Step 2: Initialize the number of states R, the number of actions W, and the Q value matrix Q R×W , learning factor η∈[0,1], discount factor γ∈[0,1], exploration factor 0<ε min <ε max <1, annealing step number τ>0, benefit weight parameters w0>0, w1>0, w2>0; initialize channel gain h (0) , total delay t (0) , overall energy consumption (0) and the long-term confidence score α (0) ; Step 3: At the kth time slot, the mobile device generates z based on the multimodal sensing data (k) Inference task, and observe the channel gain h of the previous time slot (k-1) , total delay t (k-1) , overall energy consumption (k-1) And the model long-term confidence score α, construct the current state vector s (k) =[h (k-1) ,z (k) ,t (k-1) ,e (k-1) ,α]; Step 4: At the kth time slot, according to the Q value matrix, select the compression rate with the maximum Q value in the current state with a probability of 1-ε and deep learning model split points Randomly select other actions with probability ε; Step 5: Based on the selected strategy a (k) =[c (k) ,x (k) ], for feature encoder i∈[1,M], the mobile device selects from the candidate model pool ν i Select the compression ratio as Compression model Feed the model with raw data for all inference tasks Before completion Layer calculation to obtain intermediate results After all feature encoders have completed their inference tasks, the inference latency on the mobile device is calculated by the mobile device's internal processor. and inferred energy consumption Step 6: Set the compression ratio of each feature encoder Split Point Intermediate results The information is sent to the edge server with a transmission power of P; Step 7: After receiving data from the mobile device, the edge server measures the transmission delay and transmission energy consumption And process the data to get the compression ratio Split Point Intermediate results For modality i, from the candidate model pool ν i Select the corresponding compression ratio as Model From the split point Then start executing the remaining Layer calculation; After all feature encoders have completed the inference task, each The calculated output data is merged, where 1≤i≤M, as the input layer of the fusion backend, and the final inference result ζ is calculated. (k) , and then the internal processor of the edge server calculates the total inference delay of the local end And use Softmax regression to turn the output into a probability distribution in the fusion backend to get the prediction vector y (k) , use the following formula to calculate the confidence ρ(y): Step 8: The edge server collects the inference results (k) , calculate the long-term confidence score Forming feedback information Send to mobile device; Step 9: After receiving the feedback information, the mobile device calculates the total delay of k time slots and overall energy consumption Calculate the benefit u generated by this inference (k) : Step 10: Update Q(s (k) ,a (k) ): Step 11: Repeat steps 3 to 10 until |Q(s (k+1) ,a (k+1) )-Q(s (k) ,a (k) )|<0.01, which means that the mobile device has learned a stable inference selection strategy.

2. A fast collaborative inference method for multimodal deep learning models as described in claim 1, characterized in that In step 1, the compression ratio of the original model of each feature encoder is recorded as 1.

0.

3. A fast collaborative inference method for multimodal deep learning models as claimed in claim 1, characterized in that In step 4, the exploration factor ε is calculated by ε in τ time slots. max Uniformly reduced to ε min .

Citation Information

Patent Citations

  • Deep neural network cooperative reasoning method based on end-edge cloud architecture

    CN112348172A

  • Mobile equipment collaborative inference system for deep learning Transform class model

    CN114662690A

  • Relevant redundant transformation and reinforcement learning-based multi-dimension cooperative control method

    CN108021028A

  • Adaptive optimization scheduling method for mobile terminal software based on deep reinforcement learning

    CN109002358A