Multi-modal model distributed reasoning optimization method for dynamic intelligent Internet of Things
By building a multimodal model distributed inference optimization system, the parallel scheduling strategy of DNN inference task data is determined and the long-term energy consumption is optimized using heuristic algorithms, the inference delay and energy consumption problems of multimodal tasks in resource-constrained environments are solved, and efficient distributed inference optimization is achieved.
Patent Information
- Application Number
- CN202510282215.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-08-01
AI Technical Summary
The existing inference task processing strategies cannot meet the dynamic and distributed processing requirements of multimodal tasks. Especially in resource-constrained environments, how to effectively manage the distributed inference tasks of multimodal models and optimize inference delay and energy consumption.
Build a multimodal model distributed inference optimization system, determine the parallel scheduling strategy of DNN inference task data, optimize the minimization of long-term delay under long-term energy consumption constraints through heuristic algorithms, and realize efficient model deployment and distributed inference optimization of multimodal models.
Significantly optimize inference time and energy consumption, and improve the ubiquitous inference efficiency and computing power utilization of the intelligent Internet of Things.
Smart Images

Figure CN120407096A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent Internet of Things, and in particular to a multi-modal model distributed inference optimization method for dynamic intelligent Internet of Things. Background Art
[0002] With the rapid development of the intelligent Internet of Things, the amount of data generated by Internet of Things devices has increased sharply, and the demand for data processing and inference tasks has also grown accordingly. Existing inference task processing strategies often cannot meet the dynamic and distributed processing requirements of multi-modal tasks, especially in resource-constrained environments. Therefore, how to effectively manage the distributed inference tasks of multi-modal models and optimize inference latency and energy consumption has become an urgent problem to be solved in this field. Summary of the Invention
[0003] In a first aspect, an embodiment of the present invention provides a multi-modal model distributed inference optimization method for dynamic intelligent Internet of Things, the method includes:
[0004] Construct a multi-modal model distributed inference optimization system for dynamic intelligent Internet of Things; in the multi-modal model distributed inference optimization system, it includes multiple Internet of Things devices that generate and upload DNN inference tasks and multiple servers that receive and process DNN inference tasks;
[0005] Determine a DNN inference task data parallel scheduling strategy for the multi-modal model distributed inference optimization system;
[0006] Use the DNN inference task data parallel scheduling strategy to determine the system inference time and system energy consumption of the multi-modal model distributed inference optimization system;
[0007] Construct a minimization of long-term latency under long-term energy consumption constraints based on the system inference time and system energy consumption, and use this as an optimization goal to construct a long-term distributed inference stability optimization model;
[0008] Adopt a heuristic algorithm to solve the long-term distributed inference stability optimization model to achieve multi-modal model distributed inference optimization.
[0009] [[ID=�1]]In some realizable ways of the first aspect, in the multi-modal model distributed inference optimization system, multiple Internet of Things devices are denoted as Multiple servers are denoted as
[0010] Divide the entire time domain into multiple time slots, denoted as
[0011] In each time slot, an Internet of Things device generates a type of DNN inference task, and the task type is denoted as Represent different types, i.e., modalities, and the Internet of Things device C m The DNN inference task generated within time slot τ, also known as the Internet of Things device within time slot τ Is expressed as:
[0012]
[0013] Among them, Represents the type of the DNN inference task within time slot τ And the size of the storage resources occupied by the model corresponding to the DNN inference task Represents the amount of data required for the DNN inference task; Represents the deadline of the DNN inference task; Represents the transmission power of the Internet of Things device;
[0014] The server within time slot τ Is expressed as:
[0015]
[0016] Among them, Represents the server computing power; Represents the server computing resources; Represents the server storage resources; Represents the server computing power; Represents the server credibility.
[0017] In some realizable ways of the first aspect, determining the DNN inference task data parallel scheduling strategy for the multi-modal model distributed inference optimization system includes:
[0018] Within time slot τ, whether a DNN inference task is generated from the Internet of Things device To the server The distribution of the DNN inference task is denoted as When It indicates that a DNN inference task distribution is generated. At this time, the amount of DNN inference task data to be transmitted is denoted as And when It indicates that no DNN inference task distribution is generated. At this time Let Represents the server Whether the model corresponding to the task type Is deployed. When It indicates that the model corresponding to the task type Is deployed on the server When There is no deployment. Further, if And Then and are both 0; On this basis, the DNN inference task data parallel scheduling strategy for time slot τ is expressed as:
[0019]
[0020] In some realizable ways of the first aspect, using the DNN inference task data parallel scheduling strategy to determine the system inference time and system energy consumption of the multimodal model distributed inference optimization system includes:
[0021] Using γ within time slot τ [τ] to represent the data transmission rate matrix between the server and the Internet of Things devices. The magnitude of the transmission rate is related to the transmission power of the Internet of Things devices, the distance and bandwidth between the Internet of Things devices and the server. Assuming that the Internet of Things devices move randomly within a certain range, then γ [τ] will be different for each time slot τ and is expressed as:
[0022]
[0023] For each time slot τ, if the Internet of Things device sends DNN inference task data to the server then the data transmission time is expressed as:
[0024]
[0025] where represents the data transmission rate between the Internet of Things device and the server
[0026]
[0027]
[0028]
[0029]
[0030]
[0031]
[0026]
[0027]
[0028]
[0029]
[0030]
[0031]
[0026]
[0028]
[0029]
[0030]
[0031]
[0026]
[0028]
[0029]
[0030]
[0031]
[0026]
[0028]
[0029]
[0030]
[0031]
[0026]
[0028]
[0029]
[0030]
[0031]
[0026]
[0028]
[0029]
[0030]
[0031]
[0026]
[0028]
[0029]
[0030]
[0031]
[0026]
[0028]
[0029]
[0030]
[0031]
[0026]
[0028]
[0029]
[0030]
[0031]
[0026]
[0028]
[0029]
[0030]
[0031]
[0026]
[0028]
[0029]
[0030]
[0031]
[0026]
[0028]
[0029]
[0030]
[0031]
[0026]
[0028]
[0029]
[0030]
[0031]
[0026]
[0028]
[0029]
[0030]
[0031]
[0026]
[0028]
[0029]
[0030]
[0031]
[0026]
[0028] [[Based on the data parallel scheduling strategy for DNN inference tasks, the DNN inference task data is divided into multiple sub-data to be scheduled to different servers. Therefore, the Internet of Things devices The inference time of the complete DNN inference task, that is, the system time is expressed as:
[0032]
[0033] The Internet of Things devices communication energy consumption is expressed as:
[0034]
[0035] The servers computing energy consumption is expressed as:
[0036]
[0037] Based on the communication energy consumption and the computing energy consumption determine the system energy consumption e [τ] as:
[0038]
[0039] In some realizable ways of the first aspect, the construction process of the optimization target includes:
[0040] For each time slot τ, the overall completion time of the DNN inference task is expressed as:
[0041]
[0042] Construct the minimization of the long-term delay under the long-term energy consumption constraint, as shown below:
[0043]
[0044] If and then
[0045]
[0046] e [τ] ≥0;
[0047] where ε max represents the upper limit value of the system energy consumption;
[0048] Within the time slot τ, use to represent the Internet of Things devices The amount of task data actually completed is determined by the policy Then the update of the real queue is:
[0049]
[0050] where Q [τ] represents the real queue in time slot τ, and Q [τ+1] represents the real queue in time slot τ + 1;
[0051] Correspondingly, the update of the virtual queue is:
[0052]
[0053] where represents the virtual queue in time slot τ, represents the virtual queue in time slot τ + 1;
[0054] On this basis, the Lyapunov function is expressed as:
[0055]
[0056] The Lyapunov drift Δ(τ) is expressed as:
[0057]
[0058] The drift plus penalty function is expressed as:
[0059]
[0060] where V represents a non - negative real number;
[0061] Substituting the update formulas of the real queue and the virtual queue, we get:
[0062]
[0063] The upper bound of Δ(τ) is:
[0064]
[0065] where μ represents a constant;
[0066] To minimize the drift plus penalty function, the optimization objective is transformed into an optimization objective for each time slot τ, and solved separately for each time slot:
[0067]
[0068] According to the upper bound of Δ(τ), the optimization objective is transformed into:
[0069]
[0070] If and then
[0071]
[0072] e [τ] ≥0.
[0073] In some realizable ways of the first aspect, solving the long-term distributed inference stable optimization model by using a heuristic algorithm to realize distributed inference optimization of the multimodal model includes:
[0074] Initializing a virtual queue and a real queue, and initializing the performance parameters of each Internet of Things device and each server;
[0075] Making decisions on task scheduling and resource allocation in each time slot;
[0076] Continuously optimizing the queue state and task scheduling strategy through real-time feedback to achieve the optimization goal.
[0077] In some realizable ways of the first aspect, making decisions on task scheduling and resource allocation in each time slot includes:
[0078] Comprehensively considering various characteristics of each DNN inference task and assigning priorities to the DNN inference tasks;
[0079] According to the queue state and priorities, determining the resource allocation strategy, selecting a server with resource adaptation for task allocation for servers with different computing capabilities and computing resources, and calculating the optimization goal;
[0080] Updating the virtual queue and the real queue and adjusting according to the energy consumption and task scheduling situation;
[0081] Adjusting task scheduling and resource allocation in each time slot, optimizing the objective function, and using the scheduling priority based on the queue state and system load to optimize the system performance.
[0082] In a second aspect, an embodiment of the present invention provides a distributed inference optimization device for a multimodal model for a dynamic intelligent Internet of Things, and the device includes:
[0083] A construction module, configured to construct a distributed inference optimization system for a multimodal model for a dynamic intelligent Internet of Things; in the distributed inference optimization system for the multimodal model, it includes multiple Internet of Things devices that generate and upload DNN inference tasks and multiple servers that receive and process DNN inference tasks;
[0084] A determination module, configured to determine a data parallel scheduling strategy for DNN inference tasks for a multi-modal model distributed inference optimization system;
[0085] The determination module is further configured to use the data parallel scheduling strategy for DNN inference tasks to determine the system inference time and system energy consumption of the multi-modal model distributed inference optimization system;
[0086] The construction module is further configured to construct a minimized long-term delay under long-term energy consumption constraints based on the system inference time and system energy consumption, and use this as an optimization objective to construct a long-term distributed inference stability optimization model;
[0087] A solution module, configured to solve the long-term distributed inference stability optimization model by using a heuristic algorithm to achieve multi-modal model distributed inference optimization.
[0088] In a third aspect, an embodiment of the present invention provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.
[0089] In a fourth aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method as described above.
[0090] Compared with the prior art, the present invention has at least the following technical effects:
[0091] In view of the dynamic characteristics of resource supply and demand in the intelligent Internet of Things, the present invention constructs a long-term distributed inference stability optimization model with the optimization objective of minimizing the long-term delay under long-term energy consumption constraints. By efficiently solving the model using a heuristic algorithm, efficient model deployment and distributed inference optimization of the multi-modal model are achieved. In this way, the distributed inference tasks of the multi-modal model can be effectively managed, the inference time and energy consumption can be significantly optimized, and thus the ubiquitous inference efficiency and computing power utilization rate of the intelligent Internet of Things can be improved.
[0092] It should be understood that the content described in the summary of the invention is not intended to limit the key or important features of the embodiments of the present invention, nor to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0093] In conjunction with the drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present invention will become more apparent. The drawings are used to better understand the present invention and do not constitute a limitation to the present invention. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0094] Figure 1 A flowchart of a multi-modal model distributed inference optimization method for a dynamic intelligent Internet of Things provided by an embodiment of the present invention;
[0095] Figure 2 A structural diagram of a multi-modal model distributed inference optimization device for a dynamic intelligent Internet of Things provided by an embodiment of the present invention;
[0096] Figure 3 A structural diagram of an exemplary electronic device capable of implementing the embodiments of the present invention. Detailed implementation manners
[0097] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0098] In addition, the term "and / or" in the present invention is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present invention generally represents an "or" relationship between the associated objects before and after.
[0099] To solve the technical problems in the background art, embodiments of the present invention provide a multi-modal model distributed inference optimization method, device, equipment, and storage medium for a dynamic intelligent Internet of Things. The following will describe in detail a multi-modal model distributed inference optimization method, device, equipment, and storage medium for a dynamic intelligent Internet of Things provided by embodiments of the present invention with reference to the accompanying drawings and through specific embodiments.
[0100] Figure 1 A flowchart of a multi-modal model distributed inference optimization method for a dynamic intelligent Internet of Things provided by an embodiment of the present invention. As Figure 1 shown, the multi-modal model distributed inference optimization method 100 may include:
[0101] S110, constructing a multi-modal model distributed inference optimization system for a dynamic intelligent Internet of Things.
[0102] In the multi-modal model distributed inference optimization system, it includes multiple Internet of Things devices that generate and upload DNN inference tasks and multiple servers that receive and process DNN inference tasks.
[0103] S120. Determine the data parallel scheduling strategy for the DNN inference task for the multi-modal model distributed inference optimization system.
[0104] S130. Use the data parallel scheduling strategy for the DNN inference task to determine the system inference time and system energy consumption of the multi-modal model distributed inference optimization system.
[0105] S140. Construct the minimization of the long-term delay under the long-term energy consumption constraint based on the system inference time and system energy consumption, and use this as the optimization goal to construct a long-term distributed inference stability optimization model.
[0106] S150. Solve the long-term distributed inference stability optimization model using a heuristic algorithm to achieve multi-modal model distributed inference optimization.
[0107] For the convenience of further understanding, the above steps will be described in detail below in combination with specific embodiments:
[0108] (1) Construct a multi-modal model distributed inference optimization system
[0109] In the multi-modal model distributed inference optimization system, multiple Internet of Things devices are denoted as Multiple servers (edge nodes and cloud) are denoted as
[0110] The entire time domain is divided into multiple time slots, denoted as
[0111] In each time slot, the Internet of Things device generates a kind of Deep Neural Network (DNN) inference task, and the task type is denoted as They respectively represent modalities such as text, audio, image, video, etc. Denote the Internet of Things device C m The DNN inference task generated by the Internet of Things device in time slot τ, also known as the Internet of Things device in time slot τ It is expressed as:
[0112]
[0113] Among them, Represents the type of the DNN inference task in time slot τ And the size of the storage resources occupied by the model corresponding to the DNN inference task Represents the amount of data required for the DNN inference task; Represents the deadline of the DNN inference task; Represents the transmission power of the Internet of Things device.
[0114] The server in time slot τ It is expressed as:
[0115]
[0116] Among them, represents the computing power of the server; represents the computing resources of the server; represents the storage resources of the server, and further represents the maximum number of concurrent tasks of the server within the time slot τ; represents the computing power of the server; represents the credibility of the server. It should be noted that changes with time and follows a certain specific distribution.
[0117] (2) Determine the data parallel scheduling strategy for DNN inference tasks
[0118] Within the time slot τ, whether a DNN inference task is generated from the Internet of Things device to the server is denoted as When , it means that a DNN inference task distribution is generated, that is, a transmission of task data is generated. At this time, the amount of DNN inference task data transmitted is denoted as While when , it means that no DNN inference task distribution is generated. At this time This invention focuses on multi-modal model deployment. Here, let represent whether the server has deployed the model corresponding to the task type . When , it means that the model corresponding to the task type is deployed on the server . When , there is no deployment. Further, if and , then and are both 0. On this basis, the data parallel scheduling strategy of the DNN inference task in the time slot τ is expressed as:
[0119]
[0120] (3) Determine the system inference time and system energy consumption
[0121] (3.1) System inference time
[0122] Within the time slot τ, use γ [τ]Denote the data transmission rate matrix between the server and the Internet of Things (IoT) devices. The magnitude of the transmission rate is related to the transmit power of the IoT devices, the distance between the IoT devices and the server, and the bandwidth. Assume that the IoT devices move randomly within a certain range, then γ [τ] will be different for each time slot τ, and it is expressed as:
[0123]
[0124] Therefore, for each time slot τ, if the IoT device sends DNN inference task data to the server then the data transmission time is expressed as the ratio of the amount of transmitted data to the data transmission rate, that is:
[0125]
[0126] where represents the data transmission rate between the IoT device and the server during time slot τ.
[0127] The computing time of the server is expressed as:
[0128]
[0129] where z represents the number of floating-point operations required per unit of data volume.
[0130] The IoT device sends the inference time corresponding to the DNN inference task to the server is expressed as the sum of the transmission time and the computing time:
[0131] [[ID=!50]]
[0132] Based on the data parallel scheduling strategy for DNN inference task data, the DNN inference task data is divided into multiple sub-data to be scheduled to different servers. Therefore, the inference time of the complete DNN inference task of the IoT device which is also the system time is expressed as:
[0133]
[0134] Due to the data parallel strategy, the task data of the IoT device is executed in parallel on different servers. Therefore, the inference time of the complete task should be the longest completion time of all sub-task data.
[0135] (3.2) System Energy Consumption
[0136] Internet of Things device needs to transmit its task data to the server for edge inference. Its energy consumption mainly lies in the communication with the server, and its communication energy consumption can be expressed as the product of the data transmission time and the transmission power, i.e.:
[0137]
[0138] Server is responsible for receiving the DNN inference task and completing the inference calculation. Its computing energy consumption can be expressed as the product of the calculation of subtasks and the computing power of the server:
[0139]
[0140] Based on the communication energy consumption and the computing energy consumption determine the system energy consumption e [τ] as:
[0141]
[0142] (4) Construct a long-term distributed inference stability optimization model
[0143] For each time slot τ, the overall completion time of the DNN inference task is expressed as:
[0144]
[0145] Construct the minimization of the long-term delay under the long-term energy consumption constraint, and use this as the optimization goal to construct a long-term distributed inference stability optimization model, as shown below:
[0146]
[0147] If and then
[0148]
[0149] e [τ] ≥0# (19)
[0150] where ε max represents the upper limit value of the system energy consumption.
[0151] Within the time slot τ, use to represent the amount of task data actually completed by the Internet of Things device which is determined by the policy The update of the real queue is as follows:
[0152]
[0153] Among them, Q [τ] represents the real queue within time slot τ, and Q [τ+1] represents the real queue within time slot τ + 1.
[0154] Correspondingly, the update of the virtual queue is as follows:
[0155]
[0156] Among them, represents the virtual queue within time slot τ, represents the virtual queue within time slot τ + 1.
[0157] On this basis, the Lyapunov function is expressed as:
[0158]
[0159] The Lyapunov drift Δ(τ) is expressed as:
[0160]
[0161] The drift-plus-penalty function is expressed as:
[0162]
[0163] Among them, V represents a non-negative real number.
[0164] Substituting the update formulas of the real queue and the virtual queue, we get:
[0165]
[0166] The upper bound of Δ(τ) is:
[0167]
[0168] Among them, μ represents a constant;
[0169] To minimize the drift-plus-penalty function, the optimization objective is transformed into an optimization objective for each time slot τ, and solved separately for each time slot:
[0170]
[0171] According to the upper bound of Δ(τ), the optimization objective is transformed into:
[0172]
[0173] If and then
[0174]
[0175] e [τ] ≥0#(34)
[0176] (5) Solving the long - term distributed inference stable optimization model
[0177] Optimize the Lyapunov queue in each time slot, and minimize the upper bound of the Lyapunov drift, so as to minimize the optimization objective. Here, a heuristic algorithm is used to design a specific optimization scheme. The heuristic algorithm is suitable for solving complex problems. Especially when the global optimal solution cannot be directly obtained, approximate optimal decisions are made in each time slot through heuristic strategies. Specifically, the goal of the scheme is to optimize queue management, task scheduling, resource allocation, etc. in each time slot, and optimize the overall performance of the system by minimizing the upper bound of the Lyapunov drift.
[0178] As an example, the heuristic algorithm is as follows:
[0179] Step 1: Initialize the queues and system parameters
[0180] Initialize the virtual queue and the real queue Initialize the performance parameters (such as computing power, bandwidth, energy consumption, etc.) of each IoT device and each server.
[0181] Step 2: Perform task scheduling and resource allocation in each time slot
[0182] In each time slot, make decisions on task scheduling and resource allocation.
[0183] Step 2.1 Task priority sorting
[0184] Considering multiple characteristics (such as task type, data volume, computational complexity, etc.) of each DNN inference task comprehensively, assign priorities to DNN inference tasks.
[0185] Step 2.2 Resource allocation decision
[0186] According to the queue status and priorities, determine the resource allocation strategy. For servers with different computing capabilities and computing resources, select servers with resource adaptation for task allocation, and calculate the optimization objective.
[0187] Step 2.3 Queue update
[0188] Update the virtual queue and the real queue Q [τ], adjust according to the energy consumption and task scheduling situation.
[0189]
[0190] Step 2.4 Minimize the upper bound of the optimized objective after derivation
[0191] Adjust the task scheduling and resource allocation in each time slot, optimize the objective function, and use the scheduling priority based on the queue state and system load to optimize the system performance.
[0192]
[0193] Step 3: Real-time feedback and adjustment
[0194] Continuously optimize the queue state and task scheduling strategy through the real-time monitored feedback (such as queue state, task completion, energy consumption, etc.) to achieve the optimization objectives (minimize latency, energy consumption, and reduce the data discarded due to inability to complete within the time slot).
[0195] It should be noted that the present invention can also use heuristic algorithms such as Greedy Algorithm, Simulated Annealing, and Ant Colony Optimization to solve the model.
[0196] In summary, the present invention has at least achieved the following technical effects:
[0197] The present invention constructs a long-term distributed inference stable optimization model with the optimization objective of minimizing the long-term latency under the long-term energy consumption constraint for the dynamic characteristics of resource supply and demand in the intelligent Internet of Things. By using heuristic algorithms to efficiently solve the model, efficient model deployment and distributed inference optimization of multi-modal models are achieved. In this way, the distributed inference tasks of multi-modal models can be effectively managed, the inference time and energy consumption can be significantly optimized, thereby improving the ubiquitous inference efficiency and computing power utilization rate of the intelligent Internet of Things.
[0198] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0199] The above is the introduction of the method embodiments. The following further illustrates the solution of the present invention through device embodiments.
[0200] Figure 2 The structural diagram of a multi-modal model distributed inference optimization device provided for an embodiment of the present invention for a dynamic intelligent Internet of Things is as follows Figure 2 shown. The multi-modal model distributed inference optimization device 200 may include:
[0201] A construction module 210, configured to construct a multi-modal model distributed inference optimization system for a dynamic intelligent Internet of Things; in the multi-modal model distributed inference optimization system, it includes multiple Internet of Things devices that generate and upload DNN inference tasks and multiple servers that receive and process DNN inference tasks.
[0202] A determination module 220, configured to determine a DNN inference task data parallel scheduling strategy for the multi-modal model distributed inference optimization system.
[0203] The determination module 220 is further configured to use the DNN inference task data parallel scheduling strategy to determine the system inference time and system energy consumption of the multi-modal model distributed inference optimization system.
[0204] The construction module 210 is further configured to construct a minimization of long-term delay under long-term energy consumption constraints based on the system inference time and system energy consumption, and use this as an optimization objective to construct a long-term distributed inference stable optimization model.
[0205] A solution module 230, configured to solve the long-term distributed inference stable optimization model by using a heuristic algorithm to achieve multi-modal model distributed inference optimization.
[0206] It can be understood that Figure 2 each module / unit in the multi-modal model distributed inference optimization device 200 shown has the function of implementing Figure 1 each step in the multi-modal model distributed inference optimization method 100 shown, and can achieve its corresponding technical effects. For the sake of brevity, it will not be elaborated here.
[0207] Figure 3 The structural diagram of an exemplary electronic device capable of implementing an embodiment of the present invention. The electronic device 300 is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device 300 may also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown in the present invention, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed in the present invention.
[0208] As Figure 3As shown, the electronic device 300 may include a computing unit 301, which may perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 302 or a computer program loaded from a storage unit 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 may also be stored. The computing unit 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0209] Multiple components in the electronic device 300 are connected to the I / O interface 305, including: an input unit 306, such as a keyboard, a mouse, etc.; an output unit 307, such as various types of displays, speakers, etc.; a storage unit 308, such as a magnetic disk, an optical disc, etc.; and a communication unit 309, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 309 allows the electronic device 300 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0210] The computing unit 301 may be various general-purpose and / or dedicated processing components with processing and computing capabilities. Some examples of the computing unit 301 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 301 executes the various methods and processes described above, such as method 100. For example, in some embodiments, method 100 may be implemented as a computer program product, including a computer program, which is tangibly contained in a computer-readable medium, such as the storage unit 308. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 300 via the ROM 302 and / or the communication unit 309. When the computer program is loaded into the RAM 303 and executed by the computing unit 301, one or more steps of method 100 described above may be executed. Alternatively, in other embodiments, the computing unit 301 may be configured to execute method 100 in any other appropriate manner (e.g., by means of firmware).
[0211] The various embodiments described above in the present invention can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0212] The program code for implementing the methods of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a dedicated computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0213] In the context of the present invention, a computer-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of computer-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0214] It should be noted that the present invention also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute method 100 and achieve the corresponding technical effects achieved by the method of the embodiments of the present invention. For the sake of concise description, details are not repeated herein.
[0215] In addition, the present invention also provides a computer program product, which includes a computer program that implements method 100 when executed by a processor.
[0216] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solutions disclosed in the present invention can be achieved, and the present invention is not limited herein.
[0217] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A distributed inference optimization method for multi-modal models in the context of dynamic intelligent Internet of Things, characterized in that, The method includes: Constructing a distributed inference optimization system for a multi-modal model for a dynamic intelligent Internet of Things; in the distributed inference optimization system for the multi-modal model, it includes multiple Internet of Things devices that generate and upload DNN inference tasks and multiple servers that receive and process DNN inference tasks; Determining a DNN inference task data parallel scheduling strategy for the distributed inference optimization system for the multi-modal model; Using the DNN inference task data parallel scheduling strategy to determine the system inference time and system energy consumption of the distributed inference optimization system for the multi-modal model; Constructing a minimization of long-term delay under long-term energy consumption constraints based on the system inference time and system energy consumption, and using this as an optimization objective to construct a long-term distributed inference stability optimization model; Using a heuristic algorithm to solve the long-term distributed inference stability optimization model to achieve distributed inference optimization of the multi-modal model.
2. The method according to claim 1, wherein In the multi-modal model distributed inference optimization system, multiple Internet of Things devices are denoted as Multiple servers are denoted as Divide the entire time domain into multiple time slots, denoted as In each time slot, the IoT device generates a DNN inference task, and the task type is denoted as which represent different types, i.e., modalities. The IoT device C m generates a DNN inference task within time slot τ, also known as the IoT device within time slot τ is expressed as: Among them, represents the type l of the DNN inference task within the time slot τ and the size θ of the storage resources occupied by the corresponding model of the DNN inference task [l] ; represents the amount of data required for the DNN inference task; represents the deadline of the DNN inference task; represents the transmission power of the IoT device; The server within time slot τ is represented as: Among them, represents the computing power of the server; represents the computing resources of the server; represents the storage resources of the server; represents the computing power of the server; represents the credibility of the server.
3. The method according to claim 2, wherein The determination of the DNN inference task data parallel scheduling strategy for the distributed inference optimization system for the multi-modal model includes: During time slot τ, whether a DNN inference task distribution from the Internet of Things device to the server is denoted as When it indicates that a DNN inference task distribution is generated. At this time, the amount of DNN inference task data to be transmitted is denoted as While when it indicates that no DNN inference task distribution is generated. At this time Let denote whether the model corresponding to task type e is deployed on the server When it indicates that the model corresponding to task type l is deployed on the server during time slot τ. When it means there is no deployment. Further, if and then and are both 0. On this basis, the parallel scheduling strategy of DNN inference task data for time slot τ is expressed as:
4. The method according to claim 3, characterized in that The use of the DNN inference task data parallel scheduling strategy to determine the system inference time and system energy consumption of the distributed inference optimization system for the multi-modal model includes: Use γ within the τ time slot [τ] represents the data transmission rate matrix between the server and the Internet of Things devices. The magnitude of the transmission rate is related to the transmission power of the Internet of Things devices, the distance between the Internet of Things devices and the server, and the bandwidth. Assuming that the Internet of Things devices move randomly within a certain range, then γ [τ] will be different for each time slot τ, and it is expressed as: For each time slot τ, if the Internet of Things device sends DNN inference task data to the server then the data transmission time is expressed as: Among them, represents the data transmission rate between the Internet of Things device and the server during time slot τ; Computing time of the server It is expressed as: where z represents the number of floating-point operations required per unit of data volume; Internet of Things device to the server allocate the inference time corresponding to the DNN inference task which is expressed as: Based on the data parallel scheduling strategy for DNN inference tasks, the DNN inference task data is divided into multiple sub-data to be scheduled to different servers. Therefore, the Internet of Things devices The inference time of the complete DNN inference task, which is also the system time is expressed as: Express the communication energy consumption of the Internet of Things device as follows: The server computing energy consumption is expressed as: Based on communication energy consumption and computing energy consumption determine the system energy consumption e [τ] as follows:
5. The method according to claim 4, wherein The construction process of the optimization objective includes: For each time slot τ, the overall completion time of the DNN inference task is expressed as: Constructing a minimization of long-term delay under long-term energy consumption constraints, specifically as follows: and then e [τ] ≥0; Among them, ε max represents the upper limit value of the system energy consumption; During time slot τ, use to represent the Internet of Things device The actual amount of task data completed is determined by the policy Then the update of the real queue is as follows: Among them, Q [τ] represents the real queue within time slot τ, and Q [τ+1] represents the real queue within time slot τ + 1; Correspondingly, the update of the virtual queue is: Among them, represents the virtual queue within time slot τ, represents the virtual queue within time slot τ + 1; Based on this, the Lyapunov function is expressed as: The Lyapunov drift Δ(τ) is expressed as: The drift-plus-penalty function is expressed as: where V represents a non-negative real number; Substituting the update formulas of the real queue and the virtual queue gives: The upper bound of Δ(τ) is: where μ represents a constant; To minimize the drift-plus-penalty function, the optimization objective is transformed into an optimization objective for each time slot τ, and solved separately for each time slot: According to the upper bound of Δ(τ), the optimization objective is transformed into: and then e [τ] ≥0。 6. The method according to claim 5, characterized in that The use of a heuristic algorithm to solve the long-term distributed inference stability optimization model to achieve distributed inference optimization of the multi-modal model includes: Initializing the virtual queue and the real queue, and initializing the performance parameters of each Internet of Things device and each server; Making decisions on task scheduling and resource allocation in each time slot; Continuously optimizing the queue state and task scheduling strategy through real-time feedback to achieve the optimization objective.
7. The method according to claim 6, characterized in that The making of decisions on task scheduling and resource allocation in each time slot includes: Comprehensively considering various characteristics of each DNN inference task to assign priorities to the DNN inference tasks; According to the queue state and priorities, determining the resource allocation strategy, selecting a server with appropriate resources for task allocation for servers with different computing capabilities and computing resources, and calculating the optimization objective; Updating the virtual queue and the real queue, and adjusting according to the energy consumption and task scheduling situation; Adjusting task scheduling and resource allocation in each time slot, optimizing the objective function, and using scheduling priorities based on the queue state and system load to optimize the system performance.
8. A multi-modal model distributed inference optimization device for dynamic intelligent Internet of Things, characterized in that The device includes: A building module for building a distributed inference optimization system for a multi-modal model for a dynamic intelligent Internet of Things; in the distributed inference optimization system for a multi-modal model, it includes multiple Internet of Things devices that generate and upload DNN inference tasks and multiple servers that receive and process DNN inference tasks; A determination module for determining a DNN inference task data parallel scheduling strategy for the distributed inference optimization system for a multi-modal model; The determination module is further configured to use the DNN inference task data parallel scheduling strategy to determine the system inference time and system energy consumption of the distributed inference optimization system for a multi-modal model; The building module is further configured to construct a minimization of long-term latency under long-term energy consumption constraints based on the system inference time and system energy consumption, and use this as an optimization objective to construct a long-term distributed inference stability optimization model; A solving module for solving the long-term distributed inference stability optimization model by using a heuristic algorithm to achieve distributed inference optimization of a multi-modal model.
9. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the method described in any one of claims 1-7 is implemented.
10. A computer-readable storage medium, characterized in that, Computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, the method described in any one of claims 1-7 is implemented.
Citation Information
Cited By
Multi-mode perception and optimization method and system for low-power-consumption AR equipment
CN121387072A
A multi-modal perception and optimization method and system for low-power ar devices
CN121387072B