A task offloading and resource optimization method and device for large model transmission
By building a deep semantic authorization network and deep reinforcement learning algorithm, the latency and energy consumption problems in large-scale model transmission are solved, reasonable task allocation and resource optimization are achieved, and the efficiency and performance of edge computing are improved.
Patent Information
- Application Number
- CN202311821963.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2043-12-27
AI Technical Summary
During the transmission of large models, there are problems of excessive transmission delay, energy consumption, and network overload, which limit the performance of the MEC network and may cause data privacy and security issues.
By constructing a deep semantic authorization network, obtaining semantic symbols, semantic rate and semantic similarity, combining the binary offloading method to calculate the transmission delay and energy consumption, constructing a comprehensive cost formula, and using the deep reinforcement learning algorithm to perform two-layer optimization conversion, task offloading and resource optimization are achieved.
It effectively reduces the overall processing cost, improves task processing efficiency and edge computing performance, and enhances the intelligence level of the network and the accuracy of task processing.
Smart Images

Figure CN117785463B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of edge computing, and in particular to a method and device for task offloading and resource optimization for large model transmission. Background Art
[0002] With the rapid development of artificial intelligence and machine learning, large models have become a focus of research. Large models possess a large number of parameters and complex structures, enabling them to efficiently handle large amounts of data and complex problems. Mobile edge computing (MEC) is considered a revolutionary computing model. This model allows smart devices to offload latency-sensitive and energy-intensive tasks to nearby edge servers for rapid computation, improving overall performance. This not only helps reduce local latency and energy consumption, but also alleviates the burden on core networks.
[0003] However, MEC network performance is limited by limited bandwidth resources and the sheer volume of data transmitted. As more and more smart devices and applications generate massive amounts of data, MEC networks must strike a balance between real-time transmission and efficiency for large-scale model transmission. This poses challenges to their design and operation. First, the transmission of large models increases network load, potentially leading to congestion and increased transmission latency. Second, the transmission of large models also increases energy consumption. In MEC networks, edge devices and servers must process large amounts of data, resulting in high energy costs. To mitigate this burden, effective energy management strategies must be developed to ensure that energy costs are reduced while meeting performance requirements. Furthermore, the transmission of large models can raise data privacy and security concerns, potentially compromising private information during data transmission. Summary of the Invention
[0004] The present invention provides a task offloading and resource optimization method and device for large model transmission to solve the technical problems of excessive transmission delay and energy consumption and heavy network burden when offloading edge tasks of large model transmission.
[0005] In order to solve the above technical problems, the present invention provides a task offloading and resource optimization method for large model transmission, comprising:
[0006] Acquire a deep semantic authorization network consisting of preset devices and edge computing nodes, obtain semantic symbols, semantic rates, and semantic similarities based on the devices, and transmit the semantic symbols to the edge computing nodes;
[0007] Calculating transmission delay and energy consumption according to the semantic rate and a preset binary offloading method, and calculating a comprehensive cost according to the transmission delay and the energy consumption;
[0008] A first cost formula is calculated based on the comprehensive cost and the semantic similarity, and a double-layer optimization conversion is performed on the first cost formula to obtain an upper-layer optimization problem and a lower-layer optimization problem;
[0009] According to the upper-layer optimization problem, an offloading decision is determined by a preset deep reinforcement learning algorithm, and after obtaining a first semantic rate according to a preset method, the task offloading is completed according to the edge computing node, the semantic symbol, the offloading decision, and the first semantic rate;
[0010] Determining a first formula and a second formula based on the lower-layer optimization problem and preset scenario conditions, calculating optimization data based on the first formula or the second formula, and optimizing the deep semantic authorization network based on the optimization data to complete resource optimization;
[0011] Semantic symbols, semantic rates, and semantic similarities are obtained through a preset deep semantic authorization network. The transmission delay and energy consumption of the device are calculated using a binary offloading method to obtain a comprehensive cost. This helps to reasonably allocate tasks to devices or edge computing nodes based on the semantic characteristics of the tasks, thereby improving task processing efficiency. A first cost formula is constructed based on the comprehensive cost and semantic similarity, and a two-layer optimization conversion is performed to obtain the upper-layer optimization problem and the lower-layer optimization problem and perform corresponding processing. This can achieve effective task offloading and reasonable resource allocation, thereby reducing the overall processing cost, realizing task offloading and resource optimization based on semantic perception, and improving the efficiency and performance of edge computing.
[0012] As a preferred solution, obtaining the semantic symbol, the semantic rate, and the semantic similarity according to the device, and transmitting the semantic symbol to the edge computing node includes:
[0013] After the device generates a text task, the semantic symbols and the semantic rate are extracted from the text task according to a preset encoder, and the semantic similarity is calculated according to the semantic rate, and the semantic symbols are transmitted to the edge computing node through a wireless channel;
[0014] Obtaining semantic symbols, semantic rates, and semantic similarities through the device and transmitting the semantic symbols to the edge computing node can improve the ability to understand network text, assist in computing operations of the edge computing node, and improve the intelligence level of the system and the accuracy of task processing.
[0015] As a preferred solution, the transmission delay and energy consumption are calculated according to the semantic rate and the preset binary offloading method, and the comprehensive cost is calculated according to the transmission delay and the energy consumption, including:
[0016] A binary offloading factor is obtained by the binary offloading method, and the transmission delay is calculated based on the semantic rate, the number of word encoding bits in the device, the word bit calculation period, the computing power of the device, the computing power allocated to the device by the edge computing node, and the binary offloading factor;
[0017] Calculating the energy consumption according to the semantic rate, a preset energy consumption calculation coefficient, the computing capability of the device, the binary offloading factor, and the transmission delay;
[0018] Calculating the comprehensive cost based on the transmission delay and the energy consumption;
[0019] By calculating the transmission delay and energy consumption of the device and obtaining the comprehensive cost, it can be used for performance evaluation and optimization decision-making, thereby improving the performance and efficiency of the system.
[0020] As a preferred solution, the first cost formula calculated based on the comprehensive cost and the semantic similarity includes:
[0021] The first cost formula is calculated based on the comprehensive cost, the binary offloading factor, the number of semantic symbols of each word extracted from the device, the computing power allocated to the device by the edge computing node, the bandwidth allocated to the device by the deep semantic authorization network, the computing power of the device, and the semantic similarity; wherein the first cost formula is a cost minimization formula;
[0022] By comprehensively considering cost and semantic similarity, task processing and resource allocation can be optimized to achieve the goal of cost minimization and improve the efficiency and performance of the system.
[0023] As a preferred solution, the above-mentioned method of determining the offloading decision by a preset deep reinforcement learning algorithm based on the upper-level optimization problem includes:
[0024] Constructing a Markov decision process according to the upper-level optimization problem, and using a preset time difference as a loss function, and making the unloading decision according to the Markov decision process;
[0025] By processing the upper-level optimization problem, intelligent task offloading decisions and semantic coding strategy acquisition can be achieved, which can improve system performance and processing efficiency.
[0026] As a preferred solution, the obtaining of the first semantic rate according to a preset method includes:
[0027] When the device determines to offload the task, determining that the semantic similarity is not greater than a preset value, and setting a maximum semantic rate obtained by exhaustive enumeration as the first semantic rate;
[0028] By controlling semantic similarity, determining the maximum semantic rate, and optimizing task offloading, we can improve computing efficiency, save resources, and achieve flexibility and adjustability of task offloading.
[0029] As a preferred solution, completing task offloading according to the edge computing node, the semantic symbol, the offloading decision, and the first semantic rate includes:
[0030] The edge computing node obtains the semantic symbol according to a preset wireless channel, and reconstructs and calculates the semantic symbol by a preset semantic decoder according to the offloading decision and the first semantic rate, and offloads the task according to the calculation result;
[0031] Through the acquisition, reconstruction and calculation of semantic symbols at the edge computing nodes, combined with offloading decisions and the first semantic rate, task offloading can be achieved, which helps to optimize semantic processing and resource utilization and improve computing efficiency.
[0032] As a preferred solution, determining the first formula and the second formula according to the lower-level optimization problem and the preset scenario conditions includes:
[0033] When the scenario condition is determined to prioritize system performance, the lower-level optimization problem is converted into a convex optimization problem with linear constraints to determine the first formula; when the scenario condition is determined to prioritize algorithm efficiency and low complexity, the lower-level optimization problem is converted into an upper bound problem to determine the second formula;
[0034] According to different scenario conditions, converting the lower-level optimization problem into different forms of optimization problems and determining the calculation formula can improve system performance, reduce system latency, and enable the system to quickly respond to tasks and provide efficient calculation results under limited computing resources.
[0035] As a preferred solution, obtaining optimization data according to the first formula or the second formula, and optimizing the deep semantic authorization network using the optimization data to complete resource optimization includes:
[0036] Solve the first formula using a preset convex optimization toolkit to obtain a solution to the convex optimization problem; and solve the second formula to obtain a closed-form solution for bandwidth resources and computing power resources.
[0037] Obtaining the optimization data according to the convex optimization problem solution and the closed-form solution, and optimizing the bandwidth resources and computing power resources of the deep semantic authorization network according to the optimization data;
[0038] By obtaining solutions to convex optimization problems and closed-form solutions and obtaining optimized data, the bandwidth resources and computing power resources of the deep semantic authorization network can be optimized, thereby improving system performance, reducing latency, and meeting the requirements of prioritizing system performance and algorithm efficiency.
[0039] Accordingly, the present invention also provides a task offloading and resource optimization device for large model transmission, comprising: a network acquisition module, a cost calculation module, an optimization conversion module, a task offloading module and a resource optimization module;
[0040] The network acquisition module is used to acquire a deep semantic authorization network composed of preset devices and edge computing nodes, obtain semantic symbols, semantic rates and semantic similarities according to the devices, and transmit the semantic symbols to the edge computing nodes;
[0041] The cost calculation module is used to calculate the transmission delay and energy consumption according to the semantic rate and the preset binary offloading method, and calculate the comprehensive cost according to the transmission delay and the energy consumption;
[0042] The optimization conversion module is used to calculate a first cost formula based on the comprehensive cost and the semantic similarity, and perform a two-layer optimization conversion on the first cost formula to obtain an upper-layer optimization problem and a lower-layer optimization problem;
[0043] The task offloading module is configured to determine an offloading decision according to the upper-layer optimization problem using a preset deep reinforcement learning algorithm, and after obtaining a first semantic rate according to a preset method, complete task offloading according to the edge computing node, the semantic symbol, the offloading decision, and the first semantic rate;
[0044] The resource optimization module is configured to determine a first formula and a second formula based on the lower-layer optimization problem and preset scenario conditions, calculate optimization data based on the first formula or the second formula, and optimize the deep semantic authorization network based on the optimization data to complete resource optimization;
[0045] The task offloading and resource optimization device for large model transmission obtains semantic symbols, semantic rates and semantic similarities through a preset deep semantic authorization network, and calculates the transmission delay and energy consumption of the device in combination with the binary offloading method to obtain a comprehensive cost, which helps to reasonably allocate tasks to devices or edge computing nodes according to the semantic characteristics of the tasks, thereby improving task processing efficiency; constructs a first cost formula based on the comprehensive cost and semantic similarity and performs a two-layer optimization conversion to obtain the upper-layer optimization problem and the lower-layer optimization problem and perform corresponding processing, which can realize effective task offloading and reasonable allocation of resources, thereby reducing the overall processing cost, realizing task offloading and resource optimization based on semantic perception, and improving the efficiency and performance of edge computing. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 : A flowchart of an embodiment of a noise library construction method based on cloud-edge-device collaboration in the present invention;
[0047] Figure 2 : A structural diagram of an embodiment of a noise library construction device based on cloud-edge-device collaboration in the present invention;
[0048] Figure 3 : A schematic diagram of a model system for a deep semantic authorization network in the present invention;
[0049] Figure 4 : This is a graph showing the convergence of the algorithm proposed in the present invention in a Python simulation environment under different allocation criteria;
[0050] Figure 5 : A comparison chart of the algorithm proposed in the present invention with different allocation strategies in a Python simulation environment when the weighting factor is used as a variable;
[0051] Figure 6 : A comparison diagram of the algorithm proposed in the present invention with different allocation strategies in a Python simulation environment when semantic coding is used as a variable;
[0052] Figure 7 : This is a comparison diagram of the algorithm proposed in the present invention with different allocation strategies in a Python simulation environment when the transmitter is used as a variable. DETAILED DESCRIPTION
[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0054] In order to facilitate symbolic writing, the embodiments of the present invention use "all local", "all offloading", "brute force search", "I-optimization algorithm" and "II-optimization algorithm" to represent the strategy of all tasks being calculated locally, the strategy of all tasks being calculated at the edge computing node end, exhausting all possible offloading strategies and each offloading rate passing criterion II to obtain an optimal offloading and resource allocation strategy, and combining deep reinforcement learning to obtain the offloading strategy, while applying criteria I and criteria II to optimize the bandwidth and computing power resource allocation strategy.
[0055] Example 1
[0056] Please refer to Figure 1 , which is a flow chart of an embodiment of a task offloading and resource optimization method for large model transmission provided by the present invention, including steps 101-105, each of which is specifically as follows:
[0057] Step 101: Acquire a deep semantic authorization network consisting of preset devices and edge computing nodes, obtain semantic symbols, semantic rates, and semantic similarities based on the devices, and transmit the semantic symbols to the edge computing nodes.
[0058] In this embodiment, obtaining the semantic symbol, the semantic rate, and the semantic similarity according to the device, and transmitting the semantic symbol to the edge computing node includes:
[0059] After the device generates a text task, the semantic symbol and the semantic rate are extracted from the text task according to a preset encoder, and after the semantic similarity is calculated according to the semantic rate, the semantic symbol is transmitted to the edge computing node through a wireless channel.
[0060] In this embodiment, a deep semantic authorization network consisting of M devices and one edge computing node is created. The device set can be expressed as Among them, device m will generate a label l m Text tasks. m Contains d m sentences, each sentence contains N words on average, and device m uses the encoder to extract the words from l m Extract k for each word m The edge computing node receives the semantic symbols and uses the corresponding semantic decoder to reconstruct the original sentence and perform calculations to obtain the calculation results.
[0061] In this embodiment, the deep semantic authorization network uses semantic rate as an indicator to evaluate transmission capacity. Under the premise of considering semantic similarity, the semantic rate includes:
[0062]
[0063] Among them, c m is the semantic rate; is the bandwidth allocated to device m; G is the total amount of semantic information carried in each sentence transmitted to the edge computing node; ψ(γ m ,k m ) means that when the signal-to-noise ratio is γ m And each word is encoded as k m The semantic similarity when the number of symbols is .
[0064] In this embodiment, the signal-to-noise ratio includes:
[0065]
[0066] Among them, γ m is the signal-to-noise ratio, is the transmission power of device m; δ 2 is the variance of additive white Gaussian noise (AWGN); h m ~CN(0,α) is the channel parameter.
[0067] Step 102: Calculate the transmission delay and energy consumption according to the semantic rate and the preset binary offloading method, and calculate the comprehensive cost according to the transmission delay and the energy consumption.
[0068] In this embodiment, the transmission delay and energy consumption are calculated according to the semantic rate and the preset binary offloading method, and the comprehensive cost is calculated according to the transmission delay and the energy consumption, including:
[0069] A binary offloading factor is obtained by the binary offloading method, and the transmission delay is calculated based on the semantic rate, the number of word encoding bits in the device, the word bit calculation cycle, the computing power of the device, the computing power allocated to the device by the edge computing node and the binary offloading factor.
[0070] The energy consumption is calculated according to the semantic rate, a preset energy consumption calculation coefficient, the computing capability of the device, the binary offloading factor, and the transmission delay.
[0071] In this embodiment, ψ(γ m ,k m ) is set to ψ m , the transmission delay and energy consumption of device m include:
[0072]
[0073]
[0074] in, is the transmission delay of device m; is the energy consumption of device m.
[0075] In this embodiment, a binary offloading method is used to perform calculations on device m. When the task of device m is calculated locally, the corresponding local transmission delay includes:
[0076]
[0077] in, is the transmission delay of the task of device m when it is calculated locally; ω is the average number of bits encoded per word; ∈ is the CPU cycle required to calculate each bit; is the computing capability of device m.
[0078] In this embodiment, when the task of device m is calculated at the edge computing node, the corresponding transmission delay includes:
[0079]
[0080] in, The computing power allocated to device m by the edge computing node.
[0081] In this embodiment, when the task of device m is calculated locally, its corresponding local energy consumption includes:
[0082]
[0083] in, is the energy consumption of the task of device m when it is calculated locally; k is the energy consumption calculation coefficient.
[0084] In this embodiment, since the edge computing node has a stable power supply, its energy consumption is negligible, so I m The overall transmission delay is:
[0085]
[0086] Among them, T m For I m The overall transmission delay of m is a binary unloading factor.
[0087] In this embodiment, I m The overall energy consumption is:
[0088]
[0089] Among them, E m For I m overall energy consumption.
[0090] In this embodiment, since the tasks of each device are independent of each other and can be parallelized at the edge computing node, the overall transmission delay and overall energy consumption of the system are:
[0091]
[0092]
[0093] Among them, T is the overall transmission delay of the system; E is the overall energy consumption of the system.
[0094] In this embodiment, the comprehensive cost is calculated based on the transmission delay and the energy consumption, including:
[0095] Ψ=ηT+(1-η)E
[0096] Where Ψ is the comprehensive cost, and η∈[0, 1] is the weighting factor used to evaluate the system transmission delay and system energy consumption.
[0097] Step 103: Calculate a first cost formula based on the comprehensive cost and the semantic similarity, and perform a two-layer optimization conversion on the first cost formula to obtain an upper-layer optimization problem and a lower-layer optimization problem.
[0098] In this embodiment, the first cost formula is obtained by calculating the comprehensive cost and the semantic similarity, including:
[0099] The first cost formula is calculated based on the comprehensive cost, the binary offloading factor, the number of semantic symbols of each word extracted from the device, the computing power allocated to the device by the edge computing node, the bandwidth allocated to the device by the deep semantic authorization network, the computing power of the device and the semantic similarity; wherein, the first cost formula is a cost minimization formula.
[0100] In this embodiment, the first cost formula includes:
[0101]
[0102] stC1:
[0103] C2:
[0104] C3:
[0105] C4:
[0106] C5:
[0107] C6:
[0108] Where w is the total system bandwidth; F is the total computing power of the edge computing node; β m is a binary unloading factor; k m is the number of semantic symbols of each word extracted by device m; The computing power allocated to devices by edge computing nodes; allocating bandwidth to the device for a deep semantic authorization network; The computing power of the device; and are the minimum and maximum values of computing power of device m, respectively; ψ m When the signal-to-noise ratio is γm And each word is encoded as k m The semantic similarity when the number of symbols is ψ th is the preset semantic similarity threshold.
[0109] In this embodiment, the first cost formula is subjected to a two-layer optimization conversion, including:
[0110]
[0111] StC1, C2, C6
[0112]
[0113] stC3, C4, C5
[0114] in, It is an upper-layer optimization problem used to solve the task offloading problem; It is a lower-level optimization problem used to solve resource optimization problems.
[0115] Step 104: According to the upper-layer optimization problem, the offloading decision is determined by a preset deep reinforcement learning algorithm, and after obtaining the first semantic rate according to a preset method, the task offloading is completed according to the edge computing node, the semantic symbol, the offloading decision and the first semantic rate.
[0116] In this embodiment, the offloading decision is determined by a preset deep reinforcement learning algorithm based on the upper-level optimization problem, including:
[0117] A Markov decision process is constructed according to the upper-level optimization problem, and a preset time difference is used as a loss function. The unloading decision is made according to the Markov decision process.
[0118] In this embodiment, the agent takes actions by observing the environment, and the environment rewards the agent and enables it to enter a new state. The agent continuously interacts with the environment to obtain an unloading strategy. The states that the agent can perceive include:
[0119] s k ={h k , α k-1 , d k}
[0120] Among them, s k is the state that the agent can perceive; h k is the vector of channel parameters at time k; α k-1 is the vector of the unloading strategy at time k-1; d k is the vector of the task sentence at time k.
[0121] In this embodiment, when the agent observes the state s k When , actions are selected according to the greedy criterion, including:
[0122]
[0123] Among them, a t is the selected action; ε is the search factor; Q(s k , a) is the network that approximates the Q function, which represents the cumulative reward after executing action a; θ is the neural network parameter.
[0124] In this embodiment, when the agent performs action α k , the system will give a reward, including:
[0125]
[0126] Among them, X1, X2 and X3 are all positive numbers.
[0127] In this embodiment, after several interactions with the environment, the agent uses temporal difference (TD) as the loss function, which is expressed as:
[0128]
[0129] Where L(φ) is the loss function; ξ is the discount factor; For the target network.
[0130] In this embodiment, obtaining the first semantic rate according to a preset method includes:
[0131] When the device determines to offload the task, it determines that the semantic similarity is not greater than a preset value, and sets the maximum semantic rate obtained by exhaustive enumeration as the first semantic rate.
[0132] In this embodiment, when device m determines to uninstall task I m When the exhaustive method is used to obtain the uninstall decision k m , and ensure the semantic similarity ψ m ≥ψ th Based on this, the first semantic rate c m maximum.
[0133] In this embodiment, completing task offloading according to the edge computing node, the semantic symbol, the offloading decision, and the first semantic rate includes:
[0134] The edge computing node obtains the semantic symbol according to the preset wireless channel, and according to the offloading decision and the first semantic rate, the preset semantic decoder reconstructs and calculates the semantic symbol, and then offloads the task according to the calculation result.
[0135] Step 105: Confirm the first formula and the second formula according to the lower-level optimization problem and the preset scenario conditions, calculate and obtain optimization data according to the first formula or the second formula, and optimize the deep semantic authorization network based on the optimization data to complete resource optimization.
[0136] In this embodiment, determining the first formula and the second formula according to the underlying optimization problem and the preset scenario conditions includes:
[0137] When the scenario condition is determined to prioritize system performance, the lower-level optimization problem is converted into a convex optimization problem with linear constraints, and the first formula is determined; when the scenario condition is determined to prioritize algorithm efficiency and low complexity, the lower-level optimization problem is converted into an upper bound problem, and the second formula is determined.
[0138] In this embodiment, obtaining optimization data according to the first formula or the second formula, and optimizing the deep semantic authorization network using the optimization data to complete resource optimization includes:
[0139] Using the CVX toolkit as a convex optimization toolkit, the first formula is solved to obtain a solution to the convex optimization problem; and based on solving the second formula, a closed-form solution for bandwidth resources and computing power resources is obtained.
[0140] The optimization data is obtained according to the solution of the convex optimization problem and the closed-form solution, and the bandwidth resources and computing power resources of the deep semantic authorization network are optimized according to the optimization data.
[0141] In this embodiment, when the scenario conditions determine that algorithm efficiency is prioritized and low complexity is prioritized, the lower-level optimization problem is converted into an upper-bound problem, including:
[0142]
[0143] stC3, C4, C5
[0144] in, Optimize results for optimal resources; S l and S o is a set of devices that perform calculations locally and at edge computing nodes.
[0145] From this we can see that local computing power Just and Regarding bandwidth and edge computing node computing power Just and related.
[0146] In this embodiment, for Conduct analytical derivation, including:
[0147]
[0148] in, It is the optimal solution for local computing power.
[0149] In this embodiment, for Construct a Lagrangian function with its constraints, including:
[0150]
[0151] in, is the Lagrangian function, μ and v are the two Lagrangian multiplier coefficients.
[0152] For variables Find the first-order partial derivatives of μ and v, including:
[0153]
[0154]
[0155]
[0156]
[0157] Based on the above first-order partial derivative solution, we can obtain the closed-form solution of bandwidth resources and computing power resources, including:
[0158]
[0159]
[0160] in, is the closed-form solution for bandwidth resources, It is a closed-form solution for computing resources.
[0161] Please refer to Figure 3 , which is a schematic diagram of a model system for deep semantic authorization network in the present invention.
[0162] In this embodiment, the Deep Semantic Authorization Network adopts the DeepSC framework, where the encoder and decoder of DeepSC are integrated on the user side and the EP side respectively to achieve efficient semantic communication.
[0163] In this embodiment, the system model is configured with one edge computing node (EP) and ten devices during implementation. Each device and EP has deployed a semantic encoder and decoder to achieve semantic communication. In addition, in the network environment simulation, all channels experience Rayleigh flat fading. Unless otherwise specified, the average gain of each device channel is 1, the AWGN noise power is 0.01W, and the transmission power of each user terminal is 0.1W. In this network environment, the number of task sentences n for device m is m Obey a uniform distribution, including:
[0164] d m ~u(100+50m, 150+50m)
[0165] Each sentence consists of an average of 20 words; each word can be encoded into 40 bits for calculation, or encoded into k m ~[1, 20] semantic symbols for semantic transmission; the semantic transmission threshold is set to 0.85; the total bandwidth of the system is 1MHz, and the total computing power of the EP end is 2.5×10 9 cycle / s; the computing power of each device meets
[0166] Please refer to Figure 4 , which is a convergence performance diagram of the algorithm proposed in the present invention in the Python simulation environment under different allocation criteria.
[0167] In this embodiment, the algorithms under the two criteria achieve good convergence performance after 50 training rounds; after 100 training rounds, the effect of the "II-optimization algorithm" is very close to that of the "I-optimization algorithm", verifying the effectiveness of the algorithm in the embodiment of the present invention.
[0168] Please refer to Figure 5 , which is a comparison result diagram of the algorithm proposed in the present invention with different allocation strategies in a Python simulation environment when the weighting factor is used as a variable.
[0169] In this embodiment, Figure 3 It can be seen that when the weighting factor η is different but other parameter conditions remain consistent, no matter what the weighting factor η is and under what resource optimization criteria, the strategy proposed in the embodiment of the present invention performs better than other strategy algorithms. It can be concluded that the algorithm proposed in the embodiment of the present invention can flexibly adapt to the changes in the deep semantic authorization network, thereby making accurate decisions, thereby improving the overall performance of the deep semantic authorization network system.
[0170] Please refer to Figure 6, which is a comparison result diagram of the algorithm proposed in the present invention with different allocation strategies in a Python simulation environment when semantic coding is used as a variable.
[0171] In this embodiment, Figure 4 It can be seen that in the semantic encoding k m When K increases to 20, while other parameters remain the same, the system cost of the algorithm proposed in this embodiment of the present invention gradually decreases. This shows that a larger K value helps improve semantic similarity, further improving the quality of semantic transmission and avoiding task upload failures. In this scenario, a larger K value allows for more flexible adaptation to changing network environments, thereby improving network robustness.
[0172] Please refer to Figure 7 , which is a comparison result diagram of the algorithm proposed in the present invention with different allocation strategies in a Python simulation environment when the transmitting end is used as a variable.
[0173] In this embodiment, Figure 5 It can be seen that when the SNR at the transmitter is different but other parameters remain the same, the comprehensive cost of different task offloading strategies shows a downward trend as the SNR at the transmitter increases. It can be concluded that a higher SNR helps to increase the transmission rate, which in turn enables the device to more fully transmit tasks to the EP, thereby fully unleashing its powerful computing resources.
[0174] In summary, an embodiment of the present invention provides a task offloading and resource optimization method for large model transmission. Semantic symbols, semantic rates and semantic similarities are obtained through a preset deep semantic authorization network. The transmission delay and energy consumption of the device are calculated in combination with a binary offloading method to obtain a comprehensive cost. This helps to reasonably allocate tasks to devices or edge computing nodes according to the semantic characteristics of the tasks, thereby improving task processing efficiency. A first cost formula is constructed based on the comprehensive cost and semantic similarity, and a two-layer optimization conversion is performed to obtain the upper-layer optimization problem and the lower-layer optimization problem and perform corresponding processing. This can achieve effective task offloading and reasonable resource allocation, thereby reducing the overall processing cost, realizing task offloading and resource optimization based on semantic perception, and improving the efficiency and performance of edge computing.
[0175] Example 2
[0176] Please refer to Figure 2 , which is a structural diagram of an embodiment of a task offloading and resource optimization device for large model transmission provided by the present invention, including: a network acquisition module 201, a cost calculation module 202, an optimization conversion module 203, a task offloading module 204 and a resource optimization module 205.
[0177] Among them, the network acquisition module 201 is used to obtain a deep semantic authorization network composed of preset devices and edge computing nodes, obtain semantic symbols, semantic rates and semantic similarities according to the devices, and transmit the semantic symbols to the edge computing nodes.
[0178] The cost calculation module 202 is configured to calculate the transmission delay and energy consumption according to the semantic rate and a preset binary offloading method, and calculate the comprehensive cost according to the transmission delay and the energy consumption.
[0179] The optimization conversion module 203 is used to calculate a first cost formula based on the comprehensive cost and the semantic similarity, and perform a two-layer optimization conversion on the first cost formula to obtain an upper-layer optimization problem and a lower-layer optimization problem.
[0180] The task offloading module 204 is used to determine the offloading decision according to the upper-level optimization problem by a preset deep reinforcement learning algorithm, and after obtaining the first semantic rate according to a preset method, complete the task offloading according to the edge computing node, the semantic symbol, the offloading decision and the first semantic rate.
[0181] The resource optimization module 205 is used to confirm the first formula and the second formula according to the lower-level optimization problem and the preset scenario conditions, calculate the optimization data according to the first formula or the second formula, and optimize the deep semantic authorization network according to the optimization data to complete resource optimization.
[0182] In an embodiment of the present invention, the task offloading and resource optimization device for large model transmission obtains semantic symbols, semantic rates and semantic similarity through a preset deep semantic authorization network, and calculates the transmission delay and energy consumption of the device in combination with the binary offloading method to obtain a comprehensive cost, which helps to reasonably allocate tasks to devices or edge computing nodes according to the semantic characteristics of the tasks, thereby improving task processing efficiency; constructing a first cost formula based on the comprehensive cost and semantic similarity and performing a two-layer optimization conversion to obtain the upper-level optimization problem and the lower-level optimization problem and perform corresponding processing, which can achieve effective task offloading and reasonable allocation of resources, thereby reducing the overall processing cost, realizing task offloading and resource optimization based on semantic perception, and improving the efficiency and performance of edge computing.
[0183] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
[0184] In addition, in the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of this application, unless otherwise specified, "plurality" means two or more.
Claims
1. A task offloading and resource optimization method for large model transmission, characterized in that: include: Acquire a deep semantic authorization network consisting of preset devices and edge computing nodes, obtain semantic symbols, semantic rates, and semantic similarities based on the devices, and transmit the semantic symbols to the edge computing nodes; Calculating transmission delay and energy consumption according to the semantic rate and a preset binary offloading method, and calculating a comprehensive cost according to the transmission delay and the energy consumption; A first cost formula is obtained by calculating the comprehensive cost and the semantic similarity, including: obtaining a binary offloading factor by the binary offloading method, and calculating the transmission delay according to the semantic rate, the number of word encoding bits in the device, the word bit calculation cycle, the computing power of the device, the computing power allocated to the device by the edge computing node and the binary offloading factor; calculating the energy consumption according to the semantic rate, a preset energy consumption calculation coefficient, the computing power of the device, the binary offloading factor and the transmission delay; calculating the comprehensive cost according to the transmission delay and the energy consumption; calculating the first cost formula according to the comprehensive cost, the binary offloading factor, the number of semantic symbols extracted from each word in the device, the computing power allocated to the device by the edge computing node, the bandwidth allocated to the device by the deep semantic authorization network, the computing power of the device and the semantic similarity; wherein, the first cost formula is a cost minimization formula; and performing a two-layer optimization conversion on the first cost formula to obtain an upper-layer optimization problem and a lower-layer optimization problem; According to the upper-layer optimization problem, an offloading decision is determined by a preset deep reinforcement learning algorithm, and after obtaining a first semantic rate according to a preset method, task offloading is completed according to the edge computing node, the semantic symbol, the offloading decision, and the first semantic rate, including: the edge computing node obtains the semantic symbol according to a preset wireless channel, and after reconstructing and calculating the semantic symbol by a preset semantic decoder according to the offloading decision and the first semantic rate, the task is offloaded according to the calculation result; Confirm a first formula and a second formula according to the lower-level optimization problem and preset scenario conditions, calculate optimization data according to the first formula or the second formula, and optimize the deep semantic authorization network according to the optimization data to complete resource optimization, including: when the scenario condition is determined to prioritize system performance, convert the lower-level optimization problem into a convex optimization problem with linear constraints, and determine the first formula; when the scenario condition is determined to prioritize algorithm efficiency and low complexity, convert the lower-level optimization problem into an upper bound problem, and determine the second formula; solve the first formula according to a preset convex optimization toolkit to obtain a solution to the convex optimization problem; and obtain a closed-form solution to bandwidth resources and computing power resources by solving the second formula; obtain the optimization data according to the convex optimization problem solution and the closed-form solution, and optimize the bandwidth resources and computing power resources of the deep semantic authorization network according to the optimization data.
2. A method for task offloading and resource optimization for large model transmission according to claim 1, characterized in that: The obtaining of a semantic symbol, a semantic rate, and a semantic similarity according to the device, and transmitting the semantic symbol to the edge computing node includes: After the device generates a text task, the semantic symbol and the semantic rate are extracted from the text task according to a preset encoder, and after the semantic similarity is calculated according to the semantic rate, the semantic symbol is transmitted to the edge computing node through a wireless channel.
3. The method for task offloading and resource optimization for large model transmission according to claim 1, characterized in that: The method of determining the offloading decision based on the upper-level optimization problem by a preset deep reinforcement learning algorithm includes: A Markov decision process is constructed according to the upper-level optimization problem, and a preset time difference is used as a loss function. The unloading decision is made according to the Markov decision process.
4. The method for task offloading and resource optimization for large model transmission according to claim 1, characterized in that: The obtaining of the first semantic rate according to a preset method includes: When the device determines to offload the task, it determines that the semantic similarity is not greater than a preset value, and sets the maximum semantic rate obtained by exhaustive enumeration as the first semantic rate.
5. A task offloading and resource optimization device for large model transmission, characterized in that: include: Network acquisition module, cost calculation module, optimization conversion module, task offloading module and resource optimization module; The network acquisition module is used to acquire a deep semantic authorization network composed of preset devices and edge computing nodes, obtain semantic symbols, semantic rates and semantic similarities according to the devices, and transmit the semantic symbols to the edge computing nodes; The cost calculation module is used to calculate the transmission delay and energy consumption according to the semantic rate and the preset binary offloading method, and calculate the comprehensive cost according to the transmission delay and the energy consumption; The optimization conversion module is used to calculate a first cost formula based on the comprehensive cost and the semantic similarity, and perform a two-layer optimization conversion on the first cost formula to obtain an upper-level optimization problem and a lower-level optimization problem; wherein, the first cost formula includes: obtaining a binary unloading factor by the binary unloading method, and calculating the transmission delay based on the semantic rate, the number of word encoding bits in the device, the word bit calculation cycle, the computing power of the device, the computing power allocated to the device by the edge computing node and the binary unloading factor; calculating the energy consumption based on the semantic rate, the preset energy consumption calculation coefficient, the computing power of the device, the binary unloading factor and the transmission delay; calculating the comprehensive cost based on the transmission delay and the energy consumption; calculating the first cost formula based on the comprehensive cost, the binary unloading factor, the number of semantic symbols extracted from each word in the device, the computing power allocated to the device by the edge computing node, the bandwidth allocated to the device by the deep semantic authorization network, the computing power of the device and the semantic similarity; wherein, the first cost formula is a cost minimization formula The task offloading module is used to determine the offloading decision according to the upper-layer optimization problem by a preset deep reinforcement learning algorithm, and after obtaining the first semantic rate according to a preset method, complete the task offloading according to the edge computing node, the semantic symbol, the offloading decision and the first semantic rate, including: the edge computing node obtains the semantic symbol according to a preset wireless channel, and according to the offloading decision and the first semantic rate, a preset semantic decoder reconstructs and calculates the semantic symbol, and then offloads the task according to the calculation result; The resource optimization module is used to confirm the first formula and the second formula according to the lower-level optimization problem and the preset scenario conditions, calculate the optimization data according to the first formula or the second formula, and optimize the deep semantic authorization network according to the optimization data to complete resource optimization, including: when the scenario condition is determined to prioritize system performance, convert the lower-level optimization problem into a convex optimization problem with linear constraints, and determine the first formula; when the scenario condition is determined to prioritize algorithm efficiency and low complexity, convert the lower-level optimization problem into an upper bound problem, and determine the second formula; solve the first formula according to the preset convex optimization toolkit to obtain the solution of the convex optimization problem; and obtain the closed-form solution of bandwidth resources and computing power resources according to the solution of the second formula; obtain the optimization data according to the convex optimization problem solution and the closed-form solution, and optimize the bandwidth resources and computing power resources of the deep semantic authorization network according to the optimization data.