An end-edge collaborative inference system latency optimization method based on fluid antenna assistance

By using a fluid antenna-assisted end-to-side collaborative inference system, the antenna position is dynamically adjusted and the DNN partitioning and resource allocation are optimized, which solves the problems of latency and accuracy in wireless transmission environments and achieves low-latency, high-accuracy collaborative inference.

CN122138214APending Publication Date: 2026-06-02ZHEJIANG UNIV OF TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG UNIV OF TECH
Filing Date
2026-03-30
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing edge-to-edge collaborative DNN inference technology faces the problem of increased latency and decreased accuracy due to the impact of IFD quality in wireless transmission environments. Furthermore, FA technology is not deeply integrated with the collaborative inference framework, making it impossible to simultaneously achieve low latency and high accuracy optimization.

Method used

The end-to-side collaborative inference system with fluid antenna assistance dynamically adjusts the antenna position and jointly optimizes DNN partitioning, resource allocation and transmission parameters to build a collaborative model to minimize the total inference latency while meeting accuracy and energy consumption constraints.

Benefits of technology

It significantly reduces the total inference latency by 32 percent, meets the requirements for high accuracy, and has a significant performance improvement, especially under adverse channel conditions, achieving low latency and high accuracy collaborative inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122138214A_ABST
    Figure CN122138214A_ABST
Patent Text Reader

Abstract

This invention discloses a latency optimization method for end-edge collaborative inference systems based on fluid antenna assistance (FA). The method first establishes an FA-assisted system model and constructs DNN partitioning collaboration, IFD transmission, and inference accuracy models based on the system model to determine the total inference latency, including local inference, IFD wireless transmission, and edge inference delays. Then, a joint optimization problem is constructed with the objective of minimizing the total inference latency, jointly optimizing parameters such as DNN partitioning points, transmit power of each device, and computational resource allocation. The joint optimization problem must satisfy inference accuracy and energy consumption constraints. Finally, the problem is decomposed into four sub-problems, and the block coordinate descent method is used to alternately optimize until convergence, completing the latency optimization of the end-edge collaborative inference system. This invention is the first to integrate FA and collaborative inference frameworks, achieving multi-dimensional parameter joint optimization, effectively reducing the total inference latency, and demonstrating significant performance advantages in scenarios with high accuracy thresholds and adverse channel conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of mobile edge computing and AIoT technology, specifically relating to a latency optimization method for an end-to-edge collaborative inference system based on fluid antenna assistance. Background Technology

[0002] In recent years, with the rapid development of 6G mobile communication networks, deep neural networks (DNNs) have played an increasingly important role in applications such as facial recognition, autonomous driving, and augmented reality. These applications typically require not only low latency but also high-precision processing. However, DNN models are inherently computationally intensive and memory-intensive. If wireless devices perform inference entirely locally, they will face enormous computational burdens and energy consumption, making it difficult to meet the needs of practical applications. To address these challenges, Collaborative DNN Inference (CDI) technology has emerged. This technology effectively distributes the computational load by partitioning the DNN model into a head model and a tail model, which are deployed on the wireless device and the edge server (ES) respectively. Specifically, the wireless device executes the first few layers of the DNN to extract intermediate feature data (IFD), and then uploads the IFD to the ES via a wireless link. The server then executes the remaining layers to complete the final inference and returns the result, thereby significantly reducing the device's computational latency and energy consumption.

[0003] However, despite the significant advantages that CDI technology has demonstrated in computational load balancing, it still faces key challenges in practical wireless transmission environments. First, wireless transmission channels are susceptible to fading, noise, and interference, leading to increased IFD transmission delay and distortion, which severely reduces the final inference accuracy. Second, although most existing solutions have achieved some success in reducing latency and energy consumption, they still have significant limitations in real-world wireless environments. The combined impact of DNN split point selection and transmission environment on inference accuracy is neglected, especially in high-precision scenarios where a drop in accuracy directly affects the effectiveness of the task. Therefore, to improve the wireless transmission environment, Fluid Antenna (FA) technology, as an emerging wireless transmission enhancement method, has received widespread attention in the field of edge computing in recent years.

[0004] FA (Automatic Front-End) technology, by dynamically adjusting antenna positions within a confined space, achieves flexible adaptation to channel conditions, significantly improving link quality and spatial degrees of freedom. Unlike traditional fixed-position antennas, which are limited in their ability to utilize spatial degrees of freedom, FA allows for dynamic reconfiguration of antenna positions based on changing propagation conditions. This flexibility not only improves channel quality through optimal antenna positioning but also provides additional spatial degrees of freedom. Therefore, integrating FA into CDI (Collaborative Inference Dependency Injection) has the potential to enhance inference performance. While FA technology performs well in general edge computing offloading, it has not considered the direct impact of IFD (In-Function Distortion) distortion on accuracy in collaborative inference, nor has it explored the joint optimization of FA dynamic positioning with DNN partitioning and accuracy constraints. This makes it difficult to simultaneously achieve the goals of low latency and high accuracy in collaborative inference in practical deployments, especially in scenarios with harsh channel conditions or stringent accuracy requirements, limiting system performance improvements.

[0005] In practical applications, while existing end-edge collaborative DNN inference technology can alleviate the computational burden on devices to some extent, it is limited by the impact of the wireless transmission environment on the quality and accuracy of inference function decomposition (IFD), making it difficult to achieve optimal latency control while ensuring high-precision inference. Meanwhile, although FA (Automatic Facilitation) technology possesses significant channel enhancement potential, it has not yet been deeply integrated with the collaborative inference framework, failing to fully leverage its advantages in improving IFD transmission quality and jointly optimizing latency and accuracy. Therefore, a novel technical solution is urgently needed that can comprehensively consider the coupled impact of DNN partitioning and the transmission environment on accuracy in FA-assisted collaborative inference systems, and minimize the total inference latency through multi-dimensional resource joint optimization, while simultaneously satisfying accuracy and energy consumption constraints. This joint optimization is crucial for minimizing the overall system's inference latency. In this context, an FA-assisted end-edge collaborative inference system latency optimization method enables the system to make optimization decisions in more complex environments. Summary of the Invention

[0006] This invention aims to address the problem in existing CDI inference technologies where the impact of the wireless transmission environment on IFD quality leads to increased inference latency and decreased accuracy. It also overcomes the shortcomings of existing FA (Automatic Front-End) technologies, which are not deeply integrated with the cooperative inference framework and cannot simultaneously achieve low latency and high accuracy optimization. Specifically, this invention provides a latency optimization method for end-to-edge cooperative inference systems based on fluid antenna assistance. By introducing FA to enhance transmission performance, cooperative DNN inference is performed between resource-constrained wireless devices and base stations (BS) equipped with ES (Essence Base Stations). The method jointly optimizes DNN partitioning, resource allocation, and transmission parameters to minimize total inference latency while satisfying inference accuracy and device power consumption constraints.

[0007] The technical solution adopted in this invention is as follows: A delay optimization method for an end-to-side cooperative inference system based on fluid antenna assistance includes: S1. Establish a fluid antenna FA-assisted end-edge collaborative inference system model. The system model includes a base station BS equipped with an edge server ES and several single-antenna wireless devices. The BS is deployed with several FAs (Flying Array Components) that can dynamically adjust their positions within a predetermined linear space, and single-antenna wireless devices. Intermediate feature data (IFD) is transmitted to the BS using orthogonal spectrum resources; based on the system model, a DNN-based collaborative model, an FA-assisted IFD transmission model, and an inference accuracy model are constructed to determine the total inference delay of the system, which includes local inference delay, IFD wireless transmission delay, and edge inference delay; S2. Under the system model, construct a joint optimization problem with the goal of minimizing the total inference latency of the system. Jointly optimize the DNN partitioning point, the transmit power of each device, the allocation of computing resources, the receive beamforming vector at the BS and the antenna positioning vector at the FA. The joint optimization problem must satisfy the constraints of inference accuracy and device energy consumption. The device energy consumption includes local energy consumption and IFD wireless transmission energy consumption. S3. Decompose the joint optimization problem into several sub-problems, and use the block coordinate descent method to alternately optimize each sub-problem until convergence, thereby obtaining a suboptimal solution and completing the latency optimization of the end-edge collaborative reasoning system.

[0008] Furthermore, S1 constructs a collaborative model based on DNN partitioning, specifically including: The DNN segmentation method is adopted to divide the DNN model into a head model and a tail model based on the DNN segmentation points. The single-antenna wireless device runs the DNN head model and extracts the IFD, while the ES runs the DNN tail model to complete the inference. The segmentation points determine the local computing load, edge computing load, and IFD data volume. Based on the computing capabilities of the single-antenna wireless device and the ES, combined with the computing workload of each layer of the DNN, the local inference latency, local energy consumption, and edge inference latency are derived. At the same time, the maximum computing capability of the BS is limited to ensure that the BS has available computing resources.

[0009] Furthermore, in the computational workload of each layer of the DNN, the computational workload of the convolutional layer is... : ; in, and They are the first The height and width of the input feature map of device k in layer k. and They represent the first The number of channels in the input and output feature maps of the layer. This indicates the size of the convolution kernel; the computational workload of a fully connected layer is the product of the number of channels in the input feature map and the output feature map.

[0010] Furthermore, S1 constructs a FA-assisted IFD transmission model, specifically including: remember For the position of the nth FA, construct the steering vector of the FA array, use the narrowband channel model under the far-field uniform linear array to obtain the channel between the device and the BS, combine the Shannon formula to obtain the transmission rate of the IFD, and derive the IFD wireless transmission delay and IFD wireless transmission power consumption based on the IFD data volume.

[0011] Furthermore, the steering vector of the FA array is: ; In the formula, Indicates the wavelength of the signal. Define the signal angle of arrival and perform a transpose operation; The method of obtaining the channel between the device and the BS using the narrowband channel model under a far-field uniform linear array is as follows: ; In the formula, It is the channel gain at a reference distance of 1 meter between the device and the BS. Indicates device Distance between BS, This represents the path loss index.

[0012] Furthermore, step S1 involves constructing an inference accuracy model, specifically including: Based on experiments with the classic DNN model VGG16 on the CIFAR-10 image classification dataset, regression fitting techniques were used to obtain the relationship between inference accuracy and DNN split points, device... The closed-form expression for the receiver signal-to-noise ratio (SNR) is given, and an exponential function is used to characterize the correlation among the three factors. An inference accuracy threshold is set to establish inference accuracy constraints.

[0013] Furthermore, in S3, the joint optimization problem is specifically decomposed into four sub-problems: computational resource allocation, transmit power optimization, receive beamforming design, DNN partitioning and FA localization. Solving the computational resource allocation subproblem, specifically including: fixed equipment Based on the transmit power, receive beamforming vector at BS, DNN partitioning, and FA positioning, a subproblem of computational resource allocation with the objective of minimizing the total inference delay of the system is constructed. This subproblem is a convex optimization problem with constraints including ES computational resource constraints and FA location boundary constraints. A Lagrangian function is constructed, and the closed-form optimal solution of the computational resources allocated by ES to each device is derived using the stationary point condition and complementary relaxation condition of KKT conditions. Solving the transmit power optimization subproblem involves: fixed computational resource allocation, receiver beamforming vector of the base station (BS), DNN partitioning, and FA localization. The transmit power optimization subproblem is constructed with constraints including: inference accuracy constraints, equipment power consumption constraints, and upper and lower bounds on the equipment transmit power. Auxiliary variables are introduced. Instead of the transmission rate, the non-convex constraints are transformed into convex forms through equivalent transformation and convexization, and the optimal solution for the transmit power is solved using convex optimization tools; Solving the receive beamforming design subproblem involves: fixed computational resource allocation, equipment transmit power, DNN partitioning, and FA localization, to construct the receive beamforming design subproblem; and introducing auxiliary variables. and The method employs a successive convex approximation technique, using a first-order Taylor expansion to approximate the non-convex terms, iteratively solving the convex approximation problem until convergence to obtain the optimal solution for receiving beamforming. Constraints include: introduced auxiliary variables. and Constraints, inference accuracy constraints after equivalent convex transformation, and BS receiver power constraints. Solving the DNN partitioning and FA positioning subproblem involves: establishing a fixed computational resource allocation, equipment transmit power, and BS receive beamforming vector; and constructing the DNN partitioning and FA positioning subproblem with constraints including: inference accuracy constraints, equipment power consumption constraints, FA minimum spacing constraints, DNN partition point discreteness constraints, and edge computational resource non-negativity constraints. A penalty function method is used to reconstruct the constrained DNN partitioning and FA positioning subproblem into a standard unconstrained optimization problem, quantifying the degree of constraint violation by introducing a penalty term into the objective function. The Hybrid Gravity-Turned Turtle Group Search Intelligent Optimization Algorithm (HG3S) is then used to process the standard unconstrained optimization problem, thereby obtaining the optimal solution for the DNN partitioning points and the antenna positioning vectors of the FA.

[0014] Furthermore, the Hybrid Gravity-Turbot Swarm Search Intelligent Optimization Algorithm HG3S integrates the global exploration capability of the Gravity Search Algorithm GSA with the local development capability of the Turbot Swarm Algorithm SSA.

[0015] An electronic device includes a processor and a memory, the memory storing a program that, when executed, implements any of the methods described herein.

[0016] A computer-readable storage medium storing executable instructions that, when executed, implement any of the methods described herein.

[0017] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention introduces FA (Automatic Facilitation) technology into the end-edge collaborative inference framework for the first time in an FA-assisted collaborative DNN inference system. It explicitly considers the joint impact of wireless transmission impairments and DNN partitioning on inference accuracy. Existing collaborative inference research mostly focuses on minimizing latency or energy, neglecting accuracy constraints, leading to significant performance degradation under adverse channel conditions. This invention provides additional spatial degrees of freedom by dynamically adjusting the FA position, significantly improving the transmission rate and reception quality of the IFD (Inductively Coupled Defibrillator), thereby effectively mitigating the distortion of the final inference result caused by channel impairments while ensuring the inference accuracy threshold.

[0018] 2. Compared to traditional fixed antenna or reconfigurable smart surface-assisted collaborative inference schemes, and FA edge computing research applied only to computational offloading, this invention achieves multi-dimensional joint optimization of DNN partitioning, transmit power, computational resource allocation, receive beamforming, and antenna positioning, solving the problems of complex variable coupling and difficult non-convex optimization in existing technologies. Through efficient decomposition and solution using an alternating optimization algorithm, this invention exhibits stronger robustness and scalability in practical deployments.

[0019] 3. Simulation results show that the present invention can reduce the total inference latency by up to 32% while strictly meeting accuracy requirements. The performance gain is particularly significant in scenarios with high accuracy thresholds and degraded channel conditions, far surpassing fixed antenna, random antenna positioning, and equal baseline schemes. This advantage provides a more efficient and reliable solution for resource-constrained wireless devices to achieve low-latency, high-precision collaborative inference in sixth-generation intelligent applications. Attached Figure Description

[0020] Figure 1 This is a model diagram of a FA-assisted CDI system according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating a specific embodiment of the present invention. Figure 3 This is a diagram of segmentation points in a VGG network according to an embodiment of the present invention; Figure 4 This is a data fitting result of VGG16 on the CIFAR-10 image classification dataset according to an embodiment of the present invention; Figure 5 The convergence performance of the SCA-based receiver beamforming design algorithm and the overall alternation optimization algorithm in one embodiment of the present invention is shown. Figure 6 The convergence performance of the HG3S hybrid intelligent optimization algorithm in one embodiment of the present invention; Figure 7 This invention illustrates the relationship between total task inference latency and the number of devices under a CDI scheme with FA assistance and PFA assistance according to one embodiment of the present invention. Figure 8 This illustrates the relationship between total task inference latency and accuracy threshold in one embodiment of the present invention. Figure 9 This invention illustrates the relationship between total task inference latency and device computing power in one embodiment. Figure 10 This illustrates the relationship between total task inference latency and maximum computing power of Elasticsearch in one embodiment of the present invention. Figure 11 This is an embodiment of the relationship between total task inference latency and path loss index. Figure 12 This is an embodiment of the present invention showing the relationship between total task inference latency and the number of devices. Detailed Implementation

[0021] The present invention will now be described in further detail with reference to examples and accompanying drawings.

[0022] This invention discloses a latency optimization method for an edge-to-edge collaborative inference system based on FA assistance, the specific process of which is as follows: Figure 1 and 2 As shown. The example includes the following steps: S1: Establish a model of an FA-assisted end-edge collaborative reasoning system.

[0023] Orthogonal spectrum resources are used for transmission to prevent interference between devices, and the algorithm designed in this paper is used to coordinate FA positioning to further improve transmission efficiency.

[0024] In the S1 system model, each single-antenna device needs to use the DNN model to perform inference tasks. Given the limited computing power of the devices, the DNN model is divided into two parts. Specifically, the initial layer of the DNN is processed locally on each single-antenna wireless device, while the extracted IFD is offloaded to the ES of the base station BS, where it is further processed by the remaining layers of the DNN to complete edge inference. Due to the open nature of wireless channels, wireless transmission channels are susceptible to fading, noise, and interference, leading to increased IFD transmission delay and distortion, thus severely reducing the final inference accuracy. Therefore, by controlling the position of the FA, the impact of the wireless channel environment on the final inference accuracy is reduced.

[0025] Analysis of the computational intelligence system in the S1 system model reveals that its operation is divided into four stages: local computation at the device end, IFD wireless transmission, edge computation, and feedback of inference results. Since the amount of data in the inference results is typically small, and the feedback latency is negligible, a collaborative model based on DNN partitioning, an FA-assisted IFD transmission model, and a corresponding inference accuracy model are established to determine the total inference latency of the device.

[0026] like Figure 1 As shown, the system model described in S1 includes a base station BS equipped with an ES, wherein the BS is deployed with FA, A single-antenna wireless device is used. The single-antenna wireless device first performs local inference, then the BS uses a fluid antenna to transmit data with the wireless device, and the ES equipped in the BS performs edge inference. Simultaneously, to ensure the quality and transmission accuracy of the IFD, orthogonal spectrum resources are used for transmission during the above process to prevent inter-device interference, and the algorithm designed in this paper is used to coordinate FA positioning, further improving transmission efficiency.

[0027] S11. For the collaborative model based on DNN partitioning, for ease of representation, use... This refers to a single wireless antenna device. Considering that DNN models typically consist of multiple layers, including convolutional layers and fully connected layers, this invention uses... Index each layer. (The first...) The input of the layer depends on the first The layer's output. In the processing device. When performing inference tasks, this invention reduces the computational workload of the convolutional layer to , , in, and These are the height and width of the input feature map, respectively. and These represent the number of channels in the input feature map and the output feature map, respectively. This indicates the size of the convolution kernel. Specifically, and These are the height and width of the output feature map, respectively, which depend on the input feature map and the convolution kernel, and can be calculated: , ; The computational workload of a fully connected layer is readily apparent to be the product of the number of input and output channels, expressed as: ; also, Used to indicate equipment The DNN split points indicate the distance from layer 0 to layer 1. Layer by device Execution, and from the first layer to the first The layer is processed by ES. When This means the equipment The inference task is performed solely by ES, while Indicates device Local processing of its tasks, the first The IFD of the layer needs to be transferred to ES, and its data size can be expressed as: .

[0028] To represent data size , Introduction This represents the memory usage per unit of data. In this context, it can be represented as: ; When handling inference tasks, the device The total computational workload of ES and Espada is measured in floating-point operations. , Specifically, it is expressed as follows: ; Obtain the equipment After considering the total computational load, this invention further derives its local inference latency. and local energy consumption Let the calculated capacity of the equipment be... The number of floating-point operations per CPU cycle is The coefficient related to chip architecture is From the inference formula and the local computing energy consumption formula, we can know the local inference latency and local computing energy consumption: , ; Obtain the equipment Local inference delay Subsequently, because the BS receives the device in this system After the IFD (Inference Derivative) is executed, the BS performs the CDI (Critical Inference Difference) and returns the inference result to the device. Since the size of the inference result is typically small, the feedback latency is negligible. Therefore, only the edge inference latency needs to be calculated. ,remember This indicates that the BS is assigned to the device. The computational capacity of the inference task and the edge inference latency can be expressed as: ; Finally, considering the need to ensure that the BS has available computing resources to allow the system to operate normally, this invention limits the computing power of the BS: ; in, This indicates the maximum computing power of BS.

[0029] S12. Regarding the FA-assisted IFD transmission model, when the IFD is transmitted from the device to the BS, to avoid interference between devices, which would reduce IFD quality and ultimately inference accuracy, the wireless device transmits to the BS using orthogonal spectrum resources after completing local inference. Furthermore, the BS is equipped with an FA to improve channel gain and transmission efficiency through coordinated antenna positioning. This invention is described in [reference to a specific invention]. Let the position of the nth FA be such that it can be within the length Dynamic positioning within a linear space. Based on this, the steering vector of the FA array can be expressed as: ; in, Indicates the wavelength of the signal. Define the signal angle of arrival and perform a transpose operation. After obtaining the steering vector of the FA array, this invention utilizes a narrowband channel model suitable for a far-field uniform linear array layout. This model includes two parts: large-scale path loss and small-scale array response, as detailed below: in, This is the channel gain at a reference distance of 1 meter. Indicates device Distance between BS, This represents the path loss exponent. Wherein... This indicates large-scale path loss. This represents the response of a small-scale array. Therefore, the transmission rate of the IFD can be expressed as: ; in, and These are transmission bandwidth and power, respectively. Represents the received beamforming vector. Let the noise power be represented. Then, the formulas for IFD wireless transmission delay and IFD wireless transmission power consumption can be derived as follows: , ; In this IFD transmission model, the present invention establishes a channel model and improves IFD reception quality and reduces the impact of wireless impairments on inference accuracy by jointly optimizing the device transmit power, receive beamforming vector and FA position.

[0030] The multiple FA position vectors of the BS can be dynamically adjusted, but their minimum spacing must meet physical constraints. In addition, the channel model is constructed based on the FA's steering vector, and the receiver uses beamforming vectors. To enhance signal strength. IFD transmission rate Depends on the device's transmit power Beamforming vector Antenna location and signal-to-noise ratio.

[0031] S13. Regarding the collaborative inference accuracy model, traditional CDI schemes often focus on minimizing latency or energy, but wireless channel noise can severely distort the transmitted inference function derivation (IFD), and the robustness of the IFD extracted from different DNN partitions to noise varies. If accuracy is not explicitly modeled, optimization may sacrifice inference accuracy while reducing latency, causing the system to fail in real-world, harsh channels.

[0032] To address this issue, a regression fitting technique is employed to construct a closed-form inference accuracy function. The specific method is as follows: Figure 3 and 4 Taking the classic DNN model VGG16 on the CIFAR-10 image classification dataset as an example, a certain DNN partition point is fixed. Different signal-to-noise ratios (SNRs) were simulated by artificially adding Gaussian noise of varying intensities at the receiver, and the corresponding inference accuracy was measured. Then, a logistic function was used to fit curves to these discrete experimental points, resulting in a continuous and differentiable closed-form expression. Therefore, the relationship between inference accuracy, SNR, and DNN partition points can be expressed as follows: ; Among them, equipment Receiver signal-to-noise ratio It can be represented by the following expression: ; in, , , , It depends on the partition point The DNN structure and the fitting parameters of the dataset, among which To improve the accuracy of reasoning, The slope of the curve reflects the sensitivity of the partition point's characteristics to changes in SNR. Early partition points typically... A larger value and a steeper curve indicate greater sensitivity to noise. This is the inflection point, i.e., the SNR threshold at which inference accuracy begins to rise rapidly. Early partition points typically require a higher SNR to achieve high accuracy. Baseline accuracy, which is the lowest accuracy when SNR is very low.

[0033] In addition, to ensure the effectiveness of collaborative reasoning, constraints are placed on the accuracy of the system's reasoning: ; in, The inference accuracy threshold specified by the user.

[0034] S2: Under the system model described above, the DNN partition, the transmit power of each device, the allocation of computing resources, the receive beamforming vector at BS and the antenna positioning vector at FA are jointly optimized to minimize the total inference delay while ensuring inference accuracy and energy consumption constraints. S21, The optimization problem of minimizing the total inference latency of the system in S2 is expressed as: ; in, Represents DNN split points ,gather Represents the device's transmit power ,gather ES is assigned to the processing device Computational capabilities for inference tasks , Represents the received beamforming vector, set Antenna positioning vector , The constraints include: Inference accuracy constraints: Each single-antenna wireless device is required The final collaborative inference accuracy must be no less than the user-specified threshold. This ensures that the system can achieve low latency without sacrificing inference quality.

[0035] Equipment energy consumption constraints: , This represents the maximum available energy for each device; this constraint requires each wireless device to... The total energy consumption, namely the sum of the energy consumption of local DNN computing and the energy consumption of IFD wireless transmission, must not exceed the upper limit of its battery capacity. This constraint reflects the energy-constrained characteristics of actual mobile devices, thereby preventing the device from overheating or rapidly depleting the battery due to excessive power consumption caused by optimization.

[0036] ES compute resource constraints: , requiring all The sum of the resource allocations for the tail model calculations of each device cannot exceed the maximum available resources of the server. This aligns with the limited total computing power of the ES within the BS in practical work.

[0037] BS receive power constraints: , This represents the maximum received energy of the BS, and this constraint requires the received beamforming vector to... The total power must not exceed the upper limit. This constraint is a saturation limit for the power amplifier in a practical radio frequency link.

[0038] FA minimum interval constraint: , This represents the minimum distance between two adjacent FAs. It requires that the positions of two adjacent FAs must maintain a minimum distance of [missing information]. The physical spacing is adjusted to mitigate the impact of mutual coupling between adjacent antenna arrays on antenna performance, ensuring the optimized antenna performance. This can be implemented in actual hardware.

[0039] DNN partition point discrete constraints: Requires partition points It must be an integer, and can only range from 0 to the total number of floors. The values ​​between, make the partition point This aligns with the discrete nature of DNN partitioning.

[0040] Equipment transmit power upper and lower bound constraints: , This represents the maximum available energy of each device, and the constraint requires that the transmit power of each device must be non-negative and not exceed [a certain value]. This meets the constraints of physical feasibility.

[0041] Non-negativity constraints for edge computing resources: The requirement that the edge computing frequency allocated to each device must be non-negative is also a constraint that meets physical feasibility requirements.

[0042] FA position boundary constraints: The requirement is that the position of the first FA cannot be less than the starting point of the track, and the position of the last FA cannot exceed the ending point of the track. This ensures that all antenna positions are within a movable linear space, which is the feasible domain boundary for FA positioning optimization.

[0043] S3: Optimize the algorithm design by decomposing the optimization problem into several subproblems and using the block coordinate descent method to optimize them alternately until the subproblems converge and the solution is completed.

[0044] S3 includes the following steps: S31. For the total time delay problem of a mixed-integer nonlinear programming system, an efficient iterative algorithm is designed. The algorithm decomposes the total problem into four sub-problems: computational resource allocation, transmit power optimization, receive beamforming design, DNN partitioning, and FA localization. The AO algorithm is then used to iteratively solve these sub-problems until convergence.

[0045] S32. Solve the computational resource allocation subproblem. Given the transmit power of each device, the receive beamforming vector at the BS, the DNN partitioning, and the FA localization, the computational resource allocation subproblem can be expressed as: ; The constraints include: In the context of ES computational resource constraints and FA location boundary constraints.

[0046] Since this subproblem is a convex optimization problem, and the constraints are also convex functions, the Lagrangian function can be directly constructed and derived using the KKT conditions. The optimal allocation of computing resources. First, the optimal allocation of computing resources in the process. The Lagrange function is expressed as follows: ; The present invention will Defined as a Lagrange multiplier, and let .

[0047] Among the four conditions in KKT, the key is to use the stationary point condition and the complementary relaxation condition, which yields the following two equations: ; By performing equation transformations on the stationary point condition equation, we can obtain the Lagrange multipliers. It can be represented as: ; Furthermore, to satisfy the complementary relaxation condition and the condition that the Lagrange multipliers themselves are nonnegative, the following conditions can be obtained: ; Solving the simultaneous equations is easy to see: ; Finally, the Lagrange multipliers were obtained. as well as The sum of resource allocation calculations for the tail model of each device The optimal solution is: ; S33. Solving the transmit power optimization subproblem: Given the allocation of computational resources, the receiver beamforming vector of the BS, the DNN partitioning, and the FA positioning, the transmit power optimization subproblem can be expressed as: ; The constraints include: The constraints include inference accuracy constraints, equipment energy consumption constraints, and upper and lower bound constraints on equipment transmit power.

[0048] Due to the transmit power optimization subproblem Since the problem is non-convex, we introduce an auxiliary variable to transform it into a fully convex problem. To replace variable devices Transfer rate to the server's IFD Thus eliminating the rate This results in nonconvexity. Because... yes The concave function, its lower level set As a convex set, therefore, the present invention addresses... Apply convex constraints to the auxiliary variables. Restrictions will be imposed: ; Although the objective function achieves convexity after the above substitution, the equipment energy budget constraint and inference accuracy constraint require further processing to solve the problem. This invention first processes the second term of the equipment energy budget constraint through an equivalent transformation: ; Based on this, it is equivalently rewritten as: ; Subsequently, due to the problem The inference accuracy constraint requires uplink speed To achieve a certain level of accuracy Division The relevant threshold value, which can be expressed as This needs to be rewritten to transform the constraint into a convex form as follows: ; After introducing auxiliary variables, the non-convex transmit power optimization subproblem Transform into a convex optimization problem : ; The constraints include: constraints with the introduction of auxiliary variables, constraints on the device energy budget and inference accuracy after equivalent convex transformation, and constraints on the upper and lower bounds of the device's transmit power.

[0049] Due to the optimized transmit power optimization subproblem It is a standard convex optimization problem, so it can be solved using a standard convex solver (such as CVX).

[0050] S34. For the receiver beamforming design subproblem, given the allocation of computational resources, optimization of transmit power, and DNN partitioning and FA localization, the receiver beamforming design subproblem can be expressed as: ; The constraints include: The inference accuracy constraint and the BS receive power constraint are included.

[0051] Due to the sub-problem of receiver beamforming design It is also non-convex, therefore it can be used in A similar method used in [the text] introduces auxiliary variables. Replacement equipment achievable speed to the server and rewrite constraints ; However, introducing auxiliary variables Afterwards, the above constraints are still non-convex. To transform them into a tractable convex form, another set of variables is introduced. Therefore, the lower horizontal convex set Reconstructed as: ; This key constraint lies in two auxiliary variables. and After introduction, it becomes the required convex constraint; however, due to the introduction of the second set of auxiliary variables... There is also a non-convex constraint. Unresolved.

[0052] To effectively transform it into a solvable problem, it can be obtained by using a first-order Taylor expansion. The lower bound that can be handled: ; in It is a feasible Taylor expansion point. Based on this Taylor expansion point, for non-convex constraints... Perform the corresponding variations: ; For the sub-problem of receiver beamforming design The inference accuracy constraint is also transformed into a more manageable convex form: ; After introducing auxiliary variables, the non-convex receiver beamforming design subproblem Transform into a convex optimization problem : ; The constraints include: introducing auxiliary variables. and Constraints, inference accuracy constraints after equivalent convex transformation, and BS receive power constraints.

[0053] Due to the problem of optimized receiver beamforming This is a standard convex optimization problem, therefore it can be solved efficiently using standard convex solvers (such as CVX). However, the original problem... Approximate to There is an unavoidable performance gap between them. To bridge this gap, the SCA technique is adopted. In this algorithm, the Taylor expansion parameters... Use from No. The optimal solution obtained from the next iteration is iteratively refined.

[0054] S35. For the DNN partitioning and FA localization subproblem, given the computational resource allocation, transmit power optimization, and the BS receive beamforming vector, the DNN partitioning and FA localization subproblem can be expressed as: ; The constraints include: The constraints include inference accuracy constraints, device energy budget constraints, FA minimum spacing constraints, DNN partition point discrete constraints, and edge computing resource non-negativity constraints.

[0055] Because the above problem is a high-dimensional non-convex optimization problem, the continuous FA location variables and the discrete DNN partitioning variables are prone to coupling in this subproblem. This coupling affects the total inference delay, thus making the solution difficult.

[0056] To address the challenges of handling total inference latency and constraints caused by the strong coupling between continuous FA location variables and discrete DNN partitioning variables in this high-dimensional non-convex optimization problem, this paper innovatively proposes a hybrid gravity-salp swarm search intelligent optimization algorithm. The core design concept of this algorithm lies in organically integrating the global exploration capability of the Gravitational Search Algorithm (GSA) with the local development capability of the Salp Swarm Algorithm (SSA), achieving synergistic effects between global optimization and refined local search.

[0057] Among them, GSA, a metaheuristic optimization algorithm based on Newton's law of universal gravitation, abstracts each feasible solution in the search space as a particle with mass and drives the search process by simulating the gravitational interaction between particles. In this mechanism, high-quality solutions with larger masses exert a stronger gravitational attraction on surrounding particles, thereby guiding the population to converge toward the globally optimal region, effectively avoiding the algorithm from getting trapped in local optima and achieving robust global exploration capabilities. In contrast, SSA is a swarm intelligence algorithm that simulates the chain-like swimming and foraging behavior of swarms. Its core lies in the use of a leader-follower hierarchical collaborative mechanism: the leader is located at the front of the tunic swarm chain and directly updates its own position based on the current optimal solution to explore better areas; the followers are arranged sequentially at the back of the chain and update their positions by tracking the movement trajectory of the individuals in front. This chain-like information transmission structure can effectively maintain population diversity and enhance the ability to conduct fine-grained local searches of the neighborhood of the optimal solution.

[0058] This method deeply integrates the gravity interaction mechanism of GSA for global optimal solution exploration with the leader-follower chain structure of SSA for local refinement of the optimal solution neighborhood. The HG3S algorithm achieves an effective balance between global exploration and local refinement, demonstrating its ability to solve complex constrained optimization problems. The paper explores the significant potential of this problem. To efficiently handle complex equality and inequality constraints, this paper first employs the Penalty Function Method to reconstruct the original constrained optimization problem into a standard unconstrained optimization form. This involves introducing a penalty term into the objective function to quantify the degree of constraint violation, thus significantly penalizing infeasible solutions and leading to their natural elimination. Subsequently, the HG3S algorithm is applied to solve this unconstrained problem. Using this algorithm, multi-objective constrained optimization problems can be transformed into single-objective minimization problems, significantly reducing the complexity of algorithm design while ensuring constraint satisfaction and solution quality.

[0059] To efficiently handle constraints, this invention first uses a penalty function method to... The problem is reconstructed into a standard unconstrained optimization problem, and then the proposed HG3S algorithm is applied. This method is applicable to evolutionary algorithms because it transforms constrained optimization into single-objective minimization by penalizing infeasible solutions. Based on this, the subproblems are further... Rewritten in the following standardized form: ; The constraints include: The constraints include inference accuracy constraints, device energy budget constraints, FA minimum interval constraints, and feasible region. Constraints.

[0060] in, , This represents the set of optimization variables that satisfy the DNN partition point constraints and the FA position boundary constraints. and sub-problems The constraints in the equation are transformed as follows: , ; This makes the problem easier to represent. The penalty function for this problem is then defined as... , and Therefore, the corresponding penalty problem can be derived as follows: ; in, , and It is a penalty factor; for a sufficiently large penalty factor, The solution can converge. In the proposed HG3S algorithm, this invention will... The solution is called the particle's position. Furthermore, the particle's velocity represents the direction of evolution in the search for a better solution. The position of the i-th particle in the t-th iteration is defined as: , in, , It is a particle population. It is the total number of particles in the population. Each particle will... The former 3D encoding DNN partition points, then The dimension is encoded as the FA position. Furthermore, this invention uses... It serves as a fitness function to evaluate the quality of each particle's position.

[0061] For the HG3S algorithm, the algorithm first enters the GSA phase to enhance global exploration. In this phase, particles with better fitness values ​​are assigned greater mass, while lighter particles are attracted to heavier particles with superior fitness through gravity. Gravity generates acceleration, modifying the velocity and position of each particle. In this way, the population can get closer to higher-quality positions that generate better solutions. Therefore, the mass of each particle is defined as: , ; in, It is the first The optimal fitness of the next iteration can be determined by... Calculation. Furthermore, It is the first The worst fitness of the next iteration can be expressed as Therefore, the law of universal gravitation ensures that heavier particles exert a stronger gravitational force on the motion of other particles. From particles The effects of gravity are: ; in It is a random constant uniformly distributed within [0,1]. Therefore, through force analysis, it can be seen that the resultant force causes the particle to... The acceleration is: ; Obtaining particles After the acceleration, let and They represent the time as follows Particles at time The velocity and position can be expressed as follows: ; in This represents a random constant uniformly distributed within the range [0,1]. Because this invention discretizes the DNN partition points, i.e. Therefore, the position of each particle The former The dimension needs to be discretized into , This indicates the integer division operation.

[0062] However, although GSA provides robust global exploration by exploring the gravity between heavier and lighter particles, it has the limitation of slow convergence. Therefore, this invention uses an integrated SSA algorithm to optimize it.

[0063] For the ensemble SSA algorithm, which possesses local refinement properties that enable it to obtain better solutions and accelerate convergence, the specific method is as follows: The particles in the middle are divided into two parts, of which the first part is the first part. One particle is considered the leader particle, and the remaining particles are considered follower particles. The leader particle can be considered in its optimal position. By updating its position and performing local refinement, the following formula is obtained: ; in Indicates the new position of the leader particle. yes The position of the particle, and They are The upper and lower bounds, and It is a random constant within [0,1]. It is the time-varying decay coefficient of the SSA leader particle update step size, causing it to decay exponentially. , Furthermore, by utilizing a greedy selection mechanism to obtain a better position, the new position... An update is only performed when the fitness of the element is lower than its current position. Based on this, the above formula is modified as follows: ; Therefore, after the above iterative steps, the process continues until the convergence criterion is met. By synergistically combining the global exploration of GSA with the local refinement of SSA, HG3S can efficiently explore... The solution is found in the constrained mixed variable search space, thus solving the DNN partitioning and FA localization sub-problems.

[0064] S36. After addressing the above four sub-problems, iteratively solve the problem using an alternating optimization algorithm until convergence to a suboptimal solution. The final output is the optimized DNN model with the variable for selecting the split points. The transmit power of each device Computational resource allocation Received beamforming vector at BS and the antenna positioning vector of FA This enables latency optimization of the end-edge collaborative inference system based on FA assistance.

[0065] like Figure 5-12 This fully verifies that the present invention can reduce the total inference latency by up to 32%, and that the performance gain is far superior to the traditional baseline scheme in key scenarios such as high accuracy threshold, poor channel, and large-scale device access. It provides an efficient and reliable solution for resource-constrained wireless devices to achieve low latency and high accuracy end-edge collaborative inference.

[0066] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0067] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0068] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0069] The above description is merely an embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.

Claims

1. A delay optimization method for an end-to-side cooperative inference system based on fluid antenna assistance, characterized in that, include: S1. Establish a fluid antenna FA-assisted end-edge collaborative inference system model. The system model includes a base station BS equipped with an edge server ES and several single-antenna wireless devices. The BS is deployed with several FAs (Flying Array Components) that can dynamically adjust their positions within a predetermined linear space, and single-antenna wireless devices. Intermediate feature data (IFD) is transmitted to the BS using orthogonal spectrum resources; based on the system model, a DNN-based collaborative model, an FA-assisted IFD transmission model, and an inference accuracy model are constructed to determine the total inference delay of the system, which includes local inference delay, IFD wireless transmission delay, and edge inference delay; S2. Under the system model, construct a joint optimization problem with the goal of minimizing the total inference latency of the system. Jointly optimize the DNN partitioning point, the transmit power of each device, the allocation of computing resources, the receive beamforming vector at the BS and the antenna positioning vector at the FA. The joint optimization problem must satisfy the constraints of inference accuracy and device energy consumption. The device energy consumption includes local energy consumption and IFD wireless transmission energy consumption. S3. Decompose the joint optimization problem into several sub-problems, and use the block coordinate descent method to alternately optimize each sub-problem until convergence, thereby obtaining a suboptimal solution and completing the latency optimization of the end-edge collaborative reasoning system.

2. The method according to claim 1, characterized in that, S1 constructs a collaborative model based on DNN partitioning, specifically including: The DNN segmentation method is adopted to divide the DNN model into a head model and a tail model based on the DNN segmentation points. The single-antenna wireless device runs the DNN head model and extracts the IFD, while the ES runs the DNN tail model to complete the inference. The segmentation points determine the local computing load, edge computing load, and IFD data volume. Based on the computing capabilities of the single-antenna wireless device and the ES, combined with the computing workload of each layer of the DNN, the local inference latency, local energy consumption, and edge inference latency are derived. At the same time, the maximum computing capability of the BS is limited to ensure that the BS has available computing resources.

3. The method according to claim 2, characterized in that, The computational workload of each layer in the DNN is as follows: : ; in, and They are the first The height and width of the input feature map of device k in layer k. and They represent the first The number of channels in the input and output feature maps of the layer. This indicates the size of the convolution kernel; the computational workload of a fully connected layer is the product of the number of channels in the input feature map and the output feature map.

4. The method according to claim 2, characterized in that, The FA-assisted IFD transmission model is constructed in S1, specifically including: remember For the position of the nth FA, construct the steering vector of the FA array, use the narrowband channel model under the far-field uniform linear array to obtain the channel between the device and the BS, combine the Shannon formula to obtain the transmission rate of the IFD, and derive the IFD wireless transmission delay and IFD wireless transmission power consumption based on the IFD data volume.

5. The method according to claim 4, characterized in that, The steering vector of the FA array is: ; In the formula, Indicates the wavelength of the signal. Define the signal angle of arrival and perform a transpose operation; The method of obtaining the channel between the device and the BS using the narrowband channel model under a far-field uniform linear array is as follows: ; In the formula, It is the channel gain at a reference distance of 1 meter between the device and the BS. Indicates device Distance between BS, This represents the path loss index.

6. The method according to claim 1, characterized in that, Step S1 involves constructing the inference accuracy model, specifically including: Based on experiments with the classic DNN model VGG16 on the CIFAR-10 image classification dataset, regression fitting techniques were used to obtain the relationship between inference accuracy and DNN split points, device... The closed-form expression for the receiver signal-to-noise ratio (SNR) is given, and an exponential function is used to characterize the correlation among the three factors. An inference accuracy threshold is set to establish inference accuracy constraints.

7. The method according to claim 4, characterized in that, In S3, the joint optimization problem is specifically decomposed into four sub-problems: computational resource allocation, transmit power optimization, receive beamforming design, DNN partitioning and FA localization. Solving the computational resource allocation subproblem, specifically including: fixed equipment Based on the transmit power, receive beamforming vector at BS, DNN partitioning, and FA positioning, a subproblem of computational resource allocation with the objective of minimizing the total inference delay of the system is constructed. This subproblem is a convex optimization problem with constraints including ES computational resource constraints and FA location boundary constraints. A Lagrangian function is constructed, and the closed-form optimal solution of the computational resources allocated by ES to each device is derived using the stationary point condition and complementary relaxation condition of KKT conditions. Solving the transmit power optimization subproblem involves: fixed computational resource allocation, receiver beamforming vector of the base station (BS), DNN partitioning, and FA localization. The transmit power optimization subproblem is constructed with constraints including: inference accuracy constraints, equipment power consumption constraints, and upper and lower bounds on the equipment transmit power. Auxiliary variables are introduced. Instead of the transmission rate, the non-convex constraints are transformed into convex forms through equivalent transformation and convexization, and the optimal solution for the transmit power is solved using convex optimization tools; Solving the receive beamforming design subproblem involves: fixed computational resource allocation, equipment transmit power, DNN partitioning, and FA localization, to construct the receive beamforming design subproblem; and introducing auxiliary variables. and The method employs a successive convex approximation technique, using a first-order Taylor expansion to approximate the non-convex terms, iteratively solving the convex approximation problem until convergence to obtain the optimal solution for receiving beamforming. Constraints include: introduced auxiliary variables. and Constraints, inference accuracy constraints after equivalent convex transformation, and BS receiver power constraints. Solving the DNN partitioning and FA positioning subproblem involves: establishing a fixed computational resource allocation, equipment transmit power, and BS receive beamforming vector; and constructing the DNN partitioning and FA positioning subproblem with constraints including: inference accuracy constraints, equipment power consumption constraints, FA minimum spacing constraints, DNN partition point discreteness constraints, and edge computational resource non-negativity constraints. A penalty function method is used to reconstruct the constrained DNN partitioning and FA positioning subproblem into a standard unconstrained optimization problem, quantifying the degree of constraint violation by introducing a penalty term into the objective function. The Hybrid Gravity-Turned Turtle Group Search Intelligent Optimization Algorithm (HG3S) is then used to process the standard unconstrained optimization problem, thereby obtaining the optimal solution for the DNN partitioning points and the antenna positioning vectors of the FA.

8. The method according to claim 7, characterized in that, The Hybrid Gravity-Turbot Swarm Search Intelligent Optimization Algorithm HG3S integrates the global exploration capability of the Gravity Search Algorithm GSA with the local development capability of the Turbot Swarm Algorithm SSA.

9. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program that, when executed, implements the method according to any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, It stores executable instructions that, when executed, implement the method of any one of claims 1-8.