Sensor edge end system deployment system and method based on large model network layer fusion compression

Through cross-layer heterogeneous fusion and depth-first search algorithms, and dynamically allocate the network layer with the equipment load formula, the problem of resource and accuracy balance in multimodal data processing is solved, and efficient model deployment and stability improvement is achieved.

CN120238450APending Publication Date: 2025-07-01WUHAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510391545.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

When processing multimodal data, model compression and distributed computing cannot effectively solve the problems of resource constraints, data complexity, accuracy and resource balance, resulting in a decrease in model accuracy and system reliability.

Method used

The internal network layer of the cloud model is heterogeneously fusion across layers through the network layer through the network layer fusion module, and the depth-first search algorithm is used to identify direct dependencies, combine device computing power and memory capacity to classify network layer groups, and dynamically allocate network layer to devices through load formulas to achieve accurate resource utilization and deployment.

Benefits of technology

It effectively reduces the amount of model parameters, maintains the ability to express key features, improves hardware resource utilization, reduces task queuing delay, is suitable for environments with fluctuations in computing resources, and improves the stability and efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238450A_ABST
    Figure CN120238450A_ABST
Patent Text Reader

Abstract

The invention relates to a sensor side end system deployment system based on large model network layer fusion compression, and the system comprises a network layer fusion module which carries out the fusion compression of an internal layer of a cloud large model; the network layer characteristic classification module traverses a data flow path through a DFS algorithm, clusters network layers with a direct dependency relationship into groups based on a preset dependency layer number threshold value, and performs type classification through a characteristic value threshold value in combination with the calculation amount, the memory access frequency, the equipment calculation capability, the memory capacity and the calculation characteristic value of each group; and the large model deployment module is used for calculating an equipment load value by adopting a preset load formula, and dynamically distributing the network layer to the corresponding equipment through a distribution model according to the load state and the network layer group type to finish deployment. According to the method, the large model parameter quantity is effectively reduced through the network layer fusion technology, more key feature expression capability is reserved, and the hardware resource utilization rate is improved through the method of carrying out network layer classification dynamic deployment based on the network layer characteristic value and the equipment load.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning model deployment, and particularly relates to a sensor edge system deployment and method based on large model network layer fusion compression. Background Art

[0002] In today's technology field, deep learning model deployment technology has become a key force driving the intelligent development of various industries. Among them, the rise of cloud data centers is an important development milestone in this field. Their powerful computing capabilities and rich storage resources enable the training of large-scale deep learning models, greatly improving the performance and application scope of the models. At the same time, with the continuous progress of sensor technology, integrated sensors are widely used in many fields such as intelligent transportation, smart home, industrial automation, and healthcare, and can collect multi-modal data, providing a richer information source for deep learning models, which has also become one of the key driving forces for the development of model deployment technology.

[0003] First, the model compression technology has poor adaptability to multi-modal data. The model compression technology aims to reduce the model's computational amount and memory occupancy. Common parameter pruning and quantization methods have serious defects when dealing with multi-modal data. Parameter pruning removes parameters considered unimportant based on certain rules, and quantization represents model parameters with low-precision data. When the existing technology uses model compression technology to compress the model to process multi-modal data, due to the lack of full consideration of the complexity and diversity of multi-modal data, key parameters for specific modal feature extraction may be mistakenly deleted during the parameter pruning process, and large errors are also easily introduced during the quantization process due to the large differences in the numerical ranges and distributions of different modal data. These errors accumulate and amplify during model calculations, seriously affecting the model's accuracy.

[0004] Second, the distributed computing technology faces difficulties in multi-modal data processing. The distributed computing model relies on communication protocols and task scheduling algorithms to allocate computing tasks to multiple devices for collaborative completion. However, the transmission of multi-modal data between different devices has extremely high requirements for network bandwidth and stability. When the existing technology uses distributed computing for model applications, due to the lack of an effective response mechanism for the transmission characteristics of multi-modal data, data transmission delays or even losses often occur in actual applications, making the model unable to obtain complete multi-modal information in time for calculation, seriously affecting the running efficiency and performance of the model. Moreover, the distributed computing architecture is complex, and faults are likely to occur in links such as device-to-device communication protocol design, task allocation, and coordination, reducing the reliability of the system.

[0005] Finally, it is difficult for the prior art to balance model accuracy and resource consumption. When deploying a deep learning model to edge devices and edge servers, it is necessary to take into account both model accuracy and device resource limitations. However, current model compression and distributed computing technologies cannot effectively solve this contradiction. When applying the model in the prior art, in order to adapt to the resource limitations of edge devices, model compression and distributed computing are adopted. However, in this process, the model accuracy is overly sacrificed, and some complex faults cannot be accurately identified, which cannot meet the requirements of practical applications.

[0006] In summary, when the existing model deployment technology processes multi-modal data collected by integrated sensors and deploys it to edge devices and edge servers, it cannot effectively solve problems such as resource constraints, complex data, accuracy-resource balance, and limitations of existing deployment schemes. Summary of the Invention

[0007] The object of the present invention is to address the deficiencies of the prior art and provide a sensor edge system deployment system based on cross-layer heterogeneous fusion compression of large model network layers, including: a network layer fusion module for performing cross-layer heterogeneous fusion on each internal network layer of a large model deployed in the cloud to obtain a fused large model;

[0008] a network layer characteristic classification module for traversing the number of network layers with direct dependency relationships on the data flow path of the internal network layers of the fused large model through a depth-first search algorithm, classifying the network layers with direct dependency relationships into the same network layer group through a preset layer threshold, where the direct dependency relationship is a unidirectional or bidirectional data transfer relationship formed when the output of a certain network layer directly serves as the input of another network layer or the input of a certain network layer directly originates from the output of another network layer; calculating the characteristic values of each network layer group according to the computational amount, memory access frequency of each network layer group, as well as the computational power, memory capacity of the device, and real-time and accuracy requirements of the application scenario, and classifying the types of each network layer group by combining the characteristic values of each network layer group with a preset characteristic value threshold;

[0009] a large model deployment module for calculating the load values of each device through a load formula, and allocating and deploying each network layer group to the corresponding device according to the load values of each device and the types of each network layer group using a preset network layer allocation model.

[0010] Further, in the network layer fusion module, the specific method for fusing the internal network layers of the large model deployed in the cloud is as follows:

[0011] The network layer pairs for fusion include a convolutional layer and a bias layer, a convolutional layer and a pooling layer, a fully connected layer and a convolutional layer, an activation function and a regularization layer, a convolutional layer and a recurrent neural network, an attention mechanism and a convolutional layer, an attention mechanism and a fully connected layer, a graph neural network and a convolutional layer;

[0012] The method of fusing the convolutional layer and the bias layer is to integrate the bias values into the weight matrix of the convolutional kernel through a preset mapping rule;

[0013] The method of fusing the convolutional layer and the pooling layer is that during the process of the convolutional kernel sliding along the input feature map, for the elements in each sliding window, first perform convolutional weighted summation, and then perform a pooling operation on the convolutional output within the same window;

[0014] The method of fusing the fully connected layer and the convolutional layer is to replace the fully connected layer with global average pooling;

[0015] The method of fusing the activation function and the regularization layer is to preprocess the output value of the fully connected layer according to the deterministic rule of the regularization layer in the inference stage, and directly input the preprocessed result into the activation function to complete the non-linear transformation;

[0016] The method of fusing the convolutional layer and the recurrent neural network is to first extract the image spatial features of the convolutional layer as the input for the recurrent neural network fusion, and then introduce a convolutional operation at specific steps inside the recurrent neural network fusion;

[0017] The method of fusing the attention mechanism and the convolutional layer is that when calculating the feature map in the convolutional layer, by introducing an attention weight matrix, dynamically adjust the element weights of the local feature regions corresponding to each convolutional kernel;

[0018] The method of fusing the attention mechanism and the fully connected layer is to multiply the attention weight matrix on the basis of the neuron connection weights of the fully connected layer;

[0019] The method of fusing the graph neural network and the convolutional layer is to introduce a convolutional operation mechanism in the graph convolutional network. For graph-structured data, use the adjacency matrix to define the neighborhood relationship between nodes, and aggregate the feature information of adjacent nodes to complete local feature extraction.

[0020] Furthermore, in the network layer feature classification module, the specific method of traversing the number of network layers with direct dependency relationships through the depth-first search algorithm on the data flow path between the internal network layers of the large model after network layer fusion, and classifying the network layers with direct dependency relationships into the same network layer group through a preset dependency layer threshold is as follows:

[0021] For the network layer L1 as the input layer, its dependency relationship set is an empty set;

[0022] For the network layer L i , i > 1, if its input data is sourced from the network layer L j , j < i, then add L j to the network layer L i, the set of dependencies D for i > 1 i , network layer L i and its set of dependencies D i the network layer L in j has a direct dependency relationship; if the output data of network layer L i , i > 1 becomes the input of network layer L k , k > i, then in network layer L k , the set of dependencies D for k > i k adds network layer L i , network layer L k and its set of dependencies D k the network layer L in i has a direct dependency relationship;

[0023] Based on the set of dependencies of each network layer, a data flow path is formed. Using the depth - first search algorithm, the number of transfers n of specific data between the corresponding network layers on the data flow path is traversed. The formula is expressed as:

[0024] D i = D i ∪{L j}

[0025] D k = D k ∪{L i}

[0026] D i ={L j |R i,j = 1, j ≠ i}

[0027]

[0028] where R i,j is whether there is a direct dependency relationship between network layer i and network layer j. If there is, then R i,j = 1, otherwise R i,j = 0; D i is the set of dependencies of network layer L i , n represents the number of network layers with direct dependency relationships, and N is the total number of network layers;

[0029] Set a dependency layer threshold n min , if n ≥ n min , then classify each network layer passed by the specific data during transmission into the same network layer group.

[0030] Further, in the network layer feature classification module, the specific method for calculating the feature values of each network layer group according to the computational amount, memory access frequency of each network layer group, and the computational ability, memory capacity of the device, and the real-time and accuracy requirements of the application scenario, and classifying the types of each network layer group through a preset feature value threshold is as follows:

[0031]

[0032] Among them, P is the feature value of network layer p of the network layer group, C p is the computational amount of network layer p of the network layer group, is the weight of the computational amount of network layer p of the network layer group, M p is the memory access frequency of network layer p of the network layer group, is the weight of the memory access frequency of network layer p of the network layer group, and the weight of the computational amount of network layer p of the network layer group and the weight of the memory access frequency are determined according to the computational ability, memory capacity of the device, and the real-time and accuracy requirements of the application scenario;

[0033]

[0034] Among them, O p is the computational amount of network layer p in network layer group p ′ and l is the total number of network layers in network layer group p;

[0035]

[0036] Among them, f c is the number of times that network layer group p accesses the device memory during the c-th run, and m is the number of simulated runs of the cloud-deployed model;

[0037] Set the first feature value threshold T1 and the second feature value threshold T2, T1 < T2. When P p < T1, network layer group p is determined to be memory access intensive; when P p > T2, network layer group p is determined to be computationally intensive; when T1 ≤ P p ≤ T2, network layer group p is determined to be hybrid.

[0038] Further, in the network layer dynamic allocation module, the specific method for calculating the load value of each device through a preset load formula and allocating each network layer group to the corresponding device according to the load value of each device and the type of each network layer group using a preset network layer allocation model is as follows:

[0039] S = α×U cpu + β×U mem + γ×Unet +δ×L

[0040] Among them, S is the device load value, U cpu is the CPU usage rate of the device, U mem is the memory occupancy rate of the device, U net is the network bandwidth utilization rate of the device, L is the task queue length of the device, α is the weight of the CPU usage rate of the device, β is the weight of the memory occupancy rate of the device, γ is the weight of the network bandwidth utilization rate of the device, δ is the weight of the task queue length of the device, and α + β + γ + δ = 1;

[0041] The device includes an edge device and an edge server. First, deploy the memory access-intensive network layer group on the edge device. The deployment principle is to sequentially select, from the set of edge devices where the available memory amount is greater than the average available memory amount of all edge devices, the edge device with the current minimum memory usage rate and less than the memory usage rate threshold of the memory access-intensive network layer group for deployment;

[0042] After completing the deployment of the memory access-intensive network layer group, then deploy the compute-intensive network layer group on the edge server. The deployment principle is to sequentially select, from the set of edge servers where the computing power is greater than the average computing power of all edge servers, the edge server with the current minimum load and less than the load threshold of the compute-intensive network layer group for deployment;

[0043] For the hybrid network layer group, screen out the device set D that meets the computing power threshold and available memory amount threshold of the hybrid network layer group mix , the device set D mix includes edge devices and edge servers. In the device set D that meets the computing power threshold and available memory amount threshold of the hybrid network layer group mix select the device with the lowest load value for the deployment of the hybrid network layer group.

[0044] Furthermore, if during the deployment process of the memory access-intensive network layer group, in the set of edge devices where the available memory resource amount is greater than the average available memory resource amount of all edge devices, the situation occurs that the memory usage rate of all current edge devices is greater than the memory usage rate threshold of the memory access-intensive network layer group, then deploy the remaining undeployed memory access-intensive network layer groups to the edge server. The deployment principle is to sequentially select, from the set of edge servers where the available memory amount is greater than the average available memory amount of all edge servers, the edge server with the current minimum load value and the current memory usage rate less than the memory usage rate threshold of the memory access-intensive network layer group for deployment;

[0045] If, during the deployment of a compute-intensive network layer group, in the set of edge servers whose computing power is greater than the average computing power of all edge servers, it is found that the load values of all current edge servers are greater than the load value threshold of the compute-intensive network layer group, then the remaining undeployed compute-intensive network layer groups are deployed to the end-side devices. The deployment principle is to sequentially select, from the set of end-side devices whose computing power is greater than the average available memory of all end-side devices, the end-side device with the lowest current memory usage and a current load value less than the load value threshold of the compute-intensive network layer group for deployment.

[0046] Furthermore, a network layer migration module is used to unload network layer data from the devices where the network layer group has been deployed, and it uses an adaptive network transmission protocol for transmission based on the device network connection status and bandwidth. The unloaded network layer data is transmitted to the device determined to receive the network layer group migration using data transmission methods such as multi-channel transmission and / or fragmentation transmission and / or resume from breakpoint. The device that has received the network layer group migration selects an appropriate decompression algorithm for the received network layer data according to its own computing power and storage resource status.

[0047] Furthermore, in the network layer migration module, the specific method of using an adaptive network transmission protocol to unload the network layer on the deployed devices according to the device network connection status and bandwidth is as follows:

[0048] For a device that is determined to receive the network layer group but has an unstable network connection, before starting the unloading process of the network layer group on the corresponding deployed device, the network connection status of the device that is determined to receive the network layer group but has an unstable network connection is evaluated: When the signal strength Q signal is higher than the preset signal strength threshold Q signal_threshold , and the packet loss rate Q loss is lower than the preset packet loss rate threshold Q loss_threshold , and the latency Q delay is less than the preset latency threshold Q delay_threshold , the network layer group is unloaded;

[0049] For a device that has deployed the network layer group but has insufficient bandwidth, before transmitting the network layer data to the device determined to receive the network layer group, measure the bandwidth upper limit B max and the real-time available bandwidth B avail of the device that has deployed the network layer group but has insufficient bandwidth. If the bandwidth upper limit B max meets the transmission requirements of the network layer group data to be unloaded, then transmit directly; if the bandwidth upper limit B maxIf the transmission requirements of the network layer group data to be unloaded are not met, layered compression processing is performed on the network layer group data to be unloaded. Assume that the total number of network layers in the network layer group is l, and the compression ratio of the tth network layer is r. t , and make the original data s of the tth network layer t satisfy

[0050] A sensor edge system deployment method based on large model network layer fusion compression includes the following steps:

[0051] Perform cross-layer heterogeneous fusion of each internal network layer of the large model deployed in the cloud to obtain the fused large model;

[0052] On the data flow path of the internal network layer of the fused large model, the number of network layers with direct dependencies is traversed by a depth-first search algorithm, and the network layers with direct dependencies are classified into the same network layer group by a preset layer number threshold, wherein the direct dependency is a one-way or two-way data transmission relationship formed by the output of a certain network layer directly serving as the input of another network layer or the input of a certain network layer directly derives from the output of another network layer; the characteristic value of each network layer group is calculated according to the computing amount of each network layer group, the memory access frequency, the computing power, memory capacity of the device, and the real-time and accuracy requirements of the application scenario, and the type of each network layer group is classified by using the characteristic value of each network layer group combined with the preset characteristic value threshold;

[0053] The load value of each device is calculated using the load formula, and each network layer group is allocated and deployed to the corresponding device using a preset network layer allocation model based on the load value of each device and the type of each network layer group.

[0054] A computer program product includes a computer program / instruction, which, when executed by a processor, implements the above-mentioned sensor edge system deployment method based on large model network layer fusion compression.

[0055] The beneficial effects of the present invention are:

[0056] 1. Through the network layer fusion technology, the number of large model parameters is effectively reduced, and compared with the traditional pruning method, more key feature expression capabilities can be retained. The dependency graph constructed by the depth-first search algorithm (DFS) can accurately identify the topological structure between network layers and avoid the cliff-like drop in model performance caused by incorrect pruning.

[0057] 2. The dual - dimension classification mechanism based on the characteristics of the network layer (computing volume, memory access frequency) and device capabilities (computing power, memory) enables devices with rich memory resources to deploy memory - access - intensive network layers, high - computing - power devices such as GPUs to deploy computing - intensive network layers, and devices with rich memory and computing resources to process hybrid network layers, thus improving the utilization rate of hardware resources.

[0058] 3. The dynamic allocation strategy calculated through the load formula (considering parameters such as real - time CPU / GPU utilization rate, memory occupancy rate, etc.) can reduce the task queuing delay in multi - device scenarios compared with the static allocation scheme, and is especially suitable for environments with fluctuating computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 It is a block diagram of the system of the present invention.

[0060] Figure 2 It is a schematic flow diagram of the network layer grouping.

[0061] Figure 3 It is a schematic flow diagram of establishing a device load monitoring system.

[0062] Figure 4 It is a schematic flow diagram of dynamically allocating the network layer.

[0063] Figure 5 It is a schematic flow diagram of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present application clearer, the following further describes the present application in detail with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0065] Embodiment 1

[0066] As Figure 1 shown, a sensor edge - side system deployment based on the fusion and compression of the large - model network layer includes:

[0067] A network layer fusion module, which is used to perform cross - layer heterogeneous fusion on each internal network layer of the large model deployed in the cloud to obtain the fused large model;

[0068] The network layer feature classification module is used to traverse the number of network layers with direct dependencies on the data flow path of the internal network layer of the fused large model through the depth-first search algorithm, and classify the network layers with direct dependencies into the same network layer group through a preset layer number threshold. The direct dependency is a unidirectional or bidirectional data transfer relationship formed by the output of a certain network layer directly serving as the input of another network layer or the input of a certain network layer directly originating from the output of another network layer; calculate the feature values of each network layer group according to the computational volume, memory access frequency of each network layer group, as well as the computational power, memory capacity of the device and the real-time and accuracy requirements of the application scenario, and classify the types of each network layer group by combining the feature values of each network layer group with a preset feature value threshold;

[0069] The large model deployment module is used to calculate the load values of each device through a load formula, and allocate and deploy each network layer group to the corresponding device according to the load values of each device and the types of each network layer group using a preset network layer allocation model.

[0070] The technical solution of the present invention optimizes the inter-layer structure inside the model through the network layer fusion module, reducing the risk of misdeletion of multi-modal feature extraction parameters by traditional pruning and quantization techniques. Specifically, it reduces the computational steps and memory occupancy through inter-layer fusion, and based on the depth-first search (DFS) grouping strategy of the direct dependencies between network layers, it can identify the key paths for multi-modal data feature extraction, avoiding the loss of modal features caused by one-size-fits-all parameter pruning. The network layer feature classification module dynamically adjusts the quantization strategy according to the computational power and memory capacity of the device, solving the problem of large differences in the numerical distributions of different modal data. The network layer feature classification module and the large model deployment module optimize the distributed computing efficiency and reliability based on the dynamic scheduling strategy of the load formula, combined with the computational power of the device and the network layer feature classification results. In addition, through the feature classification module, this technical solution comprehensively considers the computational power, memory capacity of the device and the scenario requirements, finely groups the network layers, and when deploying the model, dynamically adjusts the deployment device of the model according to the load conditions of each device and the remaining resource amounts of memory and computing power, achieving a balance between accuracy and resource consumption.

[0071] (1) In the network layer fusion module, the specific method for performing cross-layer heterogeneous fusion on each internal network layer pair of the large model deployed in the cloud is as follows:

[0072] The network layer pairs for fusion include a convolutional layer and a bias layer, a convolutional layer and a pooling layer, a fully connected layer and a convolutional layer, an activation function and a regularization layer, a convolutional layer and a recurrent neural network, an attention mechanism and a convolutional layer, an attention mechanism and a fully connected layer, a graph neural network and a convolutional layer;

[0073] The method of fusing the convolutional layer and the bias layer is to integrate the bias values into the weight matrix of the convolutional kernel through a preset mapping rule, so that the bias can be automatically added during the convolution operation, reducing the extra bias calculation operations. For each output channel c out (0 ≤ c out < c out ) and each position (i1, j1, c in ) on the input channel C in , the formula is expressed as:

[0074]

[0075] Among them, W (with dimensions k h × k w × C in × C out ) represents the weight matrix of the convolutional kernel in the convolutional layer, b (with dimensions C out ) represents the bias vector, W′ represents the weight matrix of the convolutional kernel after fusing the convolutional layer and the bias layer, k h represents the size of the convolutional kernel in the height direction, which determines the coverage range of the convolutional kernel in the vertical direction; k w represents the size of the convolutional kernel in the width direction, which determines the coverage range of the convolutional kernel in the horizontal direction; C in represents the number of channels of the input feature map, reflecting the dimensional characteristics of the input data; i1 represents the index of an element on the input feature map in the height direction during the fusion process of the convolutional layer and the bias layer, and j1 represents the index of an element on the input feature map in the width direction during the fusion process of the convolutional layer and the bias layer; C out represents the number of channels of the output feature map, representing the dimension of the output data after the convolution operation; x1 represents the input data, which is the original data for the convolution operation; y1 represents the output data obtained after the fused convolution operation. By combining the bias and the convolutional kernel weights, the step of calculating the bias separately is eliminated, reducing the computational amount during inference. The fused weight matrix W′ is more likely to maintain the consistency of the numerical distribution during quantization, reducing the accumulation problem of quantization errors for multi-modal data.

[0076] The method of fusing the convolutional layer and the pooling layer is that during the process of the convolutional kernel sliding along the input feature map, for each element within each sliding window, first perform the convolution weighted summation, and then perform the pooling operation on the convolution output within the same window. Taking max pooling as an example, in it, the fusion method is that during the process of the convolutional kernel sliding on the input feature map, for each element within each sliding window, while performing the convolution weighted summation calculation, compare the element values to determine the maximum value, so as to directly obtain the max pooling result during the convolution calculation, avoiding the step-by-step calculation of first convolution and then pooling. The formula is expressed as:

[0077]

[0078] y pool = max(y conv )

[0079] y2 = y pool

[0080] Where h and w respectively represent the starting positions of the current sliding window in the height and width directions on the input feature map; k1 represents the convolution kernel size; W represents the convolution kernel weight matrix; x2 represents the input feature map; i2 represents the index of the convolution kernel in the height direction during the fusion process of the convolutional layer and the pooling layer; j2 represents the index of the convolution kernel in the width direction during the fusion process of the convolutional layer and the pooling layer; y conv represents the result of the convolutional weighted summation, y pool represents the result after pooling, and y2 represents the output result of the fusion of the convolutional layer and the pooling layer. By fusing the convolutional layer and the pooling layer, the length of the data channel is shortened, and the GPU utilization rate is improved. At the same time, the pooling operation directly acts on the convolutional output to suppress noise interference, especially suitable for scenarios with large local feature differences in multimodal data.

[0081] The method of fusing the fully connected layer and the convolutional layer is to replace the fully connected layer with global average pooling. When using global average pooling to replace the fully connected layer at the end of the convolutional neural network, assume that the size of the output feature map of the convolutional layer is H1×W1×C1. First, perform global average pooling on the output of the convolutional layer to obtain the pooled feature vector g, and each of its elements The calculation formula is:

[0082]

[0083] Where H1, W1, and C1 are the size parameters of the output feature map of the convolutional layer, representing the height, width, and number of channels of the feature map respectively; x3 represents the output feature map of the convolutional layer, which is the original data for the global average pooling operation; represents the pooled feature vector after global average pooling and is used as the input data for subsequent processing. The proportion of the fully connected layer parameters is over 80%. Global average pooling reduces the number of parameters through dimensionality reduction, alleviating the pressure on resource-constrained edge devices; i3 represents the index of the channel dimension, used to identify different feature channels in the feature map. Global pooling preserves the overall distribution of the spatial features, and the two-dimensional features of each channel are compressed into one-dimensional scalars, which not only preserves the global information but also achieves efficient parameter compression.

[0084] The method of fusing the activation function and the regularization layer is to preprocess the output value of the fully connected layer according to the deterministic rules of the regularization layer in the inference stage, and directly input the preprocessed result into the activation function to complete the non-linear transformation. For example, when Dropout and ReLU are merged, the fusion method is to directly calculate and activate through ReLU according to the Dropout rule in the inference stage (i.e., not randomly setting to 0) during the calculation of the output of the fully connected layer. In the inference stage, assuming the output of the fully connected layer is x4, the formula is expressed as:

[0085] y4 = max(0, x4)

[0086] Among them, x4 represents the output value of the fully connected layer, which is the original data for the fusion operation; y4 is the output value after fusion in the inference stage. Merging the BN layer and the activation function eliminates the overhead of layer-by-layer calculation. Embedding the regularization parameters into the activation function can dynamically adapt to the distribution differences of multi-modal data.

[0087] The method of fusing the convolutional layer and the recurrent neural network is to first extract the image spatial features of the convolutional layer as the input for the fusion of the recurrent neural network, and then introduce a convolutional operation in specific steps inside the recurrent neural network fusion to achieve the joint learning and information transfer integration of spatial and temporal features. Assuming the video contains T frames, the size of each frame of the image is H2×W2×C2, and the size of the convolutional kernel is k1. First, extract the features of the convolutional layer for each frame to obtain the feature sequence F, and then input the feature sequence F into the recurrent neural network (RNN / LSTM) for temporal modeling. In specific steps inside the recurrent neural network, a convolutional operation can also be introduced, such as performing a convolutional transformation on the hidden state h t and then using it for subsequent calculations. The formula is expressed as:

[0088] F = f1, f2, …, f T

[0089]

[0090] Among them, T represents the number of video frames, reflecting the time dimension of the video data; H2, W2, C2, and k1 are image-related parameters, representing the height, width, number of channels, and size of the convolutional kernel of the image respectively; F represents the feature sequence; represents the features of each frame extracted by the convolutional layer, constituting the feature sequence input to the recurrent neural network; represents; is the hidden state of the recurrent neural network, used to store and transfer temporal information; W ih represents the weight matrix from the input to the hidden state in the recurrent neural network; W hh represents the weight matrix of the weight matrix from the hidden state to the hidden state in the recurrent neural network, W ih and W hhUsed to control the update of the hidden state; b h Represents the bias vector of the recurrent neural network; Represents the hidden state after convolutional transformation. By sharing the feature extraction path through the fusion module, multi-branch repeated calculations are avoided.

[0091] The method of fusing the attention mechanism with the convolutional layer is to introduce an attention weight matrix when calculating the feature map in the convolutional layer to dynamically adjust the element weights of the local feature regions corresponding to each convolutional kernel. The formula is expressed as:

[0092]

[0093] Among them, k1 represents the size of the convolutional kernel in the convolutional layer; W represents the convolutional kernel weight matrix; A represents the attention weight matrix, with dimensions of H3×W3×C3, used to adjust the feature importance; R is the local feature region obtained by the convolutional kernel calculation; y conv Represents the convolutional result of the convolutional layer; y5 is the result after fusing the attention mechanism with the convolutional layer; i5 represents the position index of the convolutional kernel in the height direction during the fusion of the attention mechanism with the convolutional layer; j5 represents the position index of the convolutional kernel in the width direction during the fusion of the attention mechanism with the convolutional layer; c represents the channel index of the feature map. By dynamically allocating resources through the attention weights, the problem of incorrect deletion of key parameters in multi-modal data is solved. The fusion of the attention mechanism and convolution endows the model with stronger semantic modeling ability and task adaptability while maintaining the advantages of convolutional inductive bias through local-global feature complementarity, dynamic weight allocation, and parameter efficiency optimization.

[0094] The method of fusing the attention mechanism with the fully connected layer is to multiply the attention weight matrix on the basis of the neuron connection weights in the fully connected layer, thereby changing the weight distribution of the neuron connections and highlighting the information transfer process of important neuron connections. The formula is expressed as:

[0095] q = W fc x6

[0096]

[0097] That is:

[0098] y6 = A × (W fc x6)

[0099] Among them, i6 represents the index of the output neuron, determining each dimension of the result; j6 represents the index of the intermediate feature, determining the object for adjusting the connection weights from the input to the output; W fcis the original weight matrix of the fully connected layer, which is used to control the weights of neuron connections; x6 is the input value, which is the feature data vector input to the current attention mechanism and fully connected layer fusion module after being processed by the previous network layer; y6 is the output value, which is the output feature vector of the attention mechanism and fully connected layer fusion module, and is obtained through the linear transformation of the fully connected layer and the weight adjustment of the attention mechanism. y6 contains the enhancement and screening of the input data features by the fusion module, and will be used for subsequent model task processing, such as class judgment in classification tasks, numerical prediction in regression tasks, etc. It is an important intermediate output result to promote the model to complete the established tasks; A is the attention weight matrix, which is used to adjust the neuron connection weights; N1 represents the number of neurons; q is W fc The calculation result of x6 is the intermediate feature vector after the fully connected layer linearly transforms the input value x6. It carries the information of the input data after being preliminarily processed by the fully connected layer, and provides the basic data representation for the subsequent attention mechanism to adjust the weights. In the fully connected layer, the input value x6 and the original weight matrix W fc perform a matrix multiplication operation. The generated q integrates the information of the input data in different dimensions. Its dimension is related to the number of output neurons of the fully connected layer, and it contains the feature information that is important for subsequent decision-making or feature extraction. It is a key intermediate product in the fusion process of the attention mechanism and the fully connected layer. This fusion process dynamically allocates resources through attention weights to solve the problem of misdeleting key parameters in multi-modal data. By endowing the model with the ability to dynamically select important features through the attention mechanism, it improves the expression efficiency and interpretability of the fully connected layer, and at the same time avoids the loss of expression ability caused by completely replacing the fully connected layer.

[0100] The method of fusing the graph neural network and the convolutional layer is to introduce a convolutional operation mechanism in the graph convolutional network. For graph-structured data, the adjacency matrix is used to define the neighborhood relationship between nodes, and the feature information of adjacent nodes is aggregated to complete local feature extraction. For the node i7 of the graph-structured data, its feature update formula is:

[0101]

[0102] That is:

[0103]

[0104] Among them, N2 is the number of nodes in the graph-structured data, which reflects the scale of the graph data; σ is the activation function, which is used to perform a non-linear transformation on the calculation result; N2(i7) is the set of neighborhood nodes of node i7, which defines the neighborhood range participating in the node feature update; W″ is the fused weight matrix, which is used to weight the neighborhood node features; is the feature vector of the neighborhood node j7; is the updated feature of node i7; It is the weighted sum of neighborhood node features. By aggregating adjacent node information through graph convolution, it solves the problem of data loss in distributed computing and improves robustness by fusing local convolution and global graph structure information. At the same time, through local-global feature complementarity, dynamic attention weighting, and multi-modal unified representation, it breaks through the limitations of traditional models in dealing with irregular data and computing efficiency. Its core significance lies in constructing a more general, efficient, and interpretable deep learning framework, providing new technical paths for complex scenarios (such as hyperspectral detection and multi-modal analysis).

[0105] (2) In the network layer feature classification module, on the data flow path between the internal network layers of the large model after fusion in the network layer, the specific method of classifying the types of each network layer group by traversing the number of network layers with direct dependency relationships through the depth-first search algorithm and using the characteristic values of each network layer group combined with the preset characteristic value threshold is as follows:

[0106] As Figure 2 shown, for the network layer L1 as the input layer, its dependency relationship set For the network layer L i (i > 1), if its input data comes from the network layer L j (j < i), then add L j to the dependency relationship set D i of the network layer L i , and the relationship between the network layer L i and the network layer L i in its dependency relationship set D j is a direct dependency relationship; if the output data of the network layer L i (i > 1) becomes the input of the network layer L k (k > i), then add the network layer L k to the dependency relationship set D k of the network layer L i , and the relationship between the network layer L k and the network layer L k in its dependency relationship set D i is a direct dependency relationship;

[0107] According to the dependency relationship sets of each network layer, form a data flow path, and use the depth-first search algorithm to traverse the number of times n that specific data is transmitted between the corresponding network layers on the data flow path. The formula is expressed as:

[0108] D i = D i ∪{L j}

[0109] Dk = D k ∪ {L i}

[0110] D i = {L j | R i,j = 1, j ≠ i}

[0111]

[0112] Wherein, R i,j is whether there is a direct dependency between network layer i and network layer j. If there is, then R i,j = 1, otherwise R i,j = 0; D i is the set of dependency relationships of network layer L i . n represents the number of network layers with direct dependency relationships, and n is the total number of network layers;

[0113] Set the dependency layer threshold n min , if n ≥ n min , then classify each network layer passed during the transfer of specific data into the same network layer group.

[0114] By traversing the direct dependency relationships through DFS (such as the output of the convolutional layer directly serving as the input of the activation layer), clarify the critical path of data flow, and classify continuously dependent layers into the same group, which can reduce cross-device data transmission (such as classifying 5 consecutive convolutional layers into the same group to avoid repeated transmission of intermediate feature maps between devices), and reduce the inference latency. Moreover, by restricting the maximum number of layers within the group, the quantization error is restricted to the local path, improving the inference accuracy of multi-modal tasks. After dividing the groups, it is convenient to deploy different groups according to the capabilities of the devices.

[0115] (3) In the network layer characteristic classification module, the specific method for calculating the characteristic values of each network layer group according to the computational amount, memory access frequency of each network layer group, and the computational capabilities, memory capacity, real-time and accuracy requirements of the application scenario of the device, and classifying the types of each network layer group through a preset characteristic value threshold is as follows:

[0116]

[0117] Wherein, P is the characteristic value of network layer p of the network layer group, C p is the computational amount of network layer p of the network layer group, is the weight of the computational amount of network layer p of the network layer group, M p is the memory access frequency of network layer p of the network layer group, is the weight of the memory access frequency of network layer p of the network layer group, and the weight of the computational amount of network layer p of the network layer group and the weights of memory access frequencies Determined according to the computing power of the device, the memory capacity, and the real-time and accuracy requirements of the application scenario;

[0118]

[0119] Among them, O p is the computing amount of network layer p' in network layer group p, and l is the total number of network layers in network layer group p;

[0120]

[0121] Among them, f c is the number of times that network layer group p accesses the device memory during the c-th run, and m is the number of simulated runs of the cloud-deployed model;

[0122] Set the first characteristic value threshold T1 and the second characteristic value threshold T2, where T1 < T2. When P p < T1, network layer group p is determined to be memory access intensive; when P p > T2, network layer group p is determined to be computation intensive; when T1 ≤ P p ≤ T2, network layer group p is determined to be hybrid.

[0123] Characteristic value classification (based on computing amount, memory access frequency, etc.) allows dynamic adjustment of network layer allocation according to device capabilities (such as the low memory capacity of edge devices), so as to allocate groups with different characteristics to the most suitable devices, ensure the matching of device performance and network layers, maximize resource utilization, enhance the robustness of multi-modal tasks, and balance real-time performance and accuracy. And during this process, the weights of computing amount and memory access rate can be dynamically adjusted, improving the deployment flexibility.

[0124] (4) In the network layer dynamic allocation module, the specific method of calculating the load value of each device through a preset load formula and allocating each network layer group to the corresponding device according to the load value of each device and the type of each network layer group using a preset network layer allocation model is as follows:

[0125] As Figure 3 shown, during this process, a device load detection system is established to continuously track the load status of the edge server and the end test bed, detect CPU usage rate, memory occupancy, network bandwidth, GPU utilization rate, cache hit rate, and generate a load report. The load value calculation formula is as follows:

[0126] S = α × U cpu + β × U mem + γ × U net + δ × L

[0127] Among them, S is the device load value, Ucpu is the CPU usage rate of the device, U mem is the memory occupancy rate of the device, U net is the network bandwidth utilization rate of the device, L is the task queue length of the device, α is the weight of the device's CPU usage rate, β is the weight of the device's memory occupancy rate, γ is the weight of the device's network bandwidth utilization rate, δ is the weight of the device's task queue length, and α + β + γ + δ = 1;

[0128] As Figure 4 shown, the dynamic allocation of the network layer is then implemented according to the load value of the device.

[0129] The device includes an edge device and an edge server. First, deploy the memory access-intensive network layer group on the edge device. The deployment principle is to sequentially select, from the set of edge devices where the available memory is greater than the average available memory of all edge devices, the edge device with the smallest current memory usage rate and less than the memory usage rate threshold of the memory access-intensive network layer group for deployment. The formula is expressed as:

[0130] Let the set of edge devices be E = {e1, e2,..., e z}, where e z is the z-th device in the set of edge devices E. Define M(e z ) to represent the available memory resource amount of the z-th device in the set of edge devices E, and M avg to represent the average available memory resource amount of all devices in the set of edge devices E. Its calculation formula is

[0131]

[0132] where, e assign represents the edge device finally allocated to the memory access-intensive network layer group, represents, in the set of edge devices E, screening out all edge devices with available memory resource amount M(e z ) greater than the average available memory resource amount M avg of all edge devices and finding the edge device with the largest available memory resource amount M(e z ).

[0133] After completing the deployment of the memory access-intensive network layer group, then deploy the compute-intensive network layer group on the edge server. The deployment principle is to sequentially select, from the set of edge servers where the computing power is greater than the average computing power of all edge servers, the edge server with the smallest current load and less than the load threshold of the compute-intensive network layer group for deployment. The formula is expressed as:

[0134] D = {d1, d2,..., dy}

[0135]

[0136] Among them, D is the set of edge servers, and d n is the y-th edge server in the set of edge servers D, and c power_avg is the average computing power of all edge servers in the set of edge servers D, and C power (d y ) is the computing power of the y-th edge server in the set of edge servers D, and S threshold is the preset load threshold applicable to the network layer group of compute-intensive types.

[0137] For the hybrid network layer group, select the device set D that meets the computing power threshold and available memory threshold of the hybrid network layer group mix , the device set D mix includes end-side devices and edge servers. Among the device sets D that meet the computing power threshold and available memory threshold of the hybrid network layer group mix , select the device with the lowest load value for the deployment of the hybrid network layer group. The formula is expressed as:

[0138] D mix ={d y ∈D | C power (d y )>C power_mix_threshold ∧M available (d y )>M available_mix_threshold}

[0139] Among them, C power_mix_threshold is the preset computing power threshold applicable to the hybrid network layer group, and M available_mix_threshold is the preset available memory threshold applicable to the hybrid network layer group;

[0140] Then, among the device sets D that meet the computing power and available memory requirements mix , select the device with the lowest load for the deployment of the hybrid network layer group. The formula is expressed as:

[0141]

[0142] If, during the deployment of a memory - intensive network layer group, in the set of edge devices where the available memory resource amount is greater than the average available memory resource amount of all edge devices, the memory utilization rate of all current edge devices is greater than the memory utilization rate threshold of the memory - intensive network layer group, then deploy the remaining undeployed memory - intensive network layer groups to the edge servers. The deployment principle is to sequentially select, from the set of edge servers where the available memory amount is greater than the average available memory amount of all edge servers, the edge server with the lowest current load value and a current memory utilization rate less than the memory utilization rate threshold of the memory - intensive network layer group for deployment;

[0143] If, during the deployment of a compute - intensive network layer group, in the set of edge servers where the computing power is greater than the average computing power of all edge servers, the load value of all current edge servers is greater than the load value threshold of the compute - intensive network layer group, then deploy the remaining undeployed compute - intensive network layer groups to the edge devices. The deployment principle is to sequentially select, from the set of edge devices where the computing power is greater than the average available memory amount of all edge devices, the edge device with the lowest current memory utilization rate and a current load value less than the load value threshold of the compute - intensive network layer group for deployment.

[0144] During the allocation process, the cooperation between devices and the data transmission cost also need to be considered. If device d i and d j have a data transmission rate of R ij , and the preset transmission rate threshold is R threshold , when R ij >R threshold , and their load conditions and grouping characteristics match, then preferentially allocate the relevant groups to these devices to reduce data transmission latency and improve the overall processing efficiency.

[0145] Through the above - mentioned process of dynamically allocating network layer groups based on the real - time status of devices and grouping characteristics, all key indicators of devices, different characteristics of groups, and cooperation factors between devices are comprehensively and meticulously considered. The device load assessment index S is used to quantify the device load situation. According to different grouping characteristics (compute - intensive or memory - intensive), the corresponding allocation formula is used to accurately schedule the groups to the appropriate devices. At the same time, the data transmission rate between devices is fully considered, and the groups are preferentially allocated to the device combinations with high transmission rates and matching loads and characteristics. These series of operations are closely coordinated, effectively reducing data transmission latency and greatly improving the utilization efficiency of system resources, thereby ensuring that the entire system can operate efficiently and stably in complex and changing application scenarios, providing a reliable and optimized solution for various tasks relying on network layer group processing.

[0146] Evaluate the load value through four-dimensional weighted evaluation of CPU, memory, network bandwidth, and task queue. The evaluation dimension priority can be dynamically adjusted according to different device types through weight adjustment. Generally speaking, the computing resources and computing power of edge servers are stronger than those of end-side devices. Memory-intensive tasks are preferentially deployed on end-side devices to take advantage of their low-latency characteristics, and at the same time, overflow is avoided by controlling the memory usage threshold. Compute-intensive tasks are concentrated on edge servers, and the computing resource utilization is improved through dynamic scheduling based on load thresholds. In addition, a hierarchical deployment strategy is adopted: memory-intensive tasks are preferentially deployed on end-side devices with sufficient memory (such as mobile terminals) to reduce data migration overhead by taking advantage of the local memory bandwidth advantage, and memory access latency can be reduced. Compute-intensive tasks are directed to high-computing-power edge servers, and the computing efficiency is improved through GPU / NPU acceleration units, and the task processing speed is increased. Hybrid tasks select the optimal device based on dynamic thresholds to achieve the best ratio of CPU / memory resources, and can avoid a single resource bottleneck. When the memory of the end-side device is insufficient, it is automatically migrated to the edge server, and vice versa, when the computing resources are overloaded, it is migrated back to the end-side device. Through a two-screening mechanism (first excluding devices with below-average capabilities, and then selecting the node with the lowest load), the "avalanche effect" can be effectively avoided.

[0147] (5) It also includes a network layer migration module for unloading network layer data of devices in the deployed network layer group, transmitting it using an adaptive network transmission protocol according to the device network connection status and bandwidth, and transmitting the unloaded network layer data to the device that determines to receive the network layer group migration using a data transmission method of multi-channel transmission and / or fragmentation transmission and / or resume from breakpoint. The device that has received the network layer group migration selects a corresponding decompression algorithm for deployment according to its own computing power and storage resource status for the received network layer data.

[0148] For a device that determines to receive a network layer group but has an unstable network connection, before starting the unloading process of the network layer group on the corresponding deployed device, evaluate the network connection status of the device that determines to receive the network layer group but has an unstable network connection: when the signal strength Q signal is higher than the preset signal strength threshold Q signal_threshold and the packet loss rate Q loss is lower than the preset packet loss rate threshold Q loss_threshold and the latency Q delay is less than the preset latency threshold Q delay_threshold then perform network layer group unloading;

[0149] For a device that has a deployed network layer group but insufficient bandwidth, before transmitting the network layer data to the device that determines to receive the network layer group, measure the bandwidth upper limit B max and the real-time available bandwidth Bavail If the bandwidth upper limit B max meets the transmission requirements of the network layer packet data to be offloaded, it is directly transmitted; if the bandwidth upper limit B max does not meet the transmission requirements of the network layer packet data to be offloaded, hierarchical compression processing is performed on the network layer packet data to be offloaded. Let the total number of network layers in the network layer group be l. For the t-th network layer among them, its compression ratio is set as r t , and the original data s of the t-th network layer is made to t meet

[0150] The specific solution is as follows. For devices with unstable network connections and that are determined to receive network layer packets (including end-side devices and edge servers), before starting the offloading process, a comprehensive and detailed assessment of their network connection status is required. Set the comprehensive network connection quality assessment index as Q, and this index covers signal strength Q signal , packet loss rate Q loss , and latency Q delay and other key parameters. Only when Q meets specific conditions, that is, the signal strength Q signal is higher than the pre-set strength threshold Q signal_threshold (referring to relevant industry standards and the actual performance requirements of the device, set as -70dBm), the packet loss rate Q loss is lower than the set packet loss rate threshold Q loss_threshold (for example, set as 5%), and the latency Q delay is less than the latency threshold Q delay_threshold (such as set as 100ms), can the model offloading transmission operation be carried out.

[0151] In terms of the transmission strategy, a multi-channel transmission method is adopted, and the number of channels is set as n c . Using a dynamic channel allocation strategy based on the network real-time bandwidth and data volume, the data is reasonably allocated to each channel to enhance the stability and efficiency of the transmission; at the same time, data fragmentation transmission technology is adopted, and the fragmentation size is set as s f . According to the structural characteristics and importance level of the data, an appropriate fragmentation size is determined to ensure the reliability of the data transmission. Combining with the breakpoint resumption mechanism, the maximum number of retransmissions for breakpoint resumption is set as n r (considering the instability degree of the network and the criticality of the data comprehensively, the value range is 3 - 5 times), and reliable data verification algorithms such as CRC verification A c are adopted to ensure the integrity and accuracy of the model data in an unstable network environment through a strict verification process.

[0152] For devices with limited bandwidth and that have been allocated network layer packets, before transmitting the model data, accurately measure the bandwidth upper limit B max and the real-time available bandwidth B of the deviceavail Perform hierarchical compression processing on the network layer packet data to be unloaded. Assume that the packet model is divided into L layers. For the i-th layer, its compression ratio is set to r i When determining the compression ratio, it is strictly set according to the degree of influence of each layer on the model performance. For the data part that has relatively little influence on the model performance (such as some auxiliary feature extraction layers), a higher compression ratio is adopted (for example, the compression ratio r i is set between 0.8 and 0.9), while for the key parts (such as the core feature extraction layer and the layer where important parameters are located), a low compression ratio is adopted (such as the compression ratio r i is set between 0.1 and 0.3) or no compression is performed, and it is ensured that (where s i is the size of the original data of the i-th layer).

[0153] On the device side, select an appropriate decompression algorithm and deployment strategy according to the computing power and storage resource status of the device. For example, if the device has weak computing power (such as the CPU main frequency is lower than 1 GHz and the memory is less than 512 MB) but relatively sufficient storage resources, a decompression algorithm with lower computational complexity and better data integrity preservation after decompression can be selected, such as an improved algorithm based on Huffman coding; if the device has strong computing power (such as the CPU main frequency is higher than 2 GHz and the memory is greater than 2 GB) and tight storage resources, a decompression algorithm with fast decompression speed but possible data redundancy can be selected, such as a fast decompression version of the LZ77 algorithm. After decompression, obtain the hardware information of the local device (including CPU model, number of cores, main frequency, GPU model, video memory size, memory capacity, etc.) and software environment information (operating system type and version, installed deep learning framework version, etc.). Based on this information, adapt the running environment of the model. For example, if the local GPU video memory is small, appropriately adjust the batch size parameter of the model; if the deep learning framework version is low, perform compatibility optimization on some operators of the model to ensure that the model can run efficiently on the local device.

[0154] After completing the environment adaptation, deploy the model to the target running environment locally. This includes creating the processes or services required for the model to run in the local system and configuring the corresponding startup parameters and resource limits. For example, allocate specific CPU cores and memory space for the model process to prevent the model from occupying excessive system resources and affecting other applications when running. At the same time, establish communication interfaces between the model and other relevant local system components (such as data acquisition modules, result output modules) to ensure that data can flow smoothly into and out of the model, making the model an organic part of the local application system. After the deployment is completed, send a deployment success notification containing the model version number, deployment time, deployment environment information, etc. to the model update module. The above model deployment method is also applicable to the process of deploying from the cloud to edge devices and edge servers.

[0155] Through the above offloading strategies for devices with different network conditions, the success rate and efficiency of model offloading can be effectively improved, ensuring the smooth migration and deployment of network-layer packets between different devices, thus guaranteeing the stable operation of the entire system and providing strong support for the efficient application of deep learning models in scenarios of integrating sensor multi-modal data processing.

[0156] (6) It also includes a large model update management module, which is responsible for the update work of the large model, and real-time tracks performance indicators such as accuracy, recall rate, processing speed, and resource utilization rate of the large model on edge devices and edge servers. According to factors such as performance indicator changes, data changes, and time cycles, it triggers the large model update process, conducts data collection and preprocessing, large model training and evaluation, large model deployment and update, and implements the gray release strategy to ensure the smooth transition of the large model update and system stability.

[0157] Its operation mode is based on factors such as performance indicators, data changes, and time cycles. When these triggering conditions are met, the model update process is officially started. At the same time, after receiving the deployment success notification sent by the network layer migration module, the large model update module will extract key contents such as the large model version number, deployment time, and deployment environment information from the notification to make full preparations for the subsequent large model evaluation.

[0158] The triggering of the large model update is based on the following mechanisms:

[0159] Based on performance indicators, the system continuously monitors the performance indicators of the large model in actual applications, such as accuracy A, recall rate R, mean squared error MSE, etc. Set performance thresholds through a large number of experiments and business experience. When the accuracy A of the large model is lower than the preset accuracy threshold A threshold , for example, in the target detection task, the preset accuracy threshold A threshold = 0.85, and the current large model accuracy A = 0.83; or the mean squared error MSE is higher than the preset mean squared error threshold MSE thresholdWhen this happens, it will trigger the update of the large model, which can be expressed by the mathematical formula: when A < A threshold or MSE > MSE threshold , that is, it triggers an update.

[0160] Based on data changes, regularly conduct in-depth analysis of the distribution of input data. Let the data feature vector be X = [x1, x2, …, x k . By accurately calculating the statistics of the data distribution, such as the mean μ i and variance to accurately measure data changes. When the mean or variance of the data features changes by more than a certain proportion, that is, when any of the following conditions is met:

[0161]

[0162] it will trigger the update of the large model, enabling the large model to adapt to the new data distribution in a timely manner.

[0163] Based on the time period, set a fixed time period T according to the actual business needs. For example, it can be set to once a week (T = 7 days) or once a month (T = 30 days), etc. When the time reaches the set period, regardless of whether there are obvious changes in the performance of the large model and the data, a model update check will be triggered to ensure that the large model will not lag behind in performance and adaptability due to long-term non-update. That is, when the time t ≥ T, an update check is triggered.

[0164] The large model update process is as follows:

[0165] First, adopt the gray release strategy. Select a small number of representative parts (set the selection ratio as α, for example, α = 0.1) from all users or nodes, and let the newly deployed large model start running on these selected users or nodes. During the operation of the large model, this module will collect a large amount of performance data in real time, including but not limited to the accuracy A new , recall rate R new , response time and other key indicators. At the same time, closely monitor the stability of the large model operation, and always pay attention to whether there are situations such as crashes, memory leaks, and abnormal resource occupancy to ensure that the performance of the large model in the actual usage scenario meets expectations.

[0166] Then comes the performance evaluation and decision-making stage. The large model update module will conduct a detailed comparative analysis of the performance data of the new large model with that of the original large model. If the new large model performs better than the original large model in key performance indicators such as accuracy and recall rate, and also shows good stability without any abnormal situations, then the usage scope of the new large model will be gradually expanded according to the established expansion strategy. The expansion strategy usually involves increasing a certain proportion (such as 10%) of users or nodes each time until the new large model completely replaces the original model, thus completing the key step of large model update. On the contrary, if the new large model has performance problems, such as the accuracy not reaching the expected improvement amplitude, the response time being too long affecting the user experience, or stability problems, such as frequent crashes and high memory occupancy resulting in slow system operation, the large model update module will feedback the detailed test results, including performance data, exception logs and other information to the cloud server. After receiving the feedback, the cloud server will further optimize the model. During this period, the original large model continues to run stably in the system to ensure the normal operation of the business.

[0167] Finally is the update record and management stage. During the entire large model update process, the large model update module will record every key information in detail, including the specific reasons for update triggering, the start and end times of the update, the comparison data of the performance of the new and old large models, the test results during the gray release stage, etc. These records will be stored in a dedicated database for convenient subsequent in-depth analysis of the large model performance. By studying the historical data, the effect of each large model update can be evaluated, experiences and lessons can be summarized, and thus strong data support can be provided for future large model update decisions, continuously optimizing the large model update strategy and process, and improving the performance and stability of the entire deep learning large model.

[0168] Embodiment 2

[0169] A method for deploying a sensor edge system based on cross-layer heterogeneous fusion compression of a large model network layer, as Figure 5 shown, includes the following steps:

[0170] Perform cross-layer heterogeneous fusion on each internal network layer pair of the large model deployed in the cloud to obtain the fused large model;

[0171] On the data flow path of the internal network layer of the fused large model, traverse the number of network layers with direct dependencies through the depth-first search algorithm, and classify the network layers with direct dependencies into the same network layer group through a preset layer threshold. The direct dependency is a unidirectional or bidirectional data transfer relationship formed when the output of a certain network layer is directly used as the input of another network layer or the input of a certain network layer directly comes from the output of another network layer; calculate the characteristic values of each network layer group according to the computational amount, memory access frequency of each network layer group, and the computational power, memory capacity of the device, and the real-time and accuracy requirements of the application scenario, and classify the types of each network layer group by combining the characteristic values of each network layer group with a preset characteristic value threshold;

[0172] Calculate the load values of each device through the load formula, and allocate and deploy each network layer group to the corresponding device according to the load values of each device and the types of each network layer group using a preset network layer allocation model.

[0173] Embodiment 3

[0174] A computer program product includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, the method for deploying a sensor edge system based on large model network layer fusion compression in Embodiment 2 is implemented.

[0175] The content not described in detail in this specification belongs to the prior art well-known to those skilled in the art. Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0176] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0177] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in the block or blocks.

[0178] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in the block or blocks.

[0179] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit the scope of its protection. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: after reading the present invention, those skilled in the art can still make various changes, modifications or equivalent replacements to the specific embodiments of the invention, but these changes, modifications or equivalent replacements are all within the scope of the claims of the invention pending approval.

Claims

1. A sensor edge system deployment system based on large model network layer fusion compression, characterized in that: include: The network layer fusion module is used to perform cross-layer heterogeneous fusion of each internal network layer pair of the large model deployed in the cloud to obtain the fused large model; A network layer characteristic classification module is used to traverse the number of network layers with direct dependencies on the data flow path of the internal network layer of the fused large model through a depth-first search algorithm, and classify the network layers with direct dependencies into the same network layer group through a preset layer number threshold, wherein the direct dependency is a one-way or two-way data transmission relationship formed by the output of a certain network layer directly serving as the input of another network layer or the input of a certain network layer directly derives from the output of another network layer; calculate the characteristic value of each network layer group according to the computing amount and memory access frequency of each network layer group, as well as the computing power, memory capacity of the device and the real-time and accuracy requirements of the application scenario, and classify the type of each network layer group using the characteristic value of each network layer group combined with the preset characteristic value threshold; The large model deployment module is used to calculate the load value of each device through the load formula, and allocate and deploy each network layer group to the corresponding device using a preset network layer allocation model based on the load value of each device and the type of each network layer group.

2. According to claim 1, the sensor edge system deployment system based on large model network layer fusion compression is characterized by: In the network layer fusion module, the specific method of performing cross-layer heterogeneous fusion of each internal network layer of the large model deployed in the cloud is: The network layer pairs to be fused include convolutional layer and bias layer, convolutional layer and pooling layer, fully connected layer and convolutional layer, activation function and regularization layer, convolutional layer and recurrent neural network, attention mechanism and convolutional layer, attention mechanism and fully connected layer, graph neural network and convolutional layer; The method of fusing the convolution layer with the bias layer is to integrate the bias value into the weight matrix of the convolution kernel through a preset mapping rule; The method of fusing the convolution layer and the pooling layer is to first perform the weighted summation of the convolution for each element in the sliding window while the convolution kernel slides along the input feature map, and then perform the pooling operation on the convolution output in the same window; The method of fusing the fully connected layer with the convolutional layer is to replace the fully connected layer with global average pooling; The method of fusing the activation function with the regularization layer is to preprocess the output value of the fully connected layer according to the deterministic rules of the regularization layer in the inference stage, and directly input the preprocessing result into the activation function to complete the nonlinear transformation; The method of fusing the convolutional layer with the recurrent neural network is to first extract the image spatial features of the convolutional layer as the input of the recurrent neural network fusion, and then introduce the convolution operation in a specific step inside the recurrent neural network fusion; The method of integrating the attention mechanism with the convolution layer is to dynamically adjust the element weights of the local feature area corresponding to each convolution kernel by introducing the attention weight matrix when calculating the feature map in the convolution layer; The method of fusing the attention mechanism with the fully connected layer is to multiply the attention weight matrix based on the connection weights of the neurons in the fully connected layer; The method of fusing graph neural networks with convolutional layers is to introduce a convolution operation mechanism into the graph convolutional network, use the adjacency matrix to define the neighborhood relationship between nodes for graph structure data, and aggregate the feature information of adjacent nodes to complete local feature extraction.

3. According to claim 1, the sensor edge system deployment system based on large model network layer fusion compression is characterized by: In the network layer characteristic classification module, the specific method of traversing the number of network layers with direct dependencies on the data flow path between the internal network layers of the large model after the network layer fusion by using a depth-first search algorithm and classifying the network layers with direct dependencies into the same network layer group by using a preset dependency layer number threshold is as follows: For the network layer L1 as the input layer, its dependency set is an empty set; For network layer L that is not an input layer i , where i > 1, if its input data is sourced from network layer L j , where j < i, then add L j to the dependency set D i of network layer L i , where i > 1. The relationship between network layer L i and the network layers L i in its dependency set D j is a direct dependency relationship; if the output data of network layer L i , where i > 1, becomes the input of network layer L k , where k > i, then add network layer L k to the dependency set D k of network layer L i , where k > i. The relationship between network layer L k and the network layers L k in its dependency set D i is a direct dependency relationship; The data flow path is formed according to the dependency set of each network layer, and the depth-first search algorithm is used to traverse the number of times n that specific data is transferred between the corresponding network layers on the data flow path. The formula is expressed as: D i =D i ∪{L j } D k =D k ∪{L i } D i ={L j |R i,j =1,j≠i} Among them, R i,j Is there a direct dependency relationship between network layer i and network layer j? If yes, then R i,j =1, otherwise R i,j =0;D i The network layer L i The dependency set of n represents the number of network layers with direct dependencies, and N is the total number of network layers; Set the dependency layer threshold n min , if n≥n min , then the network layers that specific data passes through when transmitted are classified into the same network layer group.

4. According to claim 1, the sensor edge system deployment system based on large model network layer fusion compression is characterized by: In the network layer characteristic classification module, the characteristic value of each network layer group is calculated according to the computing amount of each network layer group, the memory access frequency, the computing power of the device, the memory capacity, and the real-time and accuracy requirements of the application scenario, and the specific method of classifying the type of each network layer group by using the characteristic value of each network layer group combined with the preset characteristic value threshold is as follows: Where P is the characteristic value of the network layer group network layer p, C p is the computational effort of the network layer group network layer p, is the weight of the computation of the network layer p in the network layer group, M p is the memory access frequency of the network layer p in the network layer group, is the weight of the memory access frequency of the network layer group network layer p, and the weight of the computation amount of the network layer group network layer p and the weight of memory access frequency Determined based on the computing power, memory capacity of the device, and the real-time and accuracy requirements of the application scenario; Among them, O p is the computational effort of network layer p′ in network layer group p, l is the total number of network layers in network layer group p; Among them, f c is the number of times the network layer group p accesses the device memory during the cth run, and m is the number of simulation runs of the cloud-deployed model; Set the first characteristic value threshold T1 and the second characteristic value threshold T2, where T1 < T2. When P p < T1, the network layer group p is determined to be memory access intensive; when P p > T2, the network layer group p is determined to be compute intensive; when T1 ≤ P p ≤ T2, the network layer group p is determined to be hybrid.

5. According to claim 1, the sensor edge system deployment system based on large model network layer fusion compression is characterized by: In the network layer dynamic allocation module, the load value of each device is calculated by a preset load formula, and the specific method of allocating each network layer group to the corresponding device using a preset network layer allocation model according to the load value of each device and the type of each network layer group is as follows: S=α×U cpu +β×U mem +γ×U net +δ×L Among them, S is the equipment load value, U cpu is the device CPU usage, U mem is the device memory usage, U net is the device network bandwidth utilization, L is the device task queue length, α is the weight of the device CPU utilization, β is the weight of the device memory occupancy, γ is the weight of the device network bandwidth utilization, δ is the weight of the device task queue length, α+β+γ+δ=1; The devices include end-side devices and edge servers. First, the memory access-intensive network layer group is deployed on the end-side device. The deployment principle is to select the end-side devices with the smallest current memory usage rate and less than the memory usage rate threshold of the memory access-intensive network layer group from the end-side device set whose available memory is greater than the average available memory rate of all end-side devices for deployment. After the deployment of the memory access intensive network layer group is completed, the computing intensive network layer group is deployed on the edge server. The deployment principle is to select the edge server with the smallest current load and less than the load threshold of the computing intensive network layer group from the edge server set whose computing power is greater than the average computing power of all edge servers for deployment; For a mixed network layer group, filter out the device set D that meets the mixed network layer group computing power threshold and available memory threshold. mix , device set D mix Including end-side devices and edge servers, a set of devices D that meet the hybrid network layer group computing power threshold and available memory threshold mix Select the device with the lowest load value to deploy a mixed network layer team.

6. The sensor edge system deployment system based on large model network layer fusion compression according to claim 5 is characterized by: If, during the deployment of a memory access intensive network layer group, in a set of end-side devices whose available memory resources are greater than the average available memory resources of all end-side devices, the current memory usage of all end-side devices is greater than the memory usage threshold of the memory access intensive network layer group, then the remaining undeployed memory access intensive network layer groups are deployed to edge servers, and the deployment principle is to select edge servers with the smallest current load value and the current memory usage less than the memory usage threshold of the memory access intensive network layer group in turn from a set of edge servers whose available memory is greater than the average available memory of all edge servers for deployment; If during the deployment of a compute-intensive network layer group, in a set of edge servers whose computing power is greater than the average computing power of all edge servers, the load values ​​of all current edge servers are greater than the load value threshold of the compute-intensive network layer group, then the remaining undeployed compute-intensive network layer groups will be deployed to the end-side devices. The deployment principle is to select the end-side devices with the smallest current memory usage and the current load value less than the load value threshold of the compute-intensive network layer group in turn from the set of end-side devices whose computing power is greater than the average available memory of all end-side devices for deployment.

7. According to claim 1, the sensor edge system deployment system based on large model network layer fusion compression is characterized in that: Also includes: The network layer migration module is used to unload network layer data from devices that have deployed network layer groups, and use an adaptive network transmission protocol to transmit the data according to the network connection status and bandwidth of the device. The unloaded network layer data is transmitted to the device that is determined to receive the network layer group migration using multi-channel transmission and / or segmented transmission and / or breakpoint-resume transmission. The device that has received the network layer group migration selects a corresponding decompression algorithm for the received network layer data and deploys it according to its own computing power and storage resource conditions.

8. The sensor edge system deployment system based on large model network layer fusion compression according to claim 7 is characterized by: In the network layer migration module, the specific method of using an adaptive network transmission protocol to uninstall the network layer on deployed devices according to the device network connection status and bandwidth is as follows: For devices that are determined to receive the network layer group but have unstable network connections, before starting the uninstallation process of the network layer group on the corresponding deployed devices, the network connection status of the devices that are determined to receive the network layer group but have unstable network connections is evaluated: when the signal strength Q signal Higher than the preset signal strength threshold Q signal_threshold When the packet loss rate is Q loss Lower than the preset packet loss rate threshold Q loss_threshold When the delay Q delay Less than the preset delay threshold Q delay_threshold When the network layer group is unloaded; For a device that has deployed a network layer group but has insufficient bandwidth, before transmitting network layer data to the device that is determined to receive the network layer group, measure the bandwidth upper limit B of the device that has deployed the network layer group but has insufficient bandwidth max and real-time available bandwidth B avail , if the bandwidth limit B max If the bandwidth limit B meets the transmission requirements of the network layer group data that needs to be offloaded, it is directly transmitted; max If the transmission requirements of the network layer group data to be unloaded are not met, layered compression processing is performed on the network layer group data to be unloaded. Assume that the total number of network layers in the network layer group is l, and the compression ratio of the tth network layer is r. t , and make the original data s of the tth network layer t satisfy 9. A sensor edge system deployment method based on large model network layer fusion compression, characterized in that: The following steps are involved: Perform cross-layer heterogeneous fusion of each internal network layer of the large model deployed in the cloud to obtain the fused large model; On the data flow path of the internal network layer of the fused large model, the number of network layers with direct dependencies is traversed by a depth-first search algorithm, and the network layers with direct dependencies are classified into the same network layer group by a preset layer number threshold, wherein the direct dependency is a one-way or two-way data transmission relationship formed by the output of a certain network layer directly serving as the input of another network layer or the input of a certain network layer directly derives from the output of another network layer; the characteristic value of each network layer group is calculated according to the computing amount of each network layer group, the memory access frequency, the computing power, memory capacity of the device, and the real-time and accuracy requirements of the application scenario, and the type of each network layer group is classified by using the characteristic value of each network layer group combined with the preset characteristic value threshold; The load value of each device is calculated using the load formula, and each network layer group is allocated and deployed to the corresponding device using a preset network layer allocation model based on the load value of each device and the type of each network layer group.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by the processor, the sensor edge system deployment method based on large model network layer fusion compression described in claim 9 is implemented.