A communication edge-cloud segmentation inference method, device, equipment, product and medium for large language models

By segmenting large language models into client and server-side models, and using adaptive encoder and decoder for data compression and reconstruction, the problem of data privacy and communication overhead in the edge network is solved, and efficient reasoning and privacy protection is achieved.

CN120166112BActive Publication Date: 2025-07-22FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510647446.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-07-22
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

Large language models have data privacy vulnerabilities and high communication overhead when deploying in edge networks. In particular, the increase in volume after the original data is encoded into compressed data causes a large amount of time to transmit, and the autoregressive nature of LLMs leads to a secondary increase in the number of iterative tokens, affecting system efficiency.

Method used

The large language model is segmented into client model and server model, compressed at edge devices through adaptive encoder and rebuilt at cloud servers, reducing communication overhead, and transmit compressed data between edge devices and cloud servers through adaptive encoder and decoder, and partial reasoning is performed using the computing power of edge devices.

Benefits of technology

It reduces communication overhead between edge devices and cloud servers, enhances data privacy protection, adapts to different network environments and computing resources, and improves inference efficiency and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120166112B_ABST
    Figure CN120166112B_ABST
Patent Text Reader

Abstract

The present invention discloses a communication edge-cloud segmentation inference method, device, equipment, product and medium for large language models, including: segmenting a large language model into a client model and a server model based on a preset division criterion; performing inference on data to be processed through the client model, and compressing intermediate results through an adaptive encoder, where the client model and the adaptive encoder are deployed on edge devices; reconstructing the compressed intermediate results through an adaptive decoder, and performing inference on the reconstructed intermediate results through the server model to obtain new tokens, where the adaptive decoder and the server model are deployed on a cloud server; obtaining new data to be processed based on the new tokens, and repeating the foregoing steps until a preset condition is reached to obtain a final prediction result. The present invention can not only reduce communication overhead, but also enhance data privacy protection, adapt to different network environments and computing resources, and has good flexibility and scalability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and in particular, to a communication edge-cloud segmentation inference method, device, equipment, product and medium for large language models. Background Art

[0002] The emergence of large language models (LLMs) marks an important progress in the field of artificial intelligence (AI), driving the process of realizing general intelligence. These models have a huge parameter scale and a wide range of pre-trained data sets, showing significant versatility and being able to effectively adapt to a wide range of application fields.

[0003] Currently, the deployment of LLMs mainly focuses on cloud infrastructure because their huge resource requirements make local implementation among edge network users infeasible. Unfortunately, providing LLMs in the cloud introduces inherent drawbacks, most notably significant vulnerabilities in data privacy. Edge devices must transmit raw data (including personal conversations, medical records, and financial information) to remote servers for LLM inference, thus increasing the risks of data leakage, unauthorized access, and abuse. High-profile incidents, such as unauthorized data retention by AI providers and large-scale user interaction leaks, highlight these vulnerabilities.

[0004] To overcome this limitation, segmentation inference divides the AI model by depth, deploying the shallow layers on edge devices while offloading the deep layers to cloud servers. This approach alleviates computational constraints and enhances data privacy simultaneously: (i) by only retaining a small part of the model locally, edge devices utilize the powerful server-side computing capabilities to complete inferences beyond their own processing capabilities; (ii) instead of transmitting raw data, edge devices send compressed data generated at the model cut layer to the server for final processing, thus reducing privacy risks.

[0005] The current main problem is that,

[0006] (1) Raw data (such as text) usually increases in volume after being encoded into compressed data for transformer models, resulting in a large amount of time consumed when transmitting compressed data in segmentation inference.

[0007] (2) The inherent autoregressive nature of LLMs requires iterative generation of tokens until the sentence is completed. Each iteration involves edge-cloud transmission, and the data volume gradually increases. Therefore, the total communication overhead for the number of generated tokens grows quadratically, significantly affecting system efficiency. Summary of the Invention

[0008] The present invention provides a communication edge-cloud segmentation inference method, device, equipment, product and medium for large language models, aiming to effectively solve the above technical problems.

[0009] According to a first aspect of the present invention, the present invention provides a communication edge-cloud segmentation inference method for large language models, the method comprising:

[0010] A segmentation step of segmenting a large language model into a client model and a server model based on a preset partitioning criterion;

[0011] A first inference and encoding step of inferring, by the client model, data to be processed to obtain an intermediate result; and compressing the intermediate result by an adaptive encoder to obtain a compressed intermediate result, wherein the client model and the adaptive encoder are deployed on an edge device;

[0012] A decoding and second inference step of reconstructing, by an adaptive decoder, the compressed intermediate result to obtain a reconstructed intermediate result; and inferring, by the server model, the reconstructed intermediate result to obtain a new token, wherein the adaptive decoder and the server model are deployed on a cloud server;

[0013] Obtaining new data to be processed based on the new token, and repeating the first inference and encoding step to the decoding and second inference step until a preset condition is reached to obtain a final prediction result.

[0014] Further, the preset partitioning criterion includes at least one of the network structure complexity of the large language model, the computing power of the edge device, the computing power of the cloud server, or the data transmission efficiency between the edge device and the cloud server.

[0015] Further, the inferring, by the client model, data to be processed to obtain an intermediate result includes:

[0016] Performing forward propagation of the data to be processed by the client model to a cut layer to obtain activation value data;

[0017] Evaluating the importance of the activation value data and determining the intermediate result according to the evaluation result.

[0018] Further, before the first inference and encoding step, the method further includes:

[0019] Obtaining initial data and performing word segmentation or feature vectorization processing on the initial data to obtain the data to be processed.

[0020] Further, the obtaining new data to be processed based on the new token includes:

[0021] Concatenating the new token with the data to be processed to obtain the new data to be processed.

[0022] Further, the preset condition includes that the new token is a complete sentence.

[0023] According to a second aspect of the present invention, the present invention also provides a communication edge-cloud segmentation inference device for a large language model, including:

[0024] A segmentation module, configured to segment the large language model into a client model and a server model based on a preset segmentation criterion;

[0025] A first inference and encoding module, configured to perform inference on the data to be processed through the client model to obtain an intermediate result; and compress the intermediate result through an adaptive encoder to obtain a compressed intermediate result, wherein the client model and the adaptive encoder are deployed on an edge device;

[0026] A decoding and second inference module, configured to reconstruct the compressed intermediate result by using an adaptive decoder to obtain a reconstructed intermediate result; and perform inference on the reconstructed intermediate result through the server model to obtain a new token, wherein the adaptive decoder and the server model are deployed on a cloud server;

[0027] A repetition module, configured to obtain new data to be processed based on the new token, and repeat the first inference and encoding module to the decoding and second inference module until a preset condition is reached to obtain a final prediction result.

[0028] According to a third aspect of the present invention, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor runs the computer program to implement the method as described above.

[0029] According to a fourth aspect of the present invention, the present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method as described above is implemented.

[0030] According to a fifth aspect of the present invention, the present invention also provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, the method as described above is implemented.

[0031] Through one or more of the above embodiments in the present invention, at least the following technical effects can be achieved: The large language model is divided into a client model and a server model, and the client model and the server model are jointly used for collaborative inference. During the inference process, data compression is achieved through an adaptive encoder, and data reconstruction is achieved through an adaptive decoder, thereby reducing the communication overhead between the edge device and the cloud server and enhancing data privacy protection. The method provided by the present invention can adapt to different network environments and computing resources, has good flexibility and scalability, and is applicable to various application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The technical solutions and other beneficial effects of the present invention will become apparent through a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings.

[0033] Figure 1 FIG. is one of the flowcharts of the communication edge-cloud segmentation inference method for large language models provided by the embodiments of the present invention;

[0034] Figure 2 FIG. is another flowchart of the communication edge-cloud segmentation inference method for large language models provided by the embodiments of the present invention;

[0035] Figure 3 FIG. is the performance comparison diagram provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present invention.

[0037] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the front and rear associated objects without special instructions.

[0038] The following introduces a communication edge-cloud segmentation inference method, device, equipment, product and medium for large language models provided by the present invention in conjunction with the accompanying drawings.

[0039] As Figure 1As shown in the figure, an embodiment of the present invention provides a communication edge-cloud segmentation inference method for large language models. The method includes the following steps:

[0040] Segmentation step S101: Segment the large language model into a client model and a server model based on a preset segmentation criterion.

[0041] In this step, the preset segmentation criterion includes at least one of the network structure complexity of the large language model, the computing power of the edge device, the computing power of the cloud server, or the data transmission efficiency between the edge device and the cloud server. In some embodiments, segmenting the large language model into a client model and a server model can be based on the network structure complexity of the large language model. Among them, the network structure complexity can be evaluated according to the number of parameters, the number of layers, and FLOPs of the network structure, etc. Schematically, the parameters of the large language model with a network structure complexity exceeding the threshold are used as the server model, and the remaining parameters of the large language model are used as the client model. In other embodiments of the present application, the large language model can also be segmented into multiple client models and a server model, etc., which is not limited herein.

[0042] By segmenting the large language model into a client model and a server model, the client model deployed on the edge device can independently process the user data obtained on the edge device, and the server model deployed on the cloud server can complete more complex computing tasks.

[0043] First inference and encoding step S102: Infer the data to be processed through the client model to obtain an intermediate result , where represents the mapping relationship between the input data and its predicted value when the model parameters of the client model are ; and compress the intermediate result through an adaptive encoder to obtain a compressed intermediate result . Among them, the adaptive encoder uses a multi-layer perceptron to reduce the dimension of the intermediate result to with dimensional features, while the adaptive decoder reconstructs it into the original . The adaptive decoder optimizes the dynamic feature expression by integrating multiple groups of encoders and decoders with different dimensions .

[0044] Among them, the client model and the adaptive encoder are deployed on the edge device.

[0045] In this step, the client model infers the data to be processed to obtain an intermediate result, including:

[0046] Perform forward propagation on the data to be processed through the client model to the cut layer to obtain activation value data, where the cut layer is the last network layer of the client model.

[0047] Evaluate the importance of the activation value data and determine the intermediate result according to the evaluation result. Specifically, by calculating the average score of each activation value data and the previous token corresponding to each activation value in the attention module of the adaptive encoder, the obtained score represents the attention correlation degree of the token corresponding to the activation value. The higher the score, the higher the attention correlation degree, and vice versa, the lower the score, the lower the attention correlation degree.

[0048] Identify the tokens with scores exceeding the score threshold as intermediate results according to the aforementioned scores and the score threshold.

[0049] In addition, the edge device sends the compressed intermediate result to the cloud server.

[0050] Decode and the second inference step S103, use the adaptive decoder to reconstruct the compressed intermediate result to obtain the reconstructed intermediate result ; and infer the reconstructed intermediate result through the server model to obtain new tokens , where are the model parameters of the server model.

[0051] Among them, the adaptive decoder and the server model are deployed on the cloud server.

[0052] It can be understood that the cloud server sends the new tokens to the edge device after obtaining them.

[0053] S104, Obtain new data to be processed based on the new tokens, and repeat the first inference and encoding step to the decoding and second inference step until a preset condition is reached to obtain the final prediction result. Wherein, the preset condition includes one of the new tokens being a complete sentence, a preset token number threshold, and a preset end symbol.

[0054] In this step, the edge device obtains new data to be processed based on the new tokens, including:

[0055] Concatenate the new tokens with the data to be processed to obtain the new data to be processed, that is .

[0056] The communication edge-cloud segmentation inference method for large language models provided by the embodiments of the present invention divides the large language model into two parts: a client model and a server model, and jointly performs collaborative inference with the client model and the server model. During the inference process, data compression is achieved through an adaptive encoder, and data reconstruction is achieved through an adaptive decoder, thereby reducing the communication overhead between the edge device and the cloud server and enhancing data privacy protection. In addition, the edge device utilizes the powerful computing power of the server side to alleviate the limitation of the computing resources of the edge device. The method provided by the present invention can adapt to different network environments and computing resources, has good flexibility and scalability, and is applicable to a variety of application scenarios.

[0057] In some embodiments of the present invention, before the first inference and encoding step, the method further includes:

[0058] Obtain initial data, and perform word segmentation or feature vectorization on the initial data to obtain the data to be processed.

[0059] Specifically, the initial data is obtained through an edge device. When the initial data is text, word segmentation preprocessing is performed on the initial data to obtain the data to be processed; when the initial data is an image, feature vectorization processing (i.e., image embedding) is performed on the initial data to obtain the data to be processed.

[0060] In some embodiments of the present invention, the present invention conducts experiments based on the large language model Meta-Llama-3-8B-Instruct model. The Meta-Llama-3-8B-Instruct is an instruction fine-tuning version in the Meta Llama 3 series, which has 8 billion parameters and 32 layers of Transformer encoders, and it can better follow instructions and generate high-quality text.

[0061] In addition, 1 edge device is deployed. The client model on this edge device includes the Embedding layer of the large language model and the first 8 Transformer layers (i.e., one-fourth of the total number of layers). The cloud server is responsible for running the server model, which includes the remaining 24 Transformer layers (i.e., three-fourths of the total number of layers). The computing power of the edge device is 35.6 TFLOTS, and the computing power of the cloud server is set to 284.8 TFLOTS. The communication rate between the edge device and the cloud server is uniformly set to 600 Mbps. To ensure model accuracy, the torch.float32 data type is used to load the client model and the server model. The final evaluation metrics of the experiment include the accuracy on each dataset and the latency during the inference process (including computing latency and communication latency).

[0062] Baseline: To evaluate the performance, the latency of the segmented client model and server model is compared with the direct inference of the complete large language model before segmentation:

[0063] The total inference latency of the client model and the server model, which includes the computing latency of the edge device, the communication latency between the edge device and the cloud server, and the computing latency of the cloud server.

[0064] The inference latency of the large language model, which refers to the total inference time of directly running the large language model on the cloud server.

[0065] As Figure 3 can be seen, on the HumanEval dataset, the total inference latency of the client model and the server model is reduced by about 64.6% compared with the large language model, and on the SQuAD dataset, the total inference latency of the client model and the server model is reduced by about 53.6% compared with the large language model. The long inference time of the large language model is due to its huge parameter scale and computational complexity, which requires all calculations to be completed on a single device, resulting in a slow inference speed. The method provided by the present invention divides the model into two parts: a client and a server, and distributes the computing tasks to the edge device and the cloud server, thereby reducing the computing burden on a single device and improving the inference speed. By comparison, it can be seen that the method provided by the present invention can effectively utilize the computing resources of the edge device and reduce the inference burden on the server side. In addition, it can also be seen that the method provided by the present invention can significantly reduce the inference time on different datasets, indicating that the method provided by the present invention has good generalization ability.

[0066] In summary, the method provided by the present invention has the following advantages:

[0067] 1) By segmenting the large language model and using an adaptive encoder / decoder to reduce the computing and communication burdens. The large language model is divided into a client model and a server model. Among them, the edge device is responsible for the forward propagation and adaptive compression of the client model, while the cloud server is responsible for the forward propagation of the server model and Token generation. The use of the adaptive encoder / decoder significantly reduces the size of the transmitted data because it only transmits the compressed intermediate results instead of the original data. In addition, the method provided by the present invention takes advantage of edge computing by performing part of the calculations on the edge device, thereby reducing the computing pressure on the cloud server.

[0068] 2) Reduce the model size and communication overhead through an adaptive encoder / decoder without significantly affecting the model performance. The adaptive encoder / decoder enables adaptive compression, thereby reducing the amount of data transmitted and improving the inference efficiency. In addition, the method provided by the present invention reduces the amount of data that the edge device needs to process, lowers the computational resource requirements for the edge device, and enables the deployment of LLM in resource-constrained environments. In this way, different computational and resource-constrained scenarios can be adapted while maintaining the model performance.

[0069] 3) Allow the adjustment of the model splitting point and the parameters of the adaptive encoder / decoder according to the resource heterogeneity of the edge device and the cloud server. By carefully selecting the splitting point and adjusting the parameters of the encoder / decoder, the computational workload between the edge device and the cloud server can be balanced, and the amount of data that the edge device needs to transmit can be reduced. This method enables resource-constrained devices to participate in the model inference while ensuring the efficiency and effectiveness of the model inference.

[0070] An embodiment of the present invention also provides a communication edge-cloud segmentation inference device for a large language model, and the device includes:

[0071] A splitting module, configured to split the large language model into a client model and a server model based on a preset partitioning criterion;

[0072] A first inference and encoding module, configured to perform inference on the data to be processed through the client model to obtain an intermediate result; and compress the intermediate result through an adaptive encoder to obtain a compressed intermediate result, wherein the client model and the adaptive encoder are deployed on an edge device;

[0073] A decoding and second inference module, configured to reconstruct the compressed intermediate result by using an adaptive decoder to obtain a reconstructed intermediate result; and perform inference on the reconstructed intermediate result through the server model to obtain a new token, wherein the adaptive decoder and the server model are deployed on a cloud server;

[0074] A repeating module, configured to obtain new data to be processed based on the new token, and repeat the first inference and encoding module to the decoding and second inference module until a preset condition is reached to obtain a final prediction result.

[0075] A communication edge-cloud segmentation inference device for a large language model provided by an embodiment of the present invention corresponds to the communication edge-cloud segmentation inference method for a large language model provided by any of the above embodiments, and details are not described herein again.

[0076] Based on any of the above embodiments, another embodiment of the present invention further provides an electronic device, which may include: a memory, a processor, and a computer program stored on the memory and executable on the processor. Wherein, the processor may call the logical instructions in the memory to execute the above method.

[0077] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0078] On the other hand, an embodiment of the present invention also provides a storage medium, on which multiple instructions are stored, and the instructions are suitable for being loaded by a processor to execute the communication edge-cloud segmentation inference method for large language models provided in the above various embodiments.

[0079] On the other hand, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0080] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0081] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0082] In summary, although the present invention has been disclosed above with preferred embodiments, the above preferred embodiments are not intended to limit the present invention. Those of ordinary skill in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention is subject to the scope defined by the claims.

Claims

1. A communication edge-cloud segmentation inference method for large language models, characterized in that, Including: A splitting step of splitting a large language model into a client model and a server model based on a preset partitioning criterion; A first inference and encoding step of performing inference on data to be processed through the client model to obtain an intermediate result; and compressing the intermediate result through an adaptive encoder to obtain a compressed intermediate result, wherein the client model and the adaptive encoder are deployed on an edge device; A decoding and second inference step of reconstructing the compressed intermediate result by using an adaptive decoder to obtain a reconstructed intermediate result; and performing inference on the reconstructed intermediate result through the server model to obtain a new token, wherein the adaptive decoder and the server model are deployed on a cloud server; Concatenating the new token with the data to be processed to obtain new data to be processed, and repeating the first inference and encoding step to the decoding and second inference step until a preset condition is reached to obtain a final prediction result; Wherein, the performing inference on the data to be processed through the client model to obtain an intermediate result includes: Performing forward propagation on the data to be processed through the client model to a cut layer to obtain activation value data; Performing importance evaluation on the activation value data, and determining the intermediate result according to the evaluation result; The performing importance evaluation on the activation value data and determining the intermediate result according to the evaluation result includes: Calculating the average score of each activation value data and the previous token corresponding to each activation value in the attention module of the adaptive encoder, and determining the intermediate result according to the score and a score threshold.

2. The method according to claim 1, wherein The preset partitioning criterion includes at least one of the network structure complexity of the large language model, the computing power of the edge device, the computing power of the cloud server, or the data transmission efficiency between the edge device and the cloud server.

3. The method according to claim 1, wherein Before the first inference and encoding step, the method further includes: Obtaining initial data, and performing word segmentation or feature vectorization processing on the initial data to obtain the data to be processed.

4. The method according to claim 1, characterized in that, The preset condition includes that the new token is a complete sentence.

5. A communication edge-cloud segmentation inference device for large language models, characterized in that, Including: A splitting module for splitting a large language model into a client model and a server model based on a preset partitioning criterion; A first inference and encoding module for performing inference on data to be processed through the client model to obtain an intermediate result; and compressing the intermediate result through an adaptive encoder to obtain a compressed intermediate result, wherein the client model and the adaptive encoder are deployed on an edge device; A decoding and second inference module for reconstructing the compressed intermediate result by using an adaptive decoder to obtain a reconstructed intermediate result; and performing inference on the reconstructed intermediate result through the server model to obtain a new token, wherein the adaptive decoder and the server model are deployed on a cloud server; A repeating module for concatenating the new token with the data to be processed to obtain new data to be processed, and repeating the first inference and encoding module to the decoding and second inference module until a preset condition is reached to obtain a final prediction result; Among them, the first inference and encoding module is specifically configured to: Perform forward propagation on the data to be processed through the client model to the cut layer to obtain activation value data; Evaluate the importance of the activation value data, and determine the intermediate result according to the evaluation result; The evaluating the importance of the activation value data and determining the intermediate result according to the evaluation result includes: Calculating the average score of each activation value data and the previous token corresponding to each activation value in the attention module of the adaptive encoder, and determining the intermediate result according to the score and the score threshold.

6. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor runs the computer program to implement the method according to any one of claims 1 to 4.

7. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 4 is implemented.

8. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Neural network collaborative reasoning method for multi-access edge computing system

    CN114723057A

  • Cloud edge collaborative reasoning method based on auto-encoder and intelligent agent

    CN119886219A