End-cloud collaborative deployment method and device for multi-modal large model, medium and product

By implementing the end-cloud collaborative deployment method of multimodal large models between the end-side and the cloud side, dynamically adjusting the allocation of computing tasks and optimizing the transmission path, the problem of low inference efficiency of multimodal large models on the end-side devices is solved, and efficient and highly adaptable end-cloud collaborative inference is achieved.

CN120050188AActive Publication Date: 2025-05-27BEIJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510190017.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-27
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently deploy and infer multimodal large models on end-side devices, especially under limited computing power, memory and power consumption, and it is difficult to dynamically schedule computing tasks to adapt to data flows of different modes.

Method used

By implementing the end-cloud collaborative deployment method of multimodal large models between the end-side and the cloud side, the allocation of computing tasks is dynamically adjusted, the multimodal data is encoded using the optimized encoding model, and the transmission path is optimized through a dynamic scheduling algorithm, and the computing and data transmission are optimized in combination with computing network fusion technology.

Benefits of technology

The efficiency and adaptability of multimodal large models on the end-side inference are improved, memory usage and power consumption are optimized, and the utilization rate of computing resources is improved, and intelligent scheduling and efficient inference of end-cloud collaborative computing are realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050188A_ABST
    Figure CN120050188A_ABST
Patent Text Reader

Abstract

The invention discloses an end-cloud collaborative deployment method and device for a multi-modal large model, a medium and a product, and relates to the field of multi-modal large model deployment, and the method comprises the steps: a cloud side computer obtains a to-be-deployed multi-modal large model, determines an optimized coding model and a corresponding segmentation candidate point according to the to-be-deployed multi-modal large model, and transmits the optimized coding model and the corresponding segmentation candidate point to an end side computer; the end side computer obtains multi-modal data, the optimized coding model is used for coding the multi-modal data, and intermediate data and segmentation point position information are obtained; and the end-side computer compresses and packs the intermediate data and the segmentation point position information, and sends the compressed and packed intermediate data and the segmentation point position information to the cloud-side computer through a transmission path, so that the intermediate data is calculated and processed by using a processing model to obtain a calculation result, and the calculation result is sent to the end-side computer, decoded by using a decoding model and converted into an output format. And obtaining the processed multi-modal data. The method can dynamically adjust the distribution of the calculation tasks, and improves the reasoning efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of multimodal large model deployment, and particularly to a method, device, medium and product for end-cloud collaborative deployment of a multimodal large model. Background Art

[0002] In recent years, with the rapid development of technologies in the field of artificial intelligence, multimodal large models have become a research hotspot due to their excellent performance in multiple fields such as computer vision, natural language processing, and speech recognition. However, as the data scale of deep learning large models continues to grow, the number of model parameters has increased to trillions, and the computing power demand is constantly rising. How to reasonably allocate computing tasks between the edge and the cloud for efficient end-cloud hybrid deployment has become an important research direction.

[0003] Multimodal large models usually involve various types of data, including videos, pictures, voices, texts, 3D modeling information, data tables, etc. In deep learning algorithms, there are also diverse characteristics in model structures. The emergence of algorithms such as multimodality, MoE, quantization compression, and special customized operators has brought uncertainties to the model inference process. At present, multimodal large models are widely used. In application scenarios such as intelligent driving and smart cities, edge devices can usually access data earlier. However, most edge devices do not have the ability to independently complete large model computations and still rely on the end-cloud collaborative computing mode.

[0004] In the field of multimodal deep learning large models, some representative works include DALL-E, Stable Diffusion, CLIP, and Flamingo. These models have different technical characteristics and application scenarios in multimodal data processing. DALL-E and Stable Diffusion mainly focus on text-to-image generation and achieve high-quality image generation based on diffusion models. CLIP maps text and image data into the same shared vector space through contrastive learning and is applicable to tasks such as image retrieval, classification, and semantic matching. Flamingo adopts a cross-modal attention mechanism to optimize multimodal understanding tasks by fusing text and visual information and is applicable to applications such as dialogue systems and cross-modal content generation. These works have jointly promoted the development of multimodal large models, and the multimodal data processing scenario has also provided opportunities for efficient end-cloud collaborative inference.

[0005] With the joint construction of computing power and network, the network transmission bottleneck is gradually reduced, and the distance for computing power and data transmission becomes shorter. Through computing-aware network scheduling, computing network convergence realizes intelligent scheduling and optimization of computing power. Edge-cloud collaborative computing is a method that utilizes the powerful computing capabilities of the cloud and the low-latency characteristics of edge devices to achieve efficient computing task allocation. Currently, model distributed inference can split large-scale deep learning models, and deep learning frameworks (such as TensorFlow and PyTorch) support inference or training on multiple computing nodes to improve computing efficiency. Technologies such as model compression and knowledge distillation can transfer the knowledge of large models to lightweight models, enabling them to perform inference on edge devices, thereby reducing computing overhead.

[0006] Currently, some lightweight versions of large models have been proposed. For example, the 4-bit version of the LLaMA 7B model file size is 3.8GB (the full precision is 28GB), and the 4-bit version of the edge multi-modal model MiniCPM requires more than 7GB of memory (the non-quantized model requires about 17GB), but it is still a heavy burden on edge devices. Moreover, the decrease in accuracy will reduce the model performance, resulting in MiniCPM having good performance only in a few tasks (such as Optical Character Recognition (OCR)). Therefore, in the scenario of edge inference with limited power, computing power, and storage, considering factors such as the inference speed, memory occupancy, and power consumption of the model, how to improve the inference speed, optimize memory occupancy, and power consumption under limited resource conditions remains a key issue. Currently, some optimization solutions are only applicable to specific tasks or specific computing hardware and are difficult to be generally adapted to different computing environments.

[0007] In addition, inference tasks often involve heterogeneous computing resources, and existing solutions are difficult to efficiently dynamically schedule computing tasks between different devices, resulting in poor resource utilization. Collaboratively deploying large model applications on cloud and edge devices, and optimizing the resource allocation method based on the computing power network can make the allocation of cloud computing tasks more dynamic and intelligent, greatly improving the computing resource utilization rate and achieving goals such as intelligent scheduling of edge devices and cloud computing and computing network convergence.

[0008] Regarding the above problems of dynamic optimization of edge-cloud collaborative computing, adaptability to heterogeneous computing environments, and efficient inference of multi-modal large models, they can be achieved through current distributed deep learning optimization methods, edge-cloud collaborative optimization methods, distributed system cost evaluation methods, multi-modal inference acceleration methods, etc.

[0009] The optimization method for distributed execution of deep learning tasks is based on the method of computational graph optimization. It optimizes the distributed execution of deep learning tasks through device grouping and operator splitting schemes, selects the optimal splitting scheme based on the overhead model, realizes automatic optimization, and improves the execution efficiency.

[0010] The edge-cloud collaborative optimization method based on deep reinforcement learning uses deep reinforcement learning to optimize the inference scheme in edge-cloud collaboration, including early exit points, segmentation points, quantization encoding, etc., and dynamically adjusts the inference scheme according to bandwidth, latency, energy consumption, and accuracy to improve the inference efficiency.

[0011] The cost evaluation method for distributed systems adopts the cost evaluation method for distributed systems, calculates the costs at the operator level, device level, and cluster level, and evaluates the performance of distributed parallel schemes.

[0012] The method for accelerating the inference of deep learning models based on the collaboration between edge servers and mobile devices uses model segmentation and model simplification to allocate computing tasks between edge servers and mobile devices to reduce latency, estimates the running latency of network layers through a regression model, and searches for optimal exit points and segmentation points.

[0013] The high-performance multi-modal large model inference system adopts multi-modal accelerated inference units, combines search units, cache units, and databases to optimize multi-modal inference efficiency. It provides the combined inference ability of multiple multi-modal components, supports distributed deployment, and improves throughput.

[0014] However, the existing related technologies still have the following problems: (1) The methods mainly focus on the computational optimization of single deep learning tasks, do not cover the special requirements of multi-modal large models, are difficult to adapt to the dynamic changes of multi-modal inputs, and cannot perform adaptive optimization for data streams of different modalities. For example, the related technologies optimize static computational graphs, do not support the dynamic computational adaptation of multi-modal large models, and cannot solve the dynamic input-output patterns of multi-modal large models. (2) There is no optimization for the integration of computing and networking, and it may still be limited by the static allocation of network bandwidth and computing load. The related technologies focus on the cost evaluation of distributed computing rather than the actual end-cloud collaborative inference optimization, and there are also optimizations based on multi-modal caching and databases, but do not involve the collaborative optimization of edge-side computing and cloud computing. (3) In addition, the related technologies are still based on traditional distributed methods, do not involve the integration of computing and networking, only optimize computing resources, and ignore the impact of network transmission on inference efficiency. Summary of the Invention

[0015] The purpose of this application is to provide a method, device, medium, and product for end-cloud collaborative deployment of multi-modal large models, which can dynamically adjust the allocation of computing tasks and improve inference efficiency.

[0016] To achieve the above purpose, this application provides the following solutions:

[0017] In a first aspect, the present application provides a method for collaborative edge-cloud deployment of a multimodal large model, which is implemented by an edge-side computer and a cloud-side computer; the method for collaborative edge-cloud deployment of the multimodal large model includes:

[0018] The cloud-side computer obtains the multimodal large model to be deployed, determines an optimized encoding model and splitting candidate points of the optimized encoding model according to the multimodal large model to be deployed, and sends them to the edge-side computer; the optimized encoding model is obtained by optimizing the encoding model; the encoding model is initially split from the multimodal large model to be deployed according to the rules of information encoding - information processing - information decoding; the optimization process includes quantization processing, pruning processing, and fusion processing;

[0019] The edge-side computer obtains multimodal data, and encodes the multimodal data by using the optimized encoding model to obtain intermediate data and splitting point position information;

[0020] The edge-side computer compresses and packages the intermediate data and the splitting point position information, and sends them to the cloud-side computer through a transmission path; the transmission path is determined by using a dynamic scheduling algorithm according to the real-time computing power and network resource status;

[0021] The cloud-side computer calculates and processes the intermediate data by using a processing model to obtain a calculation result, and sends the calculation result to the edge-side computer; the processing model is initially split from the multimodal large model to be deployed according to the rules of information encoding - information processing - information decoding;

[0022] The edge-side computer decodes the calculation result by using a decoding model and converts it into an output format to obtain processed multimodal data; the decoding model is initially split from the multimodal large model to be deployed according to the rules of information encoding - information processing - information decoding.

[0023] Optionally, determining an optimized encoding model and splitting candidate points of the optimized encoding model according to the multimodal large model to be deployed and sending them to the edge-side computer specifically includes:

[0024] Initially split the multimodal large model to be deployed according to the rules of information encoding - information processing - information decoding to obtain an encoding model, a processing model, and a decoding model;

[0025] Perform quantization processing, pruning processing, and fusion processing on the encoding model to obtain an optimized encoding model;

[0026] Split the optimized encoding model to obtain a set of encoding model splitting schemes; the set of encoding model splitting schemes includes multiple candidate encoding models and splitting candidate points of the optimized encoding model.

[0027] Judge whether the set of encoding model splitting schemes is empty.

[0028] If it is not empty, take out the candidate encoding model from the set of encoding model splitting schemes, and use the cost model to evaluate the candidate encoding model to judge whether the candidate encoding model meets the requirements of the inference computing power, latency, and data transmission speed of the edge computer.

[0029] If it meets the requirements, send the optimized encoding model and the splitting candidate points of the optimized encoding model to the edge computer.

[0030] Optionally, after sending the optimized encoding model and the splitting candidate points of the optimized encoding model to the edge computer, it further includes:

[0031] Based on the set of encoding model splitting schemes, according to the computing power model, perform a splitting on the decoding model that is symmetric to the optimized encoding model to obtain candidate decoding models and decoding model splitting candidate points.

[0032] Optionally, obtain multimodal data, and use the optimized encoding model to perform encoding processing on the multimodal data to obtain intermediate data and splitting point position information, specifically including:

[0033] Obtain the multimodal data, and judge whether there is a qualified splitting scheme for the candidate encoding model that meets the requirements of the inference computing power, latency, and data transmission speed of the edge computer.

[0034] If there is, use the optimized encoding model to perform encoding processing on the multimodal data to obtain encoded multimodal data.

[0035] Use the computing power of the edge computer to map the encoded multimodal data into the same multimodal encoding space to obtain intermediate data, and mark the splitting point position information; the splitting point position information is the splitting point corresponding to the qualified splitting scheme.

[0036] Optionally, use the processing model to perform computing processing on the intermediate data to obtain a computing result, and send the computing result to the edge computer, specifically including:

[0037] Judge whether the encoding task is completed according to the intermediate data to obtain a first judgment result.

[0038] If the first judgment result is no, execute the encoding task of the multimodal data that the edge computer has not completed.

[0039] If the first judgment result is yes, then judge whether the intermediate data is single data to obtain a second judgment result;

[0040] If the second judgment result is yes, perform batch processing and unified calculation on the intermediate data to obtain a calculation result, and send the calculation result to the edge computer;

[0041] If the second judgment result is no, perform dynamic hybrid scheduling calculation on the intermediate data to obtain a calculation result, and send the calculation result to the edge computer.

[0042] Optionally, it further includes:

[0043] Send the candidate decoding model and the decoding model segmentation candidate points to the edge computer.

[0044] Optionally, it further includes:

[0045] Establish a dynamic feedback mechanism for computing power and network resources, and the dynamic feedback mechanism is used to update and improve the cost model.

[0046] In a second aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to implement the multi-modal large model edge-cloud collaborative deployment method described in any one of the above.

[0047] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the multi-modal large model edge-cloud collaborative deployment method described in any one of the above.

[0048] In a fourth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the multi-modal large model edge-cloud collaborative deployment method described in any one of the above.

[0049] According to the specific embodiments provided by the present application, the present application has the following technical effects:

[0050] The present application provides a method, device, medium and product for end-cloud collaborative deployment of a multi-modal large model. For a target deployment environment, the present application performs optimizations on the end side and the cloud side respectively. For the computing optimization on the end side, the present application mainly uses dynamic compilation optimization technology to complete the optimized calculation by segmentation, dynamically adapts to the computing power on the end side, and improves the inference speed. Model compression and quantization optimization are adopted: by means of model pruning, quantization and knowledge distillation, etc., the computational complexity and memory occupancy of the model are reduced to adapt to the limited computing resources on the end side. For the computing on the cloud side, the present application uses a distributed parallel computing framework to efficiently schedule tasks and realizes the full utilization of cloud computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0052] Figure 1 It is a schematic flowchart of a method for end-cloud collaborative deployment of a multi-modal large model provided by an embodiment of the present application;

[0053] Figure 2 It is a flowchart of end-cloud collaborative deployment in practical applications provided by the present application;

[0054] Figure 3 It is an example diagram of symmetric cutting of the model provided by the present application;

[0055] Figure 4 It is a flowchart of step ① of the present application;

[0056] Figure 5 It is a flowchart of the execution of end-side encoding in step ② of the present application;

[0057] Figure 6 It is a flowchart of the execution of cloud computing in step ④ of the present application;

[0058] Figure 7 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0059] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0060] To make the above objects, features, and advantages of the present application more apparent and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0061] The present application breaks through the limitations of traditional methods in aspects such as static computational graph optimization and fixed computing resource allocation, enabling multi-modal large models to be efficiently co-deployed between the edge and the cloud, and providing a more intelligent and dynamic optimization strategy for efficient inference.

[0062] In multi-modal deep learning large models, dynamic input and output require the model to adapt to different data scenarios. The model needs to adjust the allocation of computing resources in real time to adapt to the processing requirements of different modalities, which increases the complexity of system design and resource management. In the computing power network, the network transmission capacity between nodes may be insufficient to support the low-latency transmission requirements of massive data. Computing-network convergence requires effective identification and perception of computing and network resources, but it is currently difficult to achieve comprehensive resource perception and dynamic scheduling. At the same time, for the application data itself, if the original data is directly transmitted to the cloud, it may involve problems such as large data volume and data privacy leakage.

[0063] Multi-modal large models usually need to process various different input and output data types, such as generating images from text (input text) or generating text from images (input image). To achieve this, the model usually adopts an encoder-decoder architecture. In the encoding stage, the model encodes the input multi-modal information through the encoder and uniformly maps it to a representation space that can express multi-modal data simultaneously. After the large language model understands and processes the language, the decoder decodes the information in this representation space to generate the required output, such as generating text or images.

[0064] The present application proposes a method for edge-cloud collaborative deployment of multi-modal large models to achieve efficient collaborative deployment of the model between the edge and the cloud. By dividing the model into three stages: encoding, processing, and decoding, and transmitting intermediate representation data between the edge and the cloud, both the data transmission volume is reduced and user privacy is protected. This method enables multi-modal large models to dynamically adjust the allocation of computing tasks according to different input and output requirements, improving the inference efficiency. At the same time, combined with computing-network convergence technology, it intelligently optimizes the computing and data transmission paths to ensure efficient inference in the edge-cloud collaborative environment. Compared with the prior art, the present application provides a more intelligent and dynamic optimization strategy, which is suitable for the efficient deployment of multi-modal large models.

[0065] The edge-side computer is an embedded system with an Nvidia Jetson Orin NX 8GB series processor, running a lightweight Linux operating system, and having a peak AI computing power of approximately 117 TOPS. The memory configuration is about 8GB (after low-precision quantization, the edge-side model data occupies 6GB - 7GB of memory), and it is equipped with a dedicated GPU accelerator. The network interface supports WiFi and 4G / 5G communications to ensure that the data upload and download bandwidth meet the requirements.

[0066] The cloud-side computer uses a high-performance CPU and GPU (for example, a single Nvidia A800 80GB PCIE with a peak AI computing power of 1248 TOPS) cluster, and is equipped with high-speed network switching equipment. The cloud-side computer has sufficient computing power support, is equipped with a deployed distributed task scheduling system, and supports large-scale data parallel computing and resource dynamic scheduling.

[0067] Network environment: The edge side and the cloud side are connected through a broadband Internet, and the supported bandwidth ranges from 10 Mbps to 1 Gbps. The network transmission protocol (such as TCP / IP) is used, and the data transmission delay is controlled within the range of 50 - 200 milliseconds under dynamic network conditions.

[0068] In an exemplary embodiment, as Figures 1 - 3 shown, a method for edge-cloud collaborative deployment of a multi-modal large model is provided. The method for edge-cloud collaborative deployment of the multi-modal large model is implemented by an edge-side computer and a cloud-side computer; the method for edge-cloud collaborative deployment of the multi-modal large model includes the following steps:

[0069] S1: The cloud-side computer obtains the multi-modal large model to be deployed, and determines the optimized encoding model and the splitting candidate points of the optimized encoding model according to the multi-modal large model to be deployed, and sends them to the edge-side computer; the optimized encoding model is obtained by optimizing the encoding model; the encoding model is obtained by initially splitting the multi-modal large model to be deployed according to the rules of information encoding - information processing - information decoding; the optimization process includes quantization processing, pruning processing, and fusion processing.

[0070] As an optional implementation manner, determining the optimized encoding model and the splitting candidate points of the optimized encoding model according to the multi-modal large model to be deployed and sending them to the edge-side computer specifically includes:

[0071] Initially split the multi-modal large model to be deployed according to the rules of information encoding - information processing - information decoding to obtain an encoding model, a processing model, and a decoding model.

[0072] Perform quantization processing, pruning processing, and fusion processing on the encoding model to obtain the optimized encoding model.

[0073] Partition the optimized encoding model to obtain a set of encoding model partitioning schemes; the set of encoding model partitioning schemes includes multiple candidate encoding models and partitioning candidate points of the optimized encoding model.

[0074] Determine whether the set of encoding model partitioning schemes is empty.

[0075] If it is not empty, take out the candidate encoding models from the set of encoding model partitioning schemes, and use a cost model to evaluate the candidate encoding models to determine whether the candidate encoding models meet the requirements of the inference computing power, latency, and data transmission speed of the edge computer.

[0076] If they meet the requirements, send the optimized encoding model and the partitioning candidate points of the optimized encoding model to the edge computer.

[0077] It further includes:

[0078] Based on the set of encoding model partitioning schemes, according to the computing power model, perform a partitioning on the decoding model that is symmetric to the optimized encoding model to obtain candidate decoding models and decoding model partitioning candidate points.

[0079] In practical applications, as Figure 4 shown, Step ①: Model partitioning.

[0080] Execution entity: Cloud computer.

[0081] Step description: The system obtains a pre-constructed computational graph and related model data, providing a basic framework and parameter information for subsequent processing. Perform optimization method calculations in the cloud and partition the model according to the three-stage execution method.

[0082] Step 1.1: Model input and initial partitioning candidate generation.

[0083] Receive a multimodal large model, the computational graph of the multimodal large model pre-partitioned by the system, which defines the connection relationships and data flows between modules. At the same time, load model parameter data (weights, biases, etc.), and partition the multimodal large model following the rules of information encoding - information processing - information decoding.

[0084] Perform corresponding optimizations such as quantization, pruning, and fusion on the encoding model; partition the optimized model and put it into the model partitioning candidate point set (set of encoding model partitioning schemes). For a multimodal large model, in the encoding / decoding stage, information is gradually encoded into a unified encoding space identifier / decoded from it. Therefore, even if a clear encoding or decoding model can be divided in the model structure, a finer-grained partitioning can be performed to provide more candidate partitions for the system to dynamically schedule.

[0085] Step 1.2: Traverse the set of encoding model segmentation schemes.

[0086] Judge whether the set of encoding model segmentation schemes is empty. If it is not empty, select a candidate encoding model from the set and execute Step 1.3.

[0087] If it is empty, do not execute the encoding model segmentation step and continue to execute Step 1.4.

[0088] Step 1.3: Traverse the evaluation and optimization of the encoding model segmentation scheme.

[0089] Calculate and evaluate the candidate encoding model through the cost model, and judge whether the encoding model segmentation scheme meets the requirements of the edge-side inference computing power, latency, and data network transmission (upload). Among them, for the decoding model, the network data transmission bottleneck caused by the edge-side data upload bandwidth needs to be considered emphatically.

[0090] If it can be satisfied, send the segmentation points of the candidate encoding model and the optimized encoding model to the edge side, prepare for edge-side encoding inference, and continue to execute Step 1.4.

[0091] If it cannot be satisfied, return to Step 1.2 and re-select the candidate decoding model segmentation from the subsequent set.

[0092] Step 1.4: Symmetric segmentation of the edge-side decoding model.

[0093] Take out the candidate encoding model from the set of encoding model segmentation schemes, and according to the computing power model, perform symmetric segmentation on the decoding model as the encoding model. It can be assumed that the edge-side encoding and decoding computing powers are the same.

[0094] If the set of encoding model segmentation schemes is empty, do not execute the decoding model segmentation step and directly execute Step 1.5.

[0095] Step 1.5: Cloud model optimization and saving.

[0096] Perform corresponding optimization and saving on the remaining model data not allocated to the edge side in the cloud.

[0097] Keep a backup of the optimized edge-side model in the cloud for dynamic scheduling during system operation to supplement the insufficient edge-side computing power.

[0098] S2: The edge-side computer obtains multi-modal data and uses the optimized encoding model to encode the multi-modal data to obtain intermediate data and segmentation point location information.

[0099] As an optional implementation manner, obtaining multi-modal data and using the optimized encoding model to encode the multi-modal data to obtain intermediate data and segmentation point location information specifically includes:

[0100] Obtain the multi-modal data, and determine whether there is a qualified segmentation scheme for the candidate encoding model that meets the requirements of the inference computing power, latency, and data transmission speed of the edge computer.

[0101] If there is, use the optimized encoding model to encode the multi-modal data to obtain the encoded multi-modal data.

[0102] Utilize the computing power of the edge computer to map the encoded multi-modal data into the same multi-modal encoding space to obtain intermediate data, and mark the position information of the segmentation points; the position information of the segmentation points is the segmentation points corresponding to the qualified segmentation scheme.

[0103] In practical applications, such as Figure 5 shown, Step ②: Edge-side encoding calculation.

[0104] Execution entities: Edge computer, network switch.

[0105] Step description: Perform the encoding calculation of the multi-modal data in the first stage of the segmentation model. Through data preprocessing, convert each modal data into vector representation, and embed / encode it into a unified multi-modal representation space, and execute the segmented model on the edge side to generate intermediate representation data.

[0106] Step 2.1: Input data acquisition and segmentation judgment.

[0107] The edge device (edge computer) obtains the multi-modal data input, including but not limited to text, image, voice, video, etc. The text data is encoded in UTF-8 format; the resolution requirement of the image data is not less than 640×480. The edge-side data storage format is floating-point 16-bit or fixed-point quantization format (such as INT8, INT4) to meet the edge-side memory limit.

[0108] Judge whether there is a qualified segmentation scheme for the encoding model obtained through Step ①. The conditions are that the computing power idle meets the requirements and the latency meets the requirements.

[0109] If not, transmit the original input data (multi-modal data) to the cloud side. If there is, continue to execute Step 2.2.

[0110] Step 2.2: Preprocessing and encoding of multi-modal data.

[0111] Perform format conversion and noise suppression on the multi-modal data to ensure that the data quality meets the requirements of subsequent processing.

[0112] Judge whether it contains image information. If it contains image information, perform image encoding processing.

[0113] Determine whether text information is included. If text information is included, perform text encoding processing.

[0114] On the edge device, use the embedding layer or encoder to perform vector representation on the preprocessed multi-modal data, and map data such as text, images, and speech to the same multi-modal representation space respectively.

[0115] Step 2.3: Multi-modal vector mapping and subsequent calculations.

[0116] Utilize the computing power on the edge side to uniformly map the encoded results of text and images to the same multi-modal encoding space, complete the inference calculation of the first-stage segmentation, generate intermediate data, and mark the position information of the calculated model segmentation points. This intermediate data undergoes fine-grained memory scheduling and dynamic compilation optimization to ensure low-latency output under limited edge-side resources.

[0117] The embedding dimension is usually set to 256, 512, or 1024, and is adjusted according to the model complexity and task requirements.

[0118] The encoder adopts a Transformer or lightweight CNN structure, and the number of parameters is pruned and quantized to ensure the optimization of computing efficiency and resource occupancy.

[0119] The quantization ratio is controlled between 0.05 and 0.1; the inference latency target is set within 50 to 150 milliseconds.

[0120] S3: The edge-side computer compresses and packages the intermediate data and the segmentation point position information, and sends them to the cloud-side computer through the transmission path; the transmission path is determined using a dynamic scheduling algorithm according to the real-time computing power and network resource status.

[0121] In practical applications, step ③ Dynamic data transmission.

[0122] Execution entities: Edge-side computer, cloud-side computer, network switch.

[0123] Step description: Transmit the intermediate data to the cloud, and based on the dynamically input multi-modal data, calculate the computing power and network resources to determine whether the deployment service requirements can be met. Implement dynamic task allocation and data transmission between the edge side and the cloud side to ensure the improvement of the overall inference efficiency, while reducing the data transmission volume and the risk of raw data leakage.

[0124] Step 3.1: Intermediate data processing.

[0125] Perform data compression, packaging, and encryption processing (such as using the AES-256 encryption algorithm) on the intermediate data and the segmentation point position information generated on the edge side.

[0126] Step 3.2: Dynamic task scheduling and data transmission.

[0127] Based on the real-time monitored computing power and network resource status, match servers with corresponding computing power in the cloud, allocate the computing tasks through dynamic scheduling algorithms (such as scheduling strategies based on deep reinforcement learning), optimize the computing resources and network transmission paths using the technology of computing-network convergence, and then send the packaged data to the cloud through the transmission path using the adaptive transmission protocol.

[0128] The size of the data packet is controlled between several hundred KB and several MB; the compression ratio target is 30% - 70%, which is specifically determined according to the data type.

[0129] The target of network bandwidth utilization is controlled within 80%, and the latency is controlled within 50 - 200 milliseconds; the resource scheduling period can be set to 1 - 5 seconds to dynamically adjust the task allocation strategy in real time.

[0130] Step 3.3: Establishment of feedback mechanism.

[0131] Establish a dynamic feedback mechanism for modeling network and computing power conditions to update and improve the cost model, so as to adjust the model segmentation judgment conditions in step ①.

[0132] S4: The cloud-side computer uses the processing model to perform computing processing on the intermediate data, obtains the computing result, and sends the computing result to the end-side computer; the processing model is obtained by initially segmenting the multi-modal large model to be deployed according to the rules of information encoding - information processing - information decoding.

[0133] As an optional implementation, using the processing model to perform computing processing on the intermediate data, obtaining the computing result, and sending the computing result to the end-side computer specifically includes:

[0134] Judge whether the encoding task is completed according to the intermediate data to obtain a first judgment result.

[0135] If the first judgment result is no, then execute the encoding task of the multi-modal data that the end-side computer has not completed.

[0136] If the first judgment result is yes, then judge whether the intermediate data is single data to obtain a second judgment result.

[0137] If the second judgment result is yes, then perform batch processing and unified calculation on the intermediate data to obtain the computing result, and send the computing result to the end-side computer.

[0138] If the second judgment result is no, then perform dynamic hybrid scheduling calculation on the intermediate data to obtain the computing result, and send the computing result to the end-side computer.

[0139] Send the candidate decoding model and the candidate points for splitting the decoding model to the edge computer.

[0140] In practical applications, as Figure 6 shown, Step ④: Cloud computing.

[0141] Execution entity: Cloud computer.

[0142] Step description: Perform intermediate result calculation and processing in the second stage of the splitting model, perform further inference and processing on the intermediate data in the cloud, perform further calculations on the data fed back from the edge side, dynamically supplement the tasks not completed on the edge side, and perform model integrity checks. Through intelligent scheduling and flexible allocation of resources, make full use of the computing performance of the cloud, and prepare to feedback the cloud computing results to the edge side.

[0143] Step 4.1: Data reception and model preparation.

[0144] Receive the packed intermediate data and the position information of the splitting points transmitted from the edge side.

[0145] Based on the calculations already completed on the edge side, the position information of the splitting points, and the encoding computing power fed back from the edge side, infer the decoding computing power that the edge side can provide, select the splitting points of the symmetrically split decoding model, and prepare for cloud large model calculations.

[0146] Step 4.2: Confirm that the encoding inference is completed.

[0147] Judge whether the encoding task has been completed. If so, proceed to Step 4.3.

[0148] If not, continue to execute the multi-modal information encoding tasks not completed on the edge side.

[0149] Step 4.3: Single image / text task processing.

[0150] Judge whether it only contains single image / text processing tasks. If so, collect, batch process, and uniformly calculate the received single image / text information, and complete dynamic cloud computing scheduling according to the computing power network situation. If not, proceed to Step 4.4.

[0151] Step 4.4: Dynamic scheduling of multi-modal hybrid tasks.

[0152] For the computing tasks containing multi-modal hybrid information, perform dynamic hybrid scheduling calculations in the cloud task space, allocate the hybrid tasks to the computing power network, and match heterogeneous computing devices (such as GPUs, TPUs, etc.) to provide computing power to meet the computing requirements.

[0153] In the cloud, through distributed computing graph segmentation and parallel scheduling, the intermediate data is further transmitted to the subsequent parts of the model for in-depth inference calculation, making full use of the high-performance computing resources in the cloud.

[0154] Distributed computing tasks can be grouped according to the number of GPU cores, and the latency target for each group of tasks does not exceed 200 milliseconds; the overhead model automatically adjusts the task priority according to the actual running duration and resource occupancy.

[0155] Step 4.5: Kernel scheduling optimization and data transmission.

[0156] On a single cloud server, optimize the kernel launch scheduling to achieve efficient computing.

[0157] Output the calculation results, feedback the model segmentation nodes that need to be further inferred on the edge side, and prepare to transmit the data back to the edge side.

[0158] After the cloud processing is completed, the calculation results (intermediate decoding information) are encapsulated, compressed, and encrypted and then transmitted back to the edge side. The data feedback latency is kept within 400 milliseconds; the data integrity and security are guaranteed through check codes (such as CRC32).

[0159] S5: The edge-side computer decodes the calculation results using the decoding model and converts them into the output format to obtain the processed multimodal data; the decoding model is obtained by initially segmenting the to-be-deployed multimodal large model according to the rules of information encoding - information processing - information decoding.

[0160] Step ⑤: Transmit the data to the edge side and complete decoding.

[0161] Execution entities: Edge-side computer, network switch.

[0162] Step description: Perform data decoding calculation in the third stage of the segmented model, receive the calculation results and the segmentation node positions from the cloud, and generate the required output (such as text or image) after processing to complete the entire edge-cloud collaborative inference process.

[0163] Step 5.1: Receive cloud computing data.

[0164] Transmit the cloud computing data to the client in the computing power network.

[0165] Step 5.2: Edge-side data decoding and output generation.

[0166] After the edge side receives the data returned from the cloud, complete the remaining calculations of the decoder and convert it into the final output format (such as text or image).

[0167] The number of decoder parameters corresponds to the order of magnitude and the amount of calculation of the encoder parameters, and the embedding dimensions are kept consistent.

[0168] The output image resolution can be dynamically adjusted (for example, it supports the output of images with any aspect ratio), and the length of text generation can be set to a fixed number of tokens (such as 640 tokens) to control the latency.

[0169] Step 5.3: Feedback and performance model improvement.

[0170] Provide end-side computing power data to further improve the cost model in step ①.

[0171] Compared with the related technology, this application has the following significant technical advantages and improvement effects:

[0172] (1) Dynamically adapt to the computing requirements of multi-modal large models:

[0173] Traditional technologies mainly optimize for fixed computation graphs and are difficult to adapt to the dynamic changes of multi-modal inputs. Existing end-side inference methods are also usually limited by computing power and memory and are difficult to efficiently run complete large-scale multi-modal models. This application can dynamically adjust the allocation of computing resources by dynamically splitting the model according to the computing characteristics of different modal data, improving the flexibility and adaptability of inference.

[0174] (2) Combine computing-network convergence to optimize end-cloud collaborative computing:

[0175] Most existing end-cloud collaborative inference methods focus on the optimization of computing resources and ignore the impact of network transmission on inference efficiency. This application introduces computing-network convergence technology and dynamically optimizes the computing and data transmission paths according to factors such as the complexity of computing tasks, bandwidth conditions, and latency requirements to achieve intelligent computing scheduling between the end and the cloud. This application not only optimizes the allocation of computing resources but also can adjust the inference strategy in combination with the network conditions to further improve the efficiency of end-cloud collaborative inference.

[0176] (3) Support heterogeneous computing devices and achieve intelligent task allocation:

[0177] This application can adapt to the computing power provided by different hardware architectures and support heterogeneous computing devices (such as CPUs, GPUs, NPUs, FPGAs, etc.); at the same time, it can optimize the allocation of computing tasks and data transmission paths to improve the adaptability and computing efficiency of distributed inference. In traditional methods, the coordination and optimization of computing and data transmission are insufficient during large-scale distributed inference. This application not only evaluates the computing cost but also can dynamically adjust the distribution method of computing tasks based on computing-network convergence to improve the overall inference performance.

[0178] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the above-mentioned end-cloud collaborative deployment method of the multi-modal large model is implemented.

[0179] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, which, when executed by a processor, implements the above-mentioned end-cloud collaborative deployment method for a multi-modal large model.

[0180] In an exemplary embodiment, a computer program product is provided, including a computer program, which, when executed by a processor, implements the above-mentioned end-cloud collaborative deployment method for a multi-modal large model.

[0181] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structural diagram can be as Figure 7 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. The computer program, when executed by the processor, implements a method for end-cloud collaborative deployment of a multi-modal large model.

[0182] Those skilled in the art can understand that Figure 7 the structure shown in

[0183] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0184] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0185] The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.

[0186] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0187] In this article, specific examples are used to elaborate on the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A method for end-cloud collaborative deployment of a multimodal large model, characterized in that: The end-cloud collaborative deployment method of the multimodal large model is implemented by an end-side computer and a cloud-side computer; the end-cloud collaborative deployment method of the multimodal large model includes: The cloud-side computer obtains the multimodal large model to be deployed, and determines an optimized coding model and a segmentation candidate point of the optimized coding model according to the multimodal large model to be deployed, and sends them to the terminal-side computer; the optimized coding model is obtained by optimizing the coding model; the coding model is obtained by initially segmenting the multimodal large model to be deployed according to the rule of information encoding-information processing-information decoding; the optimization processing includes quantization processing, pruning processing and fusion processing; The terminal computer acquires multimodal data, and uses the optimized encoding model to encode the multimodal data to obtain intermediate data and segmentation point position information; The terminal-side computer compresses and packages the intermediate data and the location information of the splitting points, and sends them to the cloud-side computer through a transmission path; the transmission path is determined by a dynamic scheduling algorithm based on real-time computing power and network resource status; The cloud-side computer uses the processing model to calculate the intermediate data to obtain a calculation result, and sends the calculation result to the terminal-side computer; the processing model is obtained by initially segmenting the multimodal large model to be deployed according to the rule of information encoding-information processing-information decoding; The terminal computer uses a decoding model to decode the calculation result and convert it into an output format to obtain processed multimodal data; the decoding model is obtained by initially dividing the large multimodal model to be deployed according to the rules of information encoding-information processing-information decoding.

2. The method for end-cloud collaborative deployment of a multimodal large model according to claim 1, characterized in that: Determining an optimized coding model and candidate segmentation points of the optimized coding model according to the multimodal large model to be deployed, and sending the optimized coding model and candidate segmentation points to a terminal-side computer, specifically includes: Initially segmenting the multimodal large model to be deployed according to the information encoding-information processing-information decoding rule to obtain an encoding model, a processing model, and a decoding model; Performing quantization processing, pruning processing and fusion processing on the coding model to obtain an optimized coding model; The optimized coding model is segmented to obtain a set of coding model segmentation schemes; the set of coding model segmentation schemes includes multiple candidate coding models and candidate segmentation points of the optimized coding model; Determine whether the encoding model segmentation scheme set is empty; If it is not empty, the candidate coding model is taken out from the set of coding model segmentation schemes, and the candidate coding model is evaluated using the cost model to determine whether the candidate coding model meets the requirements of the terminal side computer reasoning computing power, delay and data transmission speed; If satisfied, the optimized coding model and the candidate segmentation points of the optimized coding model are sent to the terminal side computer.

3. The method for end-cloud collaborative deployment of a multimodal large model according to claim 2, characterized in that: The optimized coding model and the candidate points for segmenting the optimized coding model are sent to the client computer, and then: Based on the set of coding model segmentation schemes and according to the computing power model, the decoding model is segmented symmetrically with the optimized coding model to obtain candidate decoding models and candidate decoding model segmentation points.

4. The method for end-cloud collaborative deployment of a multimodal large model according to claim 1, characterized in that: Acquiring multimodal data, and encoding the multimodal data using the optimized encoding model to obtain intermediate data and segmentation point location information, specifically including: Acquire the multimodal data, and determine whether there is a qualified segmentation scheme for the candidate coding model that meets the requirements of the terminal side computer reasoning computing power, delay, and data transmission speed; If so, encoding the multimodal data using the optimized encoding model to obtain encoded multimodal data; The computing power of the terminal computer is used to map the encoded multimodal data into the same multimodal coding space to obtain intermediate data, and the segmentation point position information is marked; the segmentation point position information is the segmentation point corresponding to the segmentation scheme that meets the conditions.

5. The method for end-cloud collaborative deployment of a multimodal large model according to claim 1, characterized in that: The processing model is used to calculate the intermediate data to obtain a calculation result, and the calculation result is sent to the terminal computer, specifically including: Determine whether the encoding task is completed according to the intermediate data to obtain a first determination result; If the first judgment result is no, executing the encoding task of the multimodal data that has not been completed by the terminal-side computer; If the first judgment result is yes, then judging whether the intermediate data is a single data, and obtaining a second judgment result; If the second judgment result is yes, batch processing and unified calculation are performed on the intermediate data to obtain a calculation result, and the calculation result is sent to the terminal side computer; If the second judgment result is no, a dynamic hybrid scheduling calculation is performed on the intermediate data to obtain a calculation result, and the calculation result is sent to the terminal-side computer.

6. The method for end-cloud collaborative deployment of a multimodal large model according to claim 3 is characterized in that: Also includes: The candidate decoding model and the candidate points for segmenting the decoding model are sent to the terminal-side computer.

7. The method for end-cloud collaborative deployment of a multimodal large model according to claim 2, characterized in that: Also includes: A dynamic feedback mechanism for computing power and network resources is established, and the dynamic feedback mechanism is used to update and improve the cost model.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the end-cloud collaborative deployment method of a multimodal large model as described in any one of claims 1-7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the end-cloud collaborative deployment method of a multimodal large model described in any one of claims 1-7.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the end-cloud collaborative deployment method of a multimodal large model described in any one of claims 1-7.

Citation Information

Patent Citations

  • Cloud edge semantic collaborative segmentation model reasoning method

    CN118072150A

  • Privacy preserving cooperative learning in untrusted environments

    US20220300618A1

Cited By

  • Cloud edge cooperative computing framework for multi-modal data stream fusion processing and processing method

    CN120872532A

  • Multi-mode-based hardware accelerator design space search method and device and electronic equipment

    CN120995953A

  • Image generation method, scheduling equipment, system and storage medium

    CN121033213A

  • Efficient large language model adaptation method based on server-free edge computing

    CN121072646A

  • Construction method and device of segmentation model and storage medium

    CN121214922A