Large and small model collaboration method and device based on cloud link and storage medium

Through the coordinated processing of the end-side small multimodal model and the large multimodal model in the cloud, the problems of limited resources on the end-side and delay in the cloud are solved, and efficient multimodal data processing and model optimization are achieved.

CN120354934APending Publication Date: 2025-07-22SHENZHEN EXTREME VISION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510319446.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, the computing resources of the end-side device are limited and difficult to process complex multimodal data, while cloud processing faces problems such as data transmission delay and privacy leakage, and it is difficult to integrate and understand multimodal data.

Method used

The end-side small multimodal model is used for preliminary processing and pre-configured PPO strategy optimization, combined with large multimodal model in the cloud for in-depth analysis, and a tight collaborative mechanism for large-scale and model models is established to achieve parallel computing and collaborative acceleration.

Benefits of technology

Reduce data transmission delay, improve resource utilization, focus on complex task processing in the cloud, enhance the accuracy and comprehensiveness of multimodal data understanding, and optimize the end-side model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354934A_ABST
    Figure CN120354934A_ABST
Patent Text Reader

Abstract

The invention discloses a large and small model collaboration method and device based on a link and a storage medium, and the method comprises the steps: end-side equipment collects multi-modal data through a built-in sensor, carries out the preprocessing of the multi-modal data, and inputs a preprocessing result into a small multi-modal model for reasoning, and obtains a preliminary reasoning result; the preliminary reasoning result and the state description information are input into a pre-configured PPO for processing, and if PPO output needs to be uploaded to a cloud, an original data preprocessing result and the preliminary reasoning result are packaged and uploaded to a cloud server; the cloud server receives and preprocesses the data, inputs a result into a large-scale multi-modal model for reasoning analysis to generate a report, and then sends the report to the end side equipment; the end-side equipment collects feedback information of model reasoning and uploads the feedback information to the cloud server, and the cloud server updates the model in sequence and transmits updated model parameters back to the end-side equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a method, device, and storage medium for collaborative operation of large and small models based on a cloud link. Background Art

[0002] With the rapid development of technologies such as the Internet of Things and artificial intelligence, the amount of data generated by intelligent devices has increased explosively, and the data exhibits multimodal characteristics, covering images, audio, text, and various sensor data, etc. When processing this multimodal data, the traditional processing methods of a single model or a single computing environment face huge challenges.

[0003] On the one hand, although edge devices can directly access the data source, due to their relatively limited computing resources and storage capabilities, it is difficult to run complex large-scale models to process complex multimodal tasks. For example, in an intelligent security camera, if only relying on edge devices to perform complex image semantic segmentation and behavior analysis, the processing speed may be slow due to insufficient computing power, and it may not be able to respond to abnormal situations in a timely manner.

[0004] On the other hand, cloud servers have powerful computing and storage resources, but directly transmitting all data to the cloud for processing will face problems such as data transmission delay, network bandwidth pressure, and privacy leakage. For example, in some scenarios with extremely high real-time requirements, such as road condition perception in autonomous driving, if all data is uploaded to the cloud for processing, data transmission delay may occur due to network fluctuations, which will in turn affect the decision-making reaction time of the vehicle and endanger driving safety.

[0005] In addition, the fusion and understanding of multimodal data require a large amount of knowledge and data support, and a single model structure is difficult to fully exploit the information between different modal data. Summary of the Invention

[0006] This application provides a method, device, and storage medium for collaborative operation of large and small models based on a cloud link, which are used to solve the problems of computing resource allocation and delay.

[0007] In the first aspect of this application, a method for collaborative operation of large and small models based on a cloud link is provided, including:

[0008] Real-time collection of multimodal data by sensors built in edge devices and preprocessing to obtain the preprocessing result of the original data;

[0009] Inputting the preprocessing result of the original data into a small multimodal model deployed on the edge device for inference calculation to obtain a preliminary inference result, and inputting the state description information corresponding to the preliminary inference result into a pre-configured PPO for processing;

[0010] When the action output by the PPO needs to be uploaded to the cloud, the original data preprocessing result and the preliminary inference result are packaged and sent to the cloud server;

[0011] The cloud server receives data from the edge device and preprocesses the data of the edge device to obtain a preprocessing result;

[0012] The preprocessing result is input into a large multi-modal model for inference analysis to generate a data analysis report, and the data analysis report is sent back to the edge device;

[0013] The edge device collects model inference feedback information during the data processing process and then uploads the model inference feedback information to the cloud server;

[0014] The cloud server updates the model according to the model inference feedback information and then sends the updated model parameters back to the edge device.

[0015] Optionally, a policy network is pre-configured in the PPO, and the preliminary inference result and the state description information corresponding to the preliminary inference result are input into the pre-configured PPO for processing.

[0016] The state description information is constructed into a multi-dimensional feature vector;

[0017] The multi-dimensional feature vector is input into the policy network to obtain the probability distribution π(a|s; θ) output by the policy network, where θ represents network parameters, s represents the multi-dimensional feature vector, and a represents the sampled action;

[0018] The sampled action is obtained according to the probability distribution.

[0019] Optionally, the policy network includes an input layer, a hidden layer, and an output layer. Inputting the multi-dimensional feature vector into the policy network to obtain the probability distribution π(a|s; θ) output by the policy network includes:

[0020] The input layer processes the multi-dimensional feature vector through the following formula:

[0021] h1 = ReLU(W1 × s + b1);

[0022] ReLU: f(x) = max(0, x);

[0023] where h1 represents the output of the input layer, ReLU represents the activation function, W1 represents the weight matrix of the input layer, s represents the multi-dimensional feature vector, and b1 represents the bias term of the input layer;

[0024] The hidden layer processes the output of the input layer through the following formula:

[0025] h2 = ReLU(W2 × h1 + b2);

[0026] where h2 represents the output of the hidden layer, ReLU represents the activation function, W2 represents the weight matrix of the hidden layer, s represents the multi-dimensional feature vector, and b2 represents the bias term of the hidden layer;

[0027] The output layer generates an action score through the following formula:

[0028] z = W3 × h2 + b3;

[0029] where z represents the action score, W3 represents the weight matrix of the output layer, s represents the multi-dimensional feature vector, and b3 represents the bias term of the output layer;

[0030] The action score is converted into a probability distribution through the following formula:

[0031]

[0032] where π(a|s; θ) represents the probability distribution, Softmax(z i ) represents the probability of the i-th class, exp(z i ) represents the exponential value of z i , and represents the sum of the exponential values of all z j from j = 1 to k.

[0033] Optionally, the policy network is optimized through the following loss function:

[0034] L CLIP (θ) = E t [min(r t (θ)·A t , clip(r t (θ), 1 - ε, 1 + ε)·A t )];

[0035]

[0036] A t = G t - V(s t );

[0037]

[0038] where L CLIP (θ) represents the loss function, E t represents the sum over all time steps t and sampled action pairs (s t , at ) Perform expectation calculation, clip(r t (θ), 1 - ε, 1 + ε) represents restricting r t (θ) within the range of (1 - ε, 1 + ε), where ε is a hyperparameter, typically taking values ε = 0.1 or ε = 0.2;

[0039] r t (θ) represents the probability ratio of the new and old policies for a given state and sampled action, πθ(a t |s t ) represents the probability of choosing action a t under the new policy for state s t , and πθ old (a t |s t ) represents the probability of choosing action a t under the old policy for state s t ;

[0040] A t represents the advantage function, V(s t ) represents the state value function, G t represents the cumulative reward, r t represents the environmental feedback reward, and γ represents the discount factor.

[0041] Optionally, the preprocessing includes: cleaning, classifying, and feature - processing the collected raw data.

[0042] Optionally, the small - scale multimodal model performs preliminary inference through the following steps:

[0043] Classify the pre - processed results of the raw data into image data and text data;

[0044] Extract features from the image data and the text data to obtain an image modality and a text modality;

[0045] Perform a splicing and fusion operation on the feature vectors extracted from the image modality and the text modality to obtain a fused feature vector;

[0046] Input the fused feature vector into the softmax layer of the small - scale multimodal model for final inference calculation.

[0047] Optionally, for feature extraction from the image data and text data, the feature extraction method includes:

[0048] Slide a convolution kernel over the image data for convolution operation to extract local features of the image through convolution;

[0049] Input the text data into the RNN, which captures the semantic information and context relationships of the text through hidden states.

[0050] The second aspect of this application provides a collaborative system for large and small models based on a cloud link. The system includes:

[0051] An acquisition unit. The end-side device collects multi-modal data in real time through built-in sensors and performs preprocessing to obtain the preprocessing result of the raw data.

[0052] A calculation unit, which is used to input the preprocessing result of the raw data into a small multi-modal model deployed on the end-side device for inference calculation to obtain a preliminary inference result, and input the status description information corresponding to the preliminary inference result into a pre-configured PPO for processing.

[0053] A conveying unit. When the action output by the PPO needs to be uploaded to the cloud, it packs the preprocessing result of the raw data and the preliminary inference result and conveys them to the cloud server.

[0054] A receiving unit. The cloud server receives the data from the end-side device and preprocesses the data of the end-side device to obtain a preprocessing result.

[0055] An analysis unit, which is used to input the preprocessing result into a large multi-modal model for inference analysis to generate a data analysis report, and the data analysis report is transmitted back to the end-side device.

[0056] A feedback unit. The end-side device collects the model inference feedback information during the data processing process and then uploads the model inference feedback information to the cloud server.

[0057] An update unit. The cloud server updates the model according to the model inference feedback information and then transmits the updated model parameters back to the end-side device.

[0058] Optionally, the calculation unit includes:

[0059] A classification module, which is used to classify the preprocessing result of the raw data into image data and text data.

[0060] An extraction module, which is used to extract features from the image data and text data to obtain an image modality and a text modality.

[0061] A splicing and fusion module, which is used to perform a splicing and fusion operation on the feature vectors extracted from the image modality and the text modality by the image extraction module to obtain a fused feature vector.

[0062] An input module for inputting the fused feature vector into the softmax layer of the small multi-modal model for final inference calculation.

[0063] The third aspect of the present application provides a large and small model collaboration device based on a cloud link, and the device includes:

[0064] A processor, a memory, an input / output unit, and a bus;

[0065] The processor is connected to the memory, the input / output unit, and the bus;

[0066] The memory stores a program, and the processor calls the program to execute the large and small model collaboration method in the first aspect and any optional one in the first aspect.

[0067] The fourth aspect of the present application provides a computer-readable storage medium, and a program is stored on the computer-readable storage medium, and when the program is executed on a computer, it executes the large and small model collaboration method in the first aspect and any optional one in the first aspect.

[0068] It can be seen from the above technical solutions that the present application has the following advantages:

[0069] The small multi-modal model in the edge device can quickly process some simple multi-modal tasks, such as basic image classification, simple voice command recognition, etc., reducing the latency of data transmission to the cloud and the computing burden on the cloud server.

[0070] The large multi-modal model in the cloud server can focus on processing complex tasks, such as high-precision image semantic segmentation, deep natural language understanding, etc. Since the cloud server has rich computing resources, more complex model structures and algorithms can be used for in-depth inference analysis, thereby improving the processing accuracy and comprehensiveness of complex multi-modal data.

[0071] This method establishes a tight large and small model collaboration mechanism. The preliminary inference result of the small multi-modal model can be used for further analysis by the large multi-modal model, and the analysis report obtained by the large multi-modal model can in turn optimize the small multi-modal model in the edge device. The large multi-modal model and the small multi-modal model can simultaneously perform task processing in different stages, realizing parallel computing and collaborative acceleration. The small multi-modal model immediately performs preliminary processing after data collection, while the large multi-modal model can use its powerful computing resources to quickly perform in-depth inference analysis when receiving data that needs further processing. The two cooperate with each other, shortening the time of the entire data processing process. Description of the Drawings

[0072] To more clearly illustrate the technical solutions in the present application, the following will briefly introduce the accompanying drawings required in the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0073] Figure 1 It is a schematic flowchart of the first embodiment in the method for collaborative operation of large and small models based on a cloud link provided by the present application;

[0074] Figure 2 It is a schematic flowchart of the second embodiment in the method for collaborative operation of large and small models based on a cloud link provided by the present application;

[0075] Figure 3 It is a schematic flowchart of the third embodiment in the method for collaborative operation of large and small models based on a cloud link provided by the present application;

[0076] Figure 4 It is a schematic structural diagram of an embodiment of the system for collaborative operation of large and small models based on a cloud link provided by the present application;

[0077] Figure 5 It is a schematic structural diagram of an embodiment of the device for the method of collaborative operation of large and small models based on a cloud link provided by the present application. Specific Embodiments

[0078] The present application provides a method, device, and storage medium for collaborative operation of large and small models based on a cloud link, which are used to solve the problems of computing resource allocation and latency.

[0079] It should be noted that a method for collaborative operation of large and small models based on a cloud link provided by the present application can be applied to the end side or on a server. For example, the end side can be a smart phone, computer, tablet computer, smart TV, smart watch, portable computer, or the end side can also be a fixed end side such as a desktop computer. For the convenience of explanation, the present application takes the end side as the execution subject for example.

[0080] In the first embodiment, the present application provides a method for collaborative operation of large and small models based on a cloud link, which is used to solve the problems of computing resource allocation and latency. Please refer to Figure 1 . The method includes:

[0081] S101. Real-time collect multimodal data through a sensor built in the end-side device and perform preprocessing to obtain a preprocessing result of the original data;

[0082] The original multimodal data has problems such as high noise, inconsistent formats, and large data volumes. Directly using the original data for analysis and processing is not only inefficient but may also lead to inaccurate results. Therefore, it is necessary to preprocess the collected original data on the edge device. The preprocessing of the original data includes cleaning, classifying, and characterizing the collected original data.

[0083] For the collected multimodal original data, first check whether there are outliers that significantly deviate from the normal range. If there are outliers, choose to directly delete them or replace them with nearby reasonable values. If the original data is audio data, it may be mixed with environmental background noise, such as noisy machine sounds, wind sounds, etc. Use noise reduction algorithms to remove this noise. When collecting motion sensor data, if the values recorded at several consecutive sampling moments are exactly the same and do not conform to the normal motion change pattern, these repeated data need to be identified and deleted to improve the efficiency of subsequent processing.

[0084] After cleaning the original data, classify the original data into different modal categories to facilitate corresponding processing and analysis of the characteristics of different modalities.

[0085] After classifying and partitioning the original data, perform feature extraction on the original data, and encode and quantify the extracted features to facilitate storage and calculation on the edge device. For example, represent the color features of an image using binary encoding.

[0086] Clean, classify, and perform feature extraction on the collected original data to improve the data quality and processing efficiency of the original data, and reduce the pressure on subsequent data transmission and cloud processing.

[0087] S102: Input the preprocessing result of the original data into the small multimodal model deployed on the edge device for inference calculation to obtain a preliminary inference result, and input the status description information corresponding to the preliminary inference result into the pre-configured PPO for processing;

[0088] Organize the preprocessed original data according to the format and specifications required by the small multimodal model. For example, adjust the image data to an appropriate resolution, perform word segmentation on the text data, etc., and then input this processed data into the small multimodal model deployed on the edge device.

[0089] After receiving the input data, the small multimodal model will perform inference calculation according to its internal neural network structure and algorithm. During the inference process, the model will extract and fuse features of different modality data. For example, for multimodal data of images and text, the model will respectively extract the visual features of the image and the semantic features of the text, and then fuse these features to obtain more comprehensive information. After the inference calculation, the small multimodal model will output a preliminary inference result.

[0090] Perform preliminary inference calculations directly on the edge device without transmitting data to the cloud or other remote servers, greatly reducing the latency of data transmission and enabling quick obtainment of preliminary inference results. Even in the case of network connection interruption or unavailable cloud services, the edge device can still process the collected data and make corresponding responses relying on a small multi-modal model, fully utilizing the resources of the edge device and improving resource utilization.

[0091] PPO is an algorithm developed based on the reinforcement learning framework. In the interaction between the edge device data processing system and the environment, it learns the optimal strategy by continuously trial and error. First, collect the preliminary inference results output by the small multi-modal model on the edge device, and organize the state description information corresponding to the preliminary inference results, such as the data complexity index reflecting the complexity of the data itself, the confidence level reflecting the credibility of the preliminary inference results, the available network bandwidth situation, the latency time experienced from data collection to obtaining the preliminary inference throughout the entire process, and the inference time spent by the small multi-modal model for inference calculations.

[0092] Then, combine the preliminary inference results and each state description information in a predefined format and input them into the pre-configured PPO algorithm. The PPO algorithm will perform corresponding processing based on its internal set policy network structure, hyperparameters, etc., according to the input content. The policy network generates the probability distribution corresponding to different actions based on the state information, and then determines the sampled action to be taken according to the probability distribution. It continuously updates the parameters of the policy network based on the feedback obtained from the executed action, optimizing the decision-making of the policy network to adapt to different edge data processing.

[0093] S103. When the action output by PPO is to be uploaded to the cloud, then package the original data preprocessing result and the preliminary inference result and transport them to the cloud server;

[0094] When detecting the action that the output of PPO is to be uploaded to the cloud, package the original data preprocessing result and the preliminary inference result. When using the binary format, it is necessary to first define the corresponding data structure description file, clarify the message types, fields and other information corresponding to each part of the original data preprocessing result and the preliminary inference result in it, and then generate the corresponding code class according to this description file through a special compiler. Use these code classes to perform serialization operations on the actual data, converting it into a compact binary format for transmission and parsing in the cloud. After packaging is completed, determine an appropriate communication protocol according to the configuration of the cloud server and the network environment to establish a connection with the cloud and transport the packaged data to the cloud server.

[0095] S104. The cloud server receives data from the edge device, preprocesses the data of the edge device, and obtains a preprocessing result;

[0096] The edge device can initiate a data transmission request to the cloud server and send the packed data to the cloud server. The cloud server listens and receives these requests on the corresponding port. When receiving, the cloud needs to verify the data format to ensure that it can correctly parse the data content. At the same time, it is necessary to check the integrity of the data, such as through methods like checksum and data length verification, to prevent data loss, damage, etc. during the transmission process and ensure that the received data is the complete and valid data sent by the edge device.

[0097] After receiving the data, the cloud server preprocesses the data. The data sent by different edge devices may vary in format. The cloud server needs to unify these different formats of data for subsequent processing. In order to make the data collected by different edge devices or different batches of data collected by the same device comparable in terms of numerical range, data standard, etc., the cloud server will perform data standardization processing so that the data can be compared and operated under the same standard.

[0098] After the data preprocessing operation, the cloud server finally obtains a preprocessing result that meets the requirements of subsequent processing, making further preparations for the inference and analysis of the large multi-modal model.

[0099] S105. Input the preprocessing result into the large multi-modal model for inference and analysis, generate a data analysis report, and transmit the data analysis report back to the edge device;

[0100] According to the data analysis task and data characteristics, select an appropriate model from the pre-trained large multi-modal models. When the preprocessing result is input into the large multi-modal model, the model will first extract features from data of different modalities. For the image modality, the CNN layer will automatically learn features such as texture, shape, and color in the image; for the audio modality, the RNN layer will extract features such as the spectrum, rhythm, and pitch of the audio; for the text modality, semantic and syntactic features will be extracted through word vector representation and neural network layers. Then, in the feature fusion layer, the model will fuse these features of different modalities. The fusion method can be simple concatenation, weighted summation, or more complex fusion methods based on the attention mechanism. Through feature fusion, the model can comprehensively consider information of multiple modalities, thereby more comprehensively and accurately understanding the situation or object represented by the input data.

[0101] After completing multi-modal feature fusion, the model will perform inference analysis based on the fused features. This process involves a series of complex computational and decision-making mechanisms within the model. For example, the model may classify, predict, or generate new content for the input data according to the learned patterns and rules. The model will utilize its huge parameters and complex network structure to deeply analyze and understand the input data in order to obtain the inference results that best conform to the data characteristics.

[0102] Based on the inference analysis results of the large multi-modal model, the generated data analysis report will contain multiple aspects of content. The report will elaborate in detail on the inference analysis results of the model, including the classification results of the data, prediction values, abnormal situations, etc. In addition, the report also includes the evaluation of data quality, the evaluation of model performance, as well as some relevant suggestions and conclusions. The large multi-modal model can conduct in-depth inference analysis on the multi-modal data preprocessed by the cloud server and generate a comprehensive and detailed data analysis report, providing strong support for decision-making and problem-solving in related fields.

[0103] The generated analysis report is returned to the edge device through the cloud communication link. The edge device presents the parsed data analysis report to the user according to its own display capabilities and user interface design. For example, if the edge device is a mobile device such as a smartphone or a tablet, the data analysis report is displayed through a dedicated application interface in the form of charts, text lists, etc. to allow the user to intuitively understand the results, conclusions, and suggestions of the data analysis; if it is other types of edge devices, the report content will also be displayed according to the corresponding display rules and layouts, facilitating the operator to view and make corresponding decisions based on the report.

[0104] In practical applications, during the vehicle and pedestrian recognition process, for the vehicle analysis scenario, the edge device trains a small vehicle detection model for detection; for the pedestrian analysis scenario, the edge device trains a small pedestrian detection model to identify pedestrians. When it detects that the vehicle crosses the line and the pedestrian does not wear a seat belt, after the small vehicle detection model and the small pedestrian detection model obtain the images, they will transmit the image data and the input text data to the visual language large model in the cloud for analysis. By inputting the image data and text data into the visual language large model, the visual language large model finally outputs text content through training. The text content is transmitted to the edge device, and the edge device gives voice prompts through the output text content, such as "The pedestrian does not wear a seat belt, the result is unqualified", "The vehicle body has crossed the line, the result is unqualified", etc.

[0105] S106. The edge device collects the model inference feedback information during the data processing process and then uploads the model inference feedback information to the cloud server;

[0106] The edge device collects relevant information during the data processing process, such as edge model inference error cases, local data feature changes, user feedback information, etc., and organizes and packages this information and uploads it to the cloud server.

[0107] In practical applications, when the edge device performs model inference, it will record the sample data of inference errors. For example, in image recognition, if the edge small multi-modal model misidentifies a picture of a cat as a dog, the edge device will save the original data of this picture and the relevant parameters during model inference, and analyze the cause of the error based on the relevant information. During the analysis process, if it is found that multiple edge devices take too long to process the inference of a specific type of data, the cloud server will check the calculation logic of the model and try to adopt a more efficient algorithm or optimize the calculation resource allocation to reduce the inference time.

[0108] S107. The cloud server updates the model according to the model inference feedback information, and then transmits the updated model parameters back to the edge device.

[0109] The cloud server receives the model inference feedback information uploaded from each edge device, including inference error cases, local data feature changes, user feedback information, and performance metric data, etc. It summarizes this multi-source heterogeneous data to ensure the integrity and traceability of the data.

[0110] According to the model inference feedback information, calculate the difference between the current prediction result of the model and the actual situation, and use a loss function to quantify this difference. For example, in image classification, the cross-entropy loss function is commonly used to measure the difference between the predicted category of the model and the actual category. Then, update the model parameters according to the gradient of the loss function, and finally perform iterative calculations to gradually adjust the parameter values to minimize the loss function, thereby improving the accuracy of the model. After updating the model, use the validation data set to evaluate the model and check the performance of the model in terms of accuracy and other metrics. If the model performance does not meet the expectations, adjust the training parameters or retrain until the model performance meets the requirements, realizing the continuous learning and optimization of the model. The cloud server encrypts the updated model parameters or model structure optimization information and then transmits it back to the edge device. The edge device updates the small multi-modal model according to the instructions of the cloud server to ensure that the edge model can continuously adapt to new application scenarios and data changes.

[0111] In this embodiment, the small multi-modal model of the edge device can reasonably allocate tasks according to its own resource status, quickly process some simple multi-modal tasks, such as basic image classification, simple voice command recognition, etc., reducing the delay of data transmission to the cloud and the computing burden of the cloud server, improving resource utilization, and avoiding problems such as excessive resource consumption and device overheating caused by running complex models.

[0112] Large multi-modal models on cloud servers can focus on processing complex tasks, such as high-precision image semantic segmentation, deep natural language understanding, etc. Since the cloud has rich computing resources, more complex model structures and algorithms can be adopted for in-depth reasoning and analysis, thereby improving the accuracy and comprehensiveness of processing complex multi-modal data.

[0113] This method establishes a tight cooperation mechanism between large and small models. The preliminary inference results of the small multi-modal model can be used for further analysis by the large multi-modal model, and the analysis report obtained by the large multi-modal model can in turn optimize the small multi-modal model in the edge device. The large multi-modal model and the small multi-modal model can simultaneously process tasks at different stages, realizing parallel computing and collaborative acceleration. The small multi-modal model immediately performs preliminary processing after data collection, while the large multi-modal model can use its powerful computing resources to quickly perform in-depth reasoning and analysis when receiving data that needs further processing. The two cooperate with each other, shortening the time of the entire data processing process.

[0114] Embodiment 2, please refer to Figure 2 , the following is a detailed introduction to the preliminary inference results output by the small multi-modal model according to the preprocessing results of the original data:

[0115] S201. Real-time collect multi-modal data through sensors built in the edge device and perform preprocessing to obtain the preprocessing results of the original data;

[0116] In this embodiment, step S301 is similar to step S101 in the foregoing embodiment, and will not be elaborated here.

[0117] S202. Classify the preprocessing results of the original data into image data and text data;

[0118] Distinguish image data and text data according to the format, structure and content characteristics of the data itself. Image data has specific image file format identifiers, such as JPEG, PNG, etc. formats, and is a two-dimensional or multi-dimensional matrix composed of pixel points. By parsing the file header information and the storage structure of the data, it can be judged whether it is image data, so as to classify the image data.

[0119] Text data exists in the form of character encoding, such as text encodings such as ASCII, UTF-8, etc., and is a sequence composed of characters such as letters, numbers, and punctuation marks. It is distinguished by identifying these characteristics, and tools and methods such as file type judgment functions, data format parsing modules, and simple regular expressions in programming are used to complete the classification of text data.

[0120] Accurately divide the preprocessed original data into two parts: an image data set and a text data set, preparing for extracting features of different modalities later.

[0121] S203. Extract features from the image data and text data to obtain the image modality and text modality;

[0122] Extract features from the classified image data and text data respectively. For the image data, methods such as calculating color statistical information, obtaining texture features by means of a gray-level co-occurrence matrix, calculating shape features through edge detection, and extracting deep features using a convolutional neural network are used to integrate these features from different aspects to form an image modality that can represent the overall characteristics of the image.

[0123] For the text data, first convert words into word vectors to represent semantics, and then combine part-of-speech tagging, syntactic analysis to obtain syntactic features, and use term frequency and inverse document frequency to determine important words, etc. These features are comprehensively processed to construct a text modality that can represent the overall situation of the text, thereby completing the modality construction of two different types of data and preparing for subsequent operations such as feature fusion.

[0124] S204. Perform a splicing fusion operation on the feature vectors extracted from the image modality and text modality to obtain a fused feature vector;

[0125] Splice the feature vectors extracted from the image modality and the text modality in sequence in the dimension, so that the feature vectors of the image modality and the text modality are fused to form a longer vector. For example, if the dimension of the image modality feature vector is n and the dimension of the text modality feature vector is m, then the dimension of the spliced fused feature vector is n + m. This method is simpler and more intuitive and retains the complete feature information of each of the two modalities.

[0126] S205. Input the fused feature vector into the softmax layer of the small multi-modal model for the final inference calculation.

[0127] When the fused feature vector is input into the softmax layer of the small multi-modal model, a series of linear transformation and non-linear activation operations are first performed. Inside the softmax layer, the scores corresponding to each category are calculated according to the input fused feature vector, and then these scores are converted into probability values through the softmax function. The calculation formula of the softmax function is as follows:

[0128]

[0129] where x represents the input fused feature vector, i represents the category index, such as i = 1, 2... k, k is the total number of categories, z iis the score corresponding to category i obtained through the calculation of the previous network layer. p(y = i|x) is the probability of belonging to category i when inputting x. After the calculation of the softmax layer, a probability distribution vector is finally output. For example, in a three-classification task, it may output (0.3, 0.5, 0.2), indicating that the probability of belonging to the first category is 0.3, the probability of belonging to the second category is 0.2, and the probability of belonging to the third category is 0.5. According to this probability distribution, classification decisions can be made, and the category with the highest probability is selected as the final inference result, thus completing the inference calculation process of the entire small multi-modal data. Then, the preliminary inference result is input into the pre-configured PPO for the next step of processing.

[0130] S206. When the action output by the PPO needs to be uploaded to the cloud server, the original data preprocessing result and the preliminary inference result are packaged and sent to the cloud server.

[0131] S207. The cloud server receives the data from the edge device and preprocesses the data of the edge device to obtain the preprocessing result.

[0132] S208. The preprocessing result is input into the large multi-modal model for inference analysis to generate a data analysis report, and the data analysis report is sent back to the edge device.

[0133] S209. The edge device collects the model inference feedback information during the data processing process and then uploads the model inference feedback information to the cloud server.

[0134] S210. The cloud server updates the model according to the model inference feedback information and then sends the updated model parameters back to the edge device.

[0135] In this embodiment, steps S206 to S210 are similar to steps S103 to S107 in the previous embodiment and will not be elaborated here.

[0136] In this embodiment, using the small modality for preliminary inference reduces the latency of data transmission to the cloud and the computational burden on the cloud server, improves resource utilization, and provides important reference data for the cloud server to further analyze in depth, intervene manually, or make comprehensive decisions.

[0137] Embodiment 3. Please refer to Figure 3 , when the action output by the PPO needs to be uploaded to the cloud server, the original data preprocessing result and the preliminary inference result are packaged and sent to the cloud server. The following is a detailed introduction to the sampling action output by the PPO according to the preliminary inference result and each status description information:

[0138] S301. The edge device collects multi-modal data in real time through the built-in sensor and preprocesses it to obtain the original data preprocessing result.

[0139] S302. Perform inference calculation on the small multi-modal model deployed on the input end device of the preprocessed result of the original data to obtain a preliminary inference result, and input the status description information corresponding to the preliminary inference result into the pre-configured PPO for processing;

[0140] In this embodiment, steps S301 to S302 are similar to steps S101 to S102 in the foregoing embodiment, and will not be elaborated here.

[0141] S303. Construct the status description information into a multi-dimensional feature vector;

[0142] The status description information includes data complexity, confidence level, bandwidth, latency, and inference time. First, determine the arrangement order of each feature in the multi-dimensional feature vector, and then arrange each feature in sequence to form a fixed dimension order convention. According to the determined dimension order, fill the previously quantized values of each feature into the corresponding positions in the vector in sequence to form a multi-dimensional feature vector.

[0143] S304. Input the multi-dimensional feature vector into the policy network to obtain the probability distribution output by the policy network;

[0144] The following is the formula for the probability distribution:

[0145] π(a|s; θ);

[0146] where θ represents the network parameter, s represents the multi-dimensional feature vector, and a represents the sampling action.

[0147] The policy network is a multi-layer neural network that realizes the probability distribution from the state to the action through fully connected layers, activation functions, and Softmax operations. When the multi-dimensional feature vector s is fed into the policy network, the data will be calculated layer by layer according to the established network structure. In each layer, the neurons will multiply the input data by the corresponding weights and add the bias, and then perform non-linear processing through the activation function (ReLU) to extract and integrate features, continuously extracting the feature representations of the next layer.

[0148] After multiple layers of such operations, finally, the output layer of the policy network will output the probability distribution corresponding to different sampling actions. This probability distribution is presented in the form of a set of probability values, and each probability value corresponds to a sampling action a, indicating the likelihood of taking this sampling action under the state represented by the current input multi-dimensional feature vector s.

[0149] Obtaining the probability distribution π(a|s; θ) output by the policy network includes:

[0150] The input layer processes the multi-dimensional feature vector through the following formula:

[0151] h1 = ReLU(W1 × s + b1);

[0152] ReLU: f(x) = max(0, x);

[0153] Where h1 represents the output of the input layer, ReLU represents the activation function, W1 represents the weight matrix of the input layer, s represents the multi-dimensional feature vector, and b1 represents the bias term of the input layer;

[0154] The hidden layer processes the output of the input layer through the following formula:

[0155] h2 = ReLU(W2 × h1 + b2);

[0156] Where h2 represents the output of the hidden layer, ReLU represents the activation function, W2 represents the weight matrix of the hidden layer, s represents the multi-dimensional feature vector, and b2 represents the bias term of the hidden layer;

[0157] The output layer generates the action score through the following formula:

[0158] z = W3 × h2 + b3;

[0159] Where z represents the action score, W3 represents the weight matrix of the output layer, s represents the multi-dimensional feature vector, and b3 represents the bias term of the output layer;

[0160] The action score is converted into a probability distribution through the following formula:

[0161]

[0162] Where π(a|s; θ) represents the probability distribution, Softmax(z i ) represents the probability of the i-th class, exp(z i ) represents the exponential value of z i , represents the sum of the exponential values of all z j from j = 1 to k.

[0163] S305. Finally, obtain the sampled action according to the probability distribution.

[0164] According to the probability values of each sampled action output by the policy network, a parameter c is introduced to explore the relationship between exploring new actions and exploiting known optimal actions. Each time the policy network needs to select a sampled action, a random number r uniformly distributed in the interval (0, 1) is first generated. If r is less than c, the action is selected according to the probability distribution by the random sampling method; if r is greater than or equal to c, the sampled action with the highest probability is selected according to the policy. For example, setting c = 0.2 and the generated random number r = 0.15, since r < c, the random sampling method is used to select an action from the actions corresponding to the probability distribution; if the generated random number r = 0.8, since r ≥ c, the action with the highest probability in the probability distribution is directly selected as the sampled action.

[0165] S306. After the policy network output function, it is optimized through the loss function:

[0166] The loss function is used in the optimization process of the policy network in reinforcement learning. By adjusting the network parameters, the action probability distribution output by the policy network can make better decisions, thereby maximizing the cumulative reward. Overall, the loss function is constructed based on the ratio of the old and new policy probabilities, the advantage function, and other related elements, and the policy network is optimized by minimizing these losses.

[0167] The formula of the following loss function:

[0168] L CLIP (θ) = E t [min(r t (θ)·A t , clip(r t (θ), 1 - ε, 1 + ε)·A t )];

[0169]

[0170] A t = G t - V(s t );

[0171]

[0172] Among them, L CLIP (θ) represents the loss function, E t represents the expectation calculation for all time t and sampled action pairs (s t , a t ), clip(r t (θ), 1 - ε, 1 + ε) represents restricting the range of r t (θ) to be in (1 - ε, 1 + ε), ε represents a hyperparameter, usually taking values ε = 0.1 or ε = 0.2; r tπ(θ) represents the probability ratio of the new and old policies for a given state and sampled action, where πθ(a t |s t ) represents the probability of selecting action a t under the new policy for state st, and πθ old (a t |s t ) represents the probability of selecting action a t under the old policy for state s t ; A t represents the advantage function, V(s t ) represents the state-value function, G t represents the cumulative reward, r t represents the environmental feedback reward, and γ represents the discount factor.

[0173] The loss function L CLIP (θ) constructs the loss by restricting the range of the probability ratio r t (θ) of the new and old policies and combining it with the advantage function A t . While using the advantage function A t to guide the policy to select better sampled actions a, a restricted mechanism is used to ensure the update of the policy. For example, when r t (θ) is within a reasonable range, its product with A t reflects the intensity of adjusting the policy according to the degree of action advantage; once r t (θ) exceeds the defined range, the function clip(r t (θ), 1 - ε, 1 + ε) pulls r t (θ) back to the reasonable interval, avoiding out-of-range updates of the policy, and enabling the policy network to optimize parameters in a direction that considers both the advantage of sampled actions and stability.

[0174] S307. The cloud server receives data from the edge device and preprocesses the data of the edge device to obtain a preprocessing result;

[0175] S308. Input the preprocessing result into a large multi-modal model for inference and analysis to generate a data analysis report, and the data analysis report is sent back to the edge device;

[0176] S309. The edge device collects the model inference feedback information during the data processing process and then uploads the model inference feedback information to the cloud server;

[0177] S310. The cloud server updates the model according to the model inference feedback information and then sends the updated model parameters back to the edge device.

[0178] In this embodiment, steps S307 to S310 are similar to steps S304 to S307 in the foregoing embodiment, and will not be described herein again.

[0179] In this embodiment, by constructing state description information into a multi-dimensional feature vector, inputting the multi-dimensional feature vector into a policy network to obtain a sampling action, and using a specific loss function to optimize the policy network output function. Based on the multi-dimensional feature vector, the policy can comprehensively consider the conditions of data, network, and computing resources, etc., to improve the comprehensiveness of decision-making; on the other hand, the probability ratio limit in the loss function ensures the stable update of the policy, and the introduction of the advantage function guides the policy network to tend to select better actions, improving the decision-making of the policy network.

[0180] Please refer to Figure 4 , Figure 4 which is another embodiment of the large and small model collaboration system based on the cloud link provided by this application. The system includes:

[0181] An acquisition unit, the edge device collects multi-modal data in real time through built-in sensors and performs preprocessing to obtain the preprocessing result of the original data;

[0182] A calculation unit, configured to input the preprocessing result of the original data into a small multi-modal model deployed on the edge device for inference calculation to obtain a preliminary inference result, and input the state description information corresponding to the preliminary inference result into a pre-configured PPO for processing;

[0183] A transmission unit, when the action output by the PPO needs to be uploaded to the cloud, packs the preprocessing result of the original data and the preliminary inference result and transmits them to the cloud server;

[0184] A receiving unit, the cloud server receives data from the edge device and preprocesses the data of the edge device to obtain a preprocessing result;

[0185] An analysis unit, configured to input the preprocessing result into a large multi-modal model for inference analysis to generate a data analysis report, and the data analysis report is transmitted back to the edge device;

[0186] A feedback unit, the edge device collects the model inference feedback information during the data processing process, and then uploads the model inference feedback information to the cloud server;

[0187] An update unit, the cloud server updates the model according to the model inference feedback information, and then transmits the updated model parameters back to the edge device.

[0188] Optionally, the calculation unit includes:

[0189] A classification module, configured to classify the preprocessing result of the original data according to image data and text data;

[0190] An extraction module for extracting features from image data and text data to obtain an image modality and a text modality;

[0191] A splicing and fusion module for performing a splicing and fusion operation on the feature vectors extracted by the image extraction module for the image modality and the text modality to obtain a fused feature vector;

[0192] An input module for inputting the fused feature vector into the softmax layer of the small multi-modal model for final inference calculation.

[0193] The present application also provides a large and small model collaboration device based on a cloud link. Please refer to Figure 5 , Figure 5 which is an embodiment of the large and small model collaboration device based on the cloud link provided by the present application. The device includes:

[0194] A processor 501, a memory 502, an input / output unit 503, and a bus 504;

[0195] The processor 501 is connected to the memory 502, the input / output unit 503, and the bus 504;

[0196] The memory 502 stores a program, and the processor 501 calls the program to execute any one of the large and small model collaboration methods based on the cloud link as described above.

[0197] The present application also relates to a computer-readable storage medium on which a program is stored. When the program runs on a computer, the computer is caused to execute any one of the large and small model collaboration methods based on the cloud link as described above.

[0198] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0199] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, and the indirect coupling or communication connection of the devices or units may be in an electrical, mechanical, or other form.

[0200] The unit described as a separate component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0201] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0202] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

Claims

1. A method for collaborative operation of large and small models based on a cloud link, characterized in that, The method includes: Collecting and preprocessing multi-modal data in real time through sensors built in the edge device to obtain the preprocessing result of the original data; Inputting the preprocessing result of the original data into a small multi-modal model deployed on the edge device for inference calculation to obtain a preliminary inference result, and inputting the status description information corresponding to the preliminary inference result into a pre-configured PPO for processing; When the action output by the PPO needs to be uploaded to the cloud, the preprocessing result of the original data and the preliminary inference result are packaged and sent to the cloud server; The cloud server receives the data from the edge device and preprocesses the data of the edge device to obtain a preprocessing result; Inputting the preprocessing result into a large multi-modal model for inference analysis to generate a data analysis report, and the data analysis report is sent back to the edge device; The edge device collects the model inference feedback information during the data processing process, and then uploads the model inference feedback information to the cloud server; The cloud server updates the model according to the model inference feedback information, and then sends the updated model parameters back to the edge device.

2. The collaborative method for large and small models based on a cloud link according to claim 1, wherein A policy network is pre-configured in the PPO. Inputting the preliminary inference result and the status description information corresponding to the preliminary inference result into the pre-configured PPO for processing, Constructing the status description information into a multi-dimensional feature vector; Inputting the multi-dimensional feature vector into the policy network to obtain the probability distribution π(a|s;θ) output by the policy network, where θ represents the network parameter, s represents the multi-dimensional feature vector, and a represents the sampled action; Obtaining the sampled action according to the probability distribution.

3. The method for collaborative operation of large and small models based on cloud link according to claim 2, wherein The policy network includes an input layer, a hidden layer, and an output layer. Inputting the multi-dimensional feature vector into the policy network to obtain the probability distribution π(a|s;θ) output by the policy network, including: The input layer processes the multi-dimensional feature vector through the following formula: h1 = ReLU(W1×s + b1); ReLU: f(x) = max(0, x); Where h1 represents the output of the input layer, ReLU represents the activation function, W1 represents the weight matrix of the input layer, s represents the multi-dimensional feature vector, and b1 represents the bias term of the input layer; The hidden layer processes the output of the input layer through the following formula: h2 = ReLU(W2×h1 + b2); Where h2 represents the output of the hidden layer, ReLU represents the activation function, W2 represents the weight matrix of the hidden layer, s represents the multi-dimensional feature vector, and b2 represents the bias term of the hidden layer; The output layer generates an action score through the following formula: z = W3×h2 + b3; Where z represents the action score, W3 represents the weight matrix of the output layer, s represents the multi-dimensional feature vector, and b3 represents the bias term of the output layer; Converting the action score into a probability distribution through the following formula: where π(a|s; θ) represents a probability distribution, and Softmax(z i ) represents the probability of the i-th class, and exp(z i ) represents the exponential value of z i , denotes the summation of the exponential values of all z j from j = 1 to k.

4. The method for collaborative operation of large and small models based on cloud link according to claim 3, wherein, The policy network is optimized through the following loss function: L CLIP L(θ) = E t [min(r t (θ)·A t , clip(r t (θ), 1 - ε, 1 + ε)·A t )]; A t = G t - V(s t ); where L CLIP (θ) represents the loss function, and E t denotes the expectation calculation over all time steps t and sampled action pairs (s t , a t ). clip(r t (θ), 1 - ε, 1 + ε) means to limit the range of r t (θ) to (1 - ε, 1 + ε), where ε is a hyperparameter, typically taking values ε = 0.1 or ε = 0.2; r t (θ) represents the probability ratio of the new and old policies for a given state and sampled action, πθ(a t |s t ) represents the probability of selecting action a t under the new policy for state s t , and πθ old (a t |s t ) represents the probability of selecting action a t under the old policy for state s t ; A t represents the advantage function, V(s t ) represents the state-value function, G t represents the cumulative reward, r t represents the environmental feedback reward, and γ represents the discount factor.

5. The method for collaborative operation of large and small models based on cloud link according to claim 1, wherein The preprocessing includes: cleaning, classifying, and featureizing the collected original data.

6. The collaborative method for large and small models based on a cloud link according to claim 1, wherein The small multi-modal model performs preliminary inference through the following steps: Classify the preprocessing results of the original data into image data and text data; Extract features from the image data and the text data to obtain an image modality and a text modality; Perform a splicing and fusion operation on the feature vectors extracted from the image modality and the text modality to obtain a fused feature vector; Input the fused feature vector into the softmax layer of the small multi-modal model for final inference calculation.

7. A large and small model collaboration system based on a cloud link, characterized in that, The system includes: An acquisition unit, where the edge device collects multi-modal data in real time through a built-in sensor and performs preprocessing to obtain the preprocessing results of the original data; A calculation unit, configured to input the preprocessing results of the original data into a small multi-modal model deployed on the edge device for inference calculation to obtain a preliminary inference result, and input the status description information corresponding to the preliminary inference result into a pre-configured PPO for processing; A conveying unit, when the action output by the PPO needs to be uploaded to the cloud, packs the preprocessing results of the original data and the preliminary inference result and conveys them to the cloud server; A receiving unit, where the cloud server receives data from the edge device and preprocesses the data of the edge device to obtain preprocessing results; An analysis unit, configured to input the preprocessing results into a large multi-modal model for inference analysis to generate a data analysis report, and the data analysis report is transmitted back to the edge device; A feedback unit, where the edge device collects model inference feedback information during the data processing process and then uploads the model inference feedback information to the cloud server; An update unit, where the cloud server updates the model according to the model inference feedback information and then transmits the updated model parameters back to the edge device.

8. The collaborative system for large and small models based on a cloud link according to claim 7, wherein The calculation unit includes: A classification module, configured to classify the preprocessing results of the original data into image data and text data; An extraction module, configured to extract features from the image data and the text data to obtain an image modality and a text modality; A splicing and fusion module, configured to perform a splicing and fusion operation on the feature vectors extracted from the image modality and the text modality by the image extraction module to obtain a fused feature vector; An input module, configured to input the fused feature vector into the softmax layer of the small multi-modal model for final inference calculation.

9. A size model collaborative device based on a cloud link, characterized in that, The device includes: A processor, a memory, an input / output unit, and a bus; The processor is connected to the memory, the input / output unit, and the bus; The memory stores a program, and the processor calls the program to execute the method according to any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that, A program is stored on the computer-readable storage medium, and when the program is executed on a computer, it executes the method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Automatic image auditing method and system and computer readable medium

    CN120976714A

  • End side AI model OTA updating optimization method and system based on end-cloud collaboration

    CN122069260A