End side large model evaluation method and device based on artificial intelligence, equipment and medium
Through the cloud-side and end-side evaluation method, the problem that the end-side equipment cannot directly conduct large-scale model evaluation is solved, efficient evaluation and optimization of the end-side model is achieved, and the efficiency and accuracy of the model deployment on the end-side is improved.
Patent Information
- Application Number
- CN202510525927.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing model evaluation framework is mainly used in the cloud, and it is impossible to directly evaluate the end-side devices (such as mobile phones and car computers), and there is a lack of optimization solutions for the end-side model effects.
An artificial intelligence end-side large-scale model evaluation method is designed, through the coordinated work of the cloud and end-side, the cloud module and the end-side module are used to complete the evaluation task together. The specific steps include cloud devices generating evaluation data and transmitting it to the end-side device, end-side device calls local big models for reasoning and returning results, cloud devices analyze the results and generate evaluation reports.
It realizes large-scale model evaluation on end-side devices, solves the problem that end-side devices cannot directly use cloud-side evaluation tools, and improves the efficiency and accuracy of model deployment on end-side through technologies such as dynamic optimization and multimodal evaluation.
Smart Images

Figure CN120069127A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and particularly to an evaluation method, device, equipment, and medium for end-side large models of artificial intelligence. Background Art
[0002] Large models have been widely applied in different fields such as finance, education, and law. However, the increase in model parameters and training data sets has brought more and more uncertainties and new capabilities, which not only pose challenges to model training but also constitute potential risks to the legality of model answers. Therefore, it is necessary to carefully evaluate the capabilities of the model during the training process of large models to ensure their legality and effectiveness. Thus, it can be seen that the evaluation of model effects is very important for improving large models (large language models (LLMs) and large vision language models (LVLMs)). The rapid development of large models requires a lightweight and easy-to-use framework for rapid deployment and evaluation.
[0003] Most of the existing evaluation frameworks are used for cloud model evaluation. When using these evaluation tools on the end side (such as mobile phones or in-vehicle infotainment systems), there are many limitations or even they cannot be used, and the models deployed on the end side cannot directly use these frameworks for evaluation. Therefore, there is a need for an evaluation method for end-side large models. Summary of the Invention
[0004] One or more embodiments of this specification provide an evaluation method, device, equipment, and medium for end-side large models of artificial intelligence to solve the technical problems raised in the background art.
[0005] One or more embodiments of this specification adopt the following technical solutions: An evaluation method for end-side large models of artificial intelligence provided by one or more embodiments of this specification, the method is applied to an evaluation platform, the evaluation platform includes a cloud module and an end-side module, the cloud module includes a cloud device, a cloud evaluation framework, and a cloud AI scoring system, the end-side module includes an end-side device and an end-side model service, and the method includes: The cloud device runs the cloud evaluation framework, generates evaluation data through the cloud evaluation framework, and transmits the evaluation data to the end-side device through the cloud evaluation framework, and the evaluation data includes evaluation input data and evaluation tasks; The end-side device receives the evaluation data through the end-side model service, invokes the end-side large model running locally for inference, obtains a model output result, and transmits the model output result to the cloud device through the end-side model service; The cloud device analyzes the model output result to obtain an evaluation result of the end-side large model; The cloud AI scoring system analyzes the evaluation results in combination with historical evaluation results, generates a scoring analysis result, and generates an improvement strategy adapted to the edge large model based on the scoring analysis result.
[0006] Furthermore, the method further includes: The evaluation platform converts the edge model service into an API service through a preset API module, so that the cloud evaluation framework can call the edge large model through the edge model service for inference.
[0007] Furthermore, the evaluation results include the real-time metrics and task types of the edge large model, and the scoring analysis result is the weight of each dimension; the cloud AI scoring system analyzes the evaluation results in combination with historical evaluation results to generate a scoring analysis result, including: The cloud AI scoring system uses a double-layer LSTM network, inputs the real-time metrics and the task types, and outputs the weights of each dimension.
[0008] Furthermore, the cloud evaluation framework includes an edge dataset processing system. The edge dataset processing system uses intelligent screening and dynamic sampling techniques to extract a representative subset from the dataset of the cloud evaluation framework and sends the subset to the AI scoring system for analysis.
[0009] Furthermore, the edge dataset processing system includes a dataset division module, an error analysis module, and a data integration module. The dataset division module performs dynamic stratified sampling, the error analysis module analyzes the evaluation results to obtain an error result, and the data integration module integrates the error results to obtain an integrated evaluation result, so that the cloud AI scoring system can analyze the integrated evaluation results in combination with historical evaluation results to generate a scoring analysis result.
[0010] Furthermore, the method further includes: Integrate a cross-modal alignment module in the cloud evaluation framework, so that the cross-modal alignment module constructs a multi-modal semantic consistency evaluation index through feature space mapping technology, where the multi-modal semantic consistency evaluation index is used to quantify the semantic deviation value of the cross-modal output result, and the multi-modal at least includes images, speech, and text.
[0011] Furthermore, the method further includes: The cloud AI scoring system collects the resource parameters of the edge device and constructs a dynamic quantization sensitivity matrix; Identify the sensitive layers vulnerable to influence through a pre-established quantization error propagation model; Based on the dependency relationships between the sensitive layers, an error accumulation prediction model is constructed, and an error result is predicted through the error accumulation prediction model; Monitor the output features of the sensitive layer based on the error result: When the error result is greater than a preset error value, corresponding hierarchical calibration is triggered. The hierarchical calibration includes local anti-quantization at the edge side to repair the parameters of the sensitive layer and cloud federated calibration to update the quantization strategy of the sensitive layer.
[0012] An edge-side large model evaluation device based on artificial intelligence provided by one or more embodiments of this specification. The device is applied to an evaluation platform, which includes a cloud module and an edge-side module. The cloud module includes a cloud device, a cloud evaluation framework, and a cloud AI scoring system. The edge-side module includes an edge-side device and an edge-side model service, and includes: An evaluation data generation unit. The cloud device runs the cloud evaluation framework to generate evaluation data through the cloud evaluation framework, and transmits the evaluation data to the edge-side device through the cloud evaluation framework. The evaluation data includes evaluation input data and evaluation tasks; An output result generation unit. The edge-side device receives the evaluation data through the edge-side model service, invokes the edge-side large model running locally for inference to obtain a model output result, and transmits the model output result to the cloud device through the edge-side model service; An output result analysis unit. The cloud device analyzes the model output result to obtain the evaluation result of the edge-side large model; An evaluation result analysis unit. The cloud AI scoring system analyzes the evaluation result in combination with historical evaluation results to generate a scoring analysis result, and generates an improvement strategy adapted to the edge-side large model based on the scoring analysis result.
[0013] An evaluation device for an edge large model based on artificial intelligence provided by one or more embodiments of this specification. The device is applied to an evaluation platform, which includes a cloud module and an edge module. The cloud module includes a cloud device, a cloud evaluation framework, and a cloud AI scoring system. The edge module includes an edge device and an edge model service, and includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to: The cloud device runs the cloud evaluation framework, generates evaluation data through the cloud evaluation framework, and transmits the evaluation data to the edge device through the cloud evaluation framework. The evaluation data includes evaluation input data and evaluation tasks; the edge device receives the evaluation data through the edge model service, invokes the locally running edge large model for inference to obtain a model output result, and transmits the model output result to the cloud device through the edge model service; the cloud device analyzes the model output result to obtain an evaluation result of the edge large model; the cloud AI scoring system analyzes the evaluation result in combination with historical evaluation results, generates a scoring analysis result, and generates an improvement strategy adapted to the edge large model based on the scoring analysis result.
[0014] A non-volatile computer storage medium provided by one or more embodiments of this specification, applied to an evaluation platform, which includes a cloud module and an edge module. The cloud module includes a cloud device, a cloud evaluation framework, and a cloud AI scoring system. The edge module includes an edge device and an edge model service, and stores computer-executable instructions. When the computer-executable instructions are executed by a computer, they can implement: The cloud device runs the cloud evaluation framework, generates evaluation data through the cloud evaluation framework, and transmits the evaluation data to the edge device through the cloud evaluation framework. The evaluation data includes evaluation input data and evaluation tasks; the edge device receives the evaluation data through the edge model service, invokes the locally running edge large model for inference to obtain a model output result, and transmits the model output result to the cloud device through the edge model service; the cloud device analyzes the model output result to obtain an evaluation result of the edge large model; the cloud AI scoring system analyzes the evaluation result in combination with historical evaluation results, generates a scoring analysis result, and generates an improvement strategy adapted to the edge large model based on the scoring analysis result.
[0015] The above at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects: Hardware Heterogeneous Adaptation and Dynamic Optimization Capability: By dynamically matching the evaluation strategies generated by the cloud with the characteristics of the edge-side hardware (such as the computing power curve of the NPU and memory bandwidth), the adaptation problems of traditional evaluation frameworks on edge-side devices (such as mobile phones and in-vehicle devices) are solved. The cloud generates differentiated evaluation tasks based on the fluctuations of device computing power (such as activating 4-bit grouped quantization verification in low-power mode), and the edge side performs lightweight inference according to the local resource status, achieving a dynamic balance between the evaluation process and hardware limitations. This mechanism can cover the instruction set compatibility of different chip architectures (such as Dimensity 9200 and RK3588), improving the generalization ability of the model in edge-side deployment.
[0016] Edge-side Privacy Security and Real-time Guarantee: The evaluation data is inferred locally on the edge side, and user-sensitive information does not need to be transmitted over the network. Combining differential privacy federated calibration technology can achieve full-link data desensitization. At the same time, the lightweight verification module on the edge side (<100KB) can monitor the KL divergence threshold of the model output in real time, ensuring that the quantization error accumulates within a controllable range and avoiding the problem of "model dementia" caused by resource fluctuations. This design not only meets the requirements of high-privacy scenarios such as finance and healthcare but also ensures the low-latency response of the evaluation process.
[0017] Integration of Multimodal and Multidimensional Evaluation Systems: The cloud evaluation framework integrates three core indicators: functionality (text understanding, logical reasoning), energy efficiency ratio (power consumption per unit accuracy), and hardware adaptation (instruction set compatibility). For example, by introducing a multimodal semantic decoding architecture, the alignment ability of the edge-side model in visual-text cross-modal tasks (such as medical image analysis scenarios) can be verified. In addition, combining retrieval-augmented generation (RAG) technology, the evaluation task can cover the knowledge retrieval and generation ability in a network-free environment, enhancing the authenticity of the evaluation scenario.
[0018] Dynamic Feedback and Closed-loop Iteration Mechanism: The cloud AI scoring system generates pruning level suggestions and quantization parameter optimization strategies based on historical evaluation data, forming a "evaluation - compression - verification" closed loop. For example, for the weak links of the edge-side model in mathematical reasoning tasks, the system can direc tionally increase high-order logic training data and update the model parameters through federated learning. At the same time, the dual-parameter backup mechanism (standard / emergency configuration) supports second-level fallback in case of quantization anomalies, ensuring the stability of the evaluation process.
[0019] Lightweight Deployment and Scenario Adaptation: An adaptive multi-dimensional compression algorithm (such as Dongni-AMDC) can be adopted to compress the volume of the evaluation module to the range that can be borne by edge devices (for example, a 1.5B parameter model is adapted to the RK3588 chip), and the accuracy in core scenarios is retained. The evaluation task can be customized according to industry needs. For example, in the medical scenario, an authoritative medical knowledge base is integrated to verify the diagnostic accuracy, or in the education scenario, an interactive question-and-answer evaluation is constructed for the subject knowledge base. This flexibility enables the evaluation framework to quickly adapt to emerging edge devices such as AI PCs and intelligent cockpits. Brief Description of the Drawings
[0020] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings: Figure 1 It is a schematic flowchart of a method for evaluating an edge large model of artificial intelligence provided by one or more embodiments of this specification; Figure 2 It is a schematic structural diagram of an edge evaluation platform provided by one or more embodiments of this specification; Figure 3 It is a schematic structural diagram of the interaction between a client and an edge model service provided by one or more embodiments of this specification; Figure 4 It is a schematic structural diagram of a device for evaluating an edge large model of artificial intelligence provided by one or more embodiments of this specification; Figure 5 It is a schematic structural diagram of a device for evaluating an edge large model of artificial intelligence provided by one or more embodiments of this specification. Detailed Embodiments
[0021] The embodiments of this specification provide a method, device, device, and medium for evaluating an edge large model of artificial intelligence.
[0022] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only some embodiments of this specification, rather than all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.
[0023] Figure 1A schematic flowchart of an end - side large - model evaluation method for artificial intelligence provided by one or more embodiments of this specification. This process can be executed by an end - side large - model evaluation system. Some input parameters or intermediate results in the process allow manual intervention and adjustment to help improve accuracy.
[0024] It should be noted that the end - side large - model evaluation method in the embodiments of this specification can be applied to an evaluation platform. The evaluation platform includes a cloud module and an end - side module. The cloud module includes a cloud device, a cloud evaluation framework, and a cloud AI scoring system. The end - side module includes an end - side device and an end - side model service. The method flow steps of the embodiments of this specification are as follows: S101, the cloud device runs the cloud evaluation framework, generates evaluation data through the cloud evaluation framework, and transmits the evaluation data to the end - side device through the cloud evaluation framework. The evaluation data includes evaluation input data and evaluation tasks.
[0025] In the embodiments of this specification, regarding the construction of the above - mentioned cloud evaluation framework and data generation (S101), the following specific implementation schemes can be adopted: The embodiments of this specification can generate adaptable evaluation data based on the type of end - side device (such as smartphones, in - vehicle devices, etc.) and hardware status (memory, computing power) by using dynamic stratified sampling technology. For example: Multi - modal scenarios: For visual - speech - text joint instruction tasks, cross - modal samples (such as the instruction "take a picture of a document and read it aloud") are extracted proportionally, and semantic consistency is aligned through feature - space mapping.
[0026] Resource - aware strategy: Combining the remaining battery power of the device and the NPU computing power, dynamically adjust the data complexity (such as preferentially generating lightweight text tasks in low - power mode).
[0027] The embodiments of this specification can adopt local feature extraction technology and only upload structured semantic vectors (such as 128 - dimensional visual features) to meet the requirements of privacy protection.
[0028] In addition, based on the capabilities of the cloud testing platform, the embodiments of this specification can also support automated task - flow configuration: including multi - dimensional indicators such as functional testing (instruction response accuracy), energy efficiency ratio (power consumption per unit accuracy), and hardware adaptation (instruction set compatibility). At the same time, according to the real - time load of the end - side device (such as when the memory occupancy rate > 80%, delay the issuance of high - computing - power tasks), the evaluation data is encrypted and transmitted through the MQTT protocol.
[0029] S102, the end - side device receives the evaluation data through the end - side model service, invokes the end - side large - model running locally for inference, obtains the model output result, and transmits the model output result to the cloud device through the end - side model service.
[0030] In the embodiments of this specification, regarding the above-mentioned end-side model inference optimization (S102), the following specific implementation solutions can be adopted: 1. Deployment of lightweight inference engine Mixed-precision computing: An FP16 (visual encoding), INT8 (speech processing), INT4 (text generation) mixed pipeline can be adopted, which can reduce the memory occupancy by more than 40%.
[0031] Resource-aware scheduling: By dynamically routing selection (such as disabling multimodal parallel computing when the battery is low), combined with spatio-temporal interleaved processing technology (key frame extraction + differential caching), the decoding delay can be optimized to the 50ms level.
[0032] 2. Localized fault tolerance mechanism Real-time monitoring of sensitive layers: Based on the quantization error propagation model, highly sensitive modules such as the Attention layer and residual connections are identified. When the error accumulation exceeds the threshold, end-side dequantization repair is triggered (such as restoring FP16 precision parameters).
[0033] Memory block management: Dynamic probing technology is used to identify redundant features, and storage allocation is optimized through the CSI interface, reducing the peak memory requirement by 64%.
[0034] S103. The cloud device analyzes the model output result to obtain the evaluation result of the end-side large model.
[0035] In the embodiments of this specification, regarding the above-mentioned evaluation result analysis and feedback (S103), the following specific implementation solutions can be adopted: 1. Multi-dimensional index calculation Functional evaluation: The quality of text generation can be quantified through ROUGE / BLEU metrics, and cross-modal ablation tests are combined to locate the source of semantic deviation (such as the misalignment rate between voice commands and visual understanding).
[0036] Energy efficiency ratio analysis: Calculate the power consumption curve per unit semantic accuracy (mW / accuracy percentage point) to identify high-energy-consuming modules (such as the multimodal joint inference layer).
[0037] 2. Dynamic threshold management Task scenario adaptation: For high-precision demand scenarios such as medical diagnosis, strict deviation thresholds are set (such as KL divergence < 0.3%); for smart home scenarios, it can be relaxed to 1.5% to improve the response speed.
[0038] Abnormal fuse mechanism: When detecting semantic mutations in financial instructions or hardware abnormalities (memory fragmentation rate > 30%), the evaluation process is immediately terminated and the local security protocol is activated.
[0039] S104. The cloud AI scoring system analyzes the evaluation results in combination with historical evaluation results, generates a scoring analysis result, and generates an improvement strategy adapted to the end-side large model based on the scoring analysis result.
[0040] In the embodiments of this specification, regarding the above-mentioned closed-loop optimization strategy generation (S104), the following specific implementation solutions can be adopted: 1. Federated learning-driven iteration Parameter encryption and aggregation: Through homomorphic encryption technology, aggregate the quantization strategies of the sensitive layers of multiple devices (such as scaling factors, truncation thresholds), and update the global model library.
[0041] Incremental model compression: Based on the historical error distribution, generate dynamic pruning suggestions (such as adjusting the weight pruning rate of the Attention layer), and adapt to different end-side chip architectures (Dimensity 9200, RK3588, etc.).
[0042] 2. End-cloud collaborative deployment optimization Elastic inference engine matching: Automatically select the optimal inference framework (TensorFlowLite / MNN) according to the device chip type to improve the instruction set compatibility.
[0043] Industry scenario customization: Orientally optimize the decoding delay threshold for the autonomous driving scenario, and strengthen multi-modal hallucination suppression in the medical scenario (RAG enhancement + contrast learning verification).
[0044] It should be noted that through the above content, the embodiments of this specification have the following beneficial effects: Hardware heterogeneous adaptation and dynamic optimization capabilities: Dynamically match the evaluation strategies generated by the cloud with the characteristics of the end-side hardware (such as the NPU computing power curve, memory bandwidth), solving the adaptation problems of traditional evaluation frameworks for end-side devices (such as mobile phones, car machines, etc.). The cloud generates differentiated evaluation tasks based on the device computing power fluctuations (such as activating 4-bit grouped quantization verification in the low battery mode), and the end-side performs lightweight inference according to the local resource status, achieving a dynamic balance between the evaluation process and hardware limitations. This mechanism can cover the instruction set compatibility of different chip architectures (such as Dimensity 9200, RK3588), and improve the generalization ability of the model in end-side deployment.
[0045] End-side privacy security and real-time guarantee: The evaluation data completes inference locally on the end-side, and user-sensitive information does not need to be transmitted over the network. Combining differential privacy federated calibration technology can achieve full-link data desensitization. At the same time, the lightweight verification module on the end-side (<100KB) can monitor the KL divergence threshold of the model output in real time to ensure that the quantization error accumulates within a controllable range and avoid the "model dementia" problem caused by resource fluctuations. This design not only meets the requirements of high-privacy scenarios such as finance and medical care, but also guarantees the low-latency response of the evaluation process.
[0046] Integration of Multimodal and Multidimensional Evaluation Systems: The cloud evaluation framework integrates three core indicators: functionality (text understanding, logical reasoning), energy efficiency ratio (power consumption per unit of precision), and hardware adaptability (instruction set compatibility). For example, by introducing a multimodal semantic decoding architecture, the alignment ability of the edge-side model in visual-text cross-modal tasks (such as in medical image analysis scenarios) can be verified. In addition, by combining Retrieval-Augmented Generation (RAG) technology, the evaluation tasks can cover knowledge retrieval and generation capabilities in offline environments, enhancing the authenticity of the evaluation scenarios.
[0047] Dynamic Feedback and Closed-Loop Iteration Mechanism: The cloud AI scoring system generates pruning level suggestions and quantization parameter optimization strategies based on historical evaluation data, forming a "evaluation - compression - verification" closed loop. For example, for the weak links of the edge-side model in mathematical reasoning tasks, the system can specifically increase high-order logic training data and update the model parameters through federated learning. At the same time, the dual-parameter backup mechanism (standard / emergency configuration) supports second-level fallback in case of quantization anomalies, ensuring the stability of the evaluation process.
[0048] Lightweight Deployment and Scenario Adaptation: An adaptive multi-dimensional compression algorithm (such as Dongni-AMDC) can be used to compress the volume of the evaluation module to the range that can be borne by edge-side devices (for example, a 1.5B parameter model is adapted to the RK3588 chip), and the accuracy in core scenarios is retained. The evaluation tasks can be customized according to industry needs. For example, in the medical scenario, an authoritative medical knowledge base is integrated to verify the diagnostic accuracy, or in the education scenario, an interactive Q&A evaluation is constructed for subject knowledge bases. This flexibility enables the evaluation framework to quickly adapt to emerging edge-side devices such as AI PCs and intelligent cockpits.
[0049] Furthermore, the evaluation platform can convert the edge-side model service into an API service through a preset API module, so that the cloud evaluation framework can call the edge-side large model through the edge-side model service for inference.
[0050] It should be noted that through the above content, the embodiments of this specification have the following beneficial effects: Improved Cross-Platform Compatibility and Heterogeneous Hardware Adaptability: By encapsulating the edge-side model service through a standardized API interface, the underlying hardware differences (such as differences in Qualcomm NPU and Intel CPU instruction sets) are shielded, enabling seamless migration of evaluation tasks across multiple terminal devices (Android / iOS phones, in-vehicle infotainment systems).
[0051] Dynamic Resource Scheduling and Real-Time Performance Optimization: The API service layer integrates a computing power awareness module to monitor the NPU utilization rate, memory occupancy, and remaining battery power of edge-side devices in real time, and dynamically adjust the computing power allocation strategy for evaluation tasks.
[0052] Furthermore, the evaluation results may include the real-time metrics and task types of the edge large model, and the scoring analysis result is the weight of each dimension; when the cloud AI scoring system analyzes the evaluation results in combination with historical evaluation results to generate the scoring analysis result, the cloud AI scoring system may use a double-layer LSTM network, input the real-time metrics and the task types, and output the weights of each dimension.
[0053] It should be noted that through the above content, the embodiments of this specification have the following beneficial effects: Temporal feature modeling and dynamic weight allocation ability: The double-layer LSTM network captures the fluctuation rules of the real-time metrics of the edge large model (such as sudden increase in latency, computing power fluctuation) through temporal memory units (such as forget gate, input gate), and dynamically adjusts the weights of each dimension in combination with task type features (such as text generation, logical reasoning). For example, in a mathematical reasoning task, the weight of logical accuracy is automatically enhanced, while in a content creation task, the focus is on the output coherence index. This mechanism solves the problem of the disconnection between traditional static weight allocation and task scenarios.
[0054] Modeling of long-term dependence and error propagation: The long short-term memory characteristics of LSTM can effectively track the error accumulation effect in historical evaluation results (such as cross-layer propagation of quantization errors), and transfer the statistical characteristics of historical evaluation data through hidden states to predict the impact of current weights on model stability. For example, when it is detected that the activation distribution of the edge model shifts during continuous inference, the system automatically reduces the hardware adaptation weight to avoid misjudgment.
[0055] Lightweight computing and edge-cloud collaborative optimization: Compared with the traditional Transformer architecture, the number of parameters of the double-layer LSTM network is reduced to about 1 / 70, meeting the real-time requirements of the cloud AI scoring system. Its hierarchical structure supports task-level feature decoupling: the bottom LSTM processes real-time hardware metrics (such as NPU utilization rate, memory occupancy rate), and the top LSTM fuses task semantic features (such as syntax compliance of programming tasks), realizing the collaborative optimization of edge-side resource status and task goals.
[0056] Furthermore, the cloud evaluation framework includes an edge-side dataset processing system. The edge-side dataset processing system uses intelligent screening and dynamic sampling techniques to extract a representative subset from the dataset of the cloud evaluation framework, and sends the subset to the AI scoring system for analysis.
[0057] It should be noted that through the above content, the embodiments of this specification have the following beneficial effects: Efficient processing of massive data and real-time response capabilities: By using dynamic sampling techniques (such as continuous sampling, stratified sampling, probability proportional sampling) to extract representative subsets from cloud datasets, and combining intelligent screening algorithms (such as feature screening based on decision trees and deep learning), the system can quickly complete data preprocessing and feature extraction. This combination of techniques solves the latency problem of traditional full-scale exploration, supports second-level update of evaluation results, and at the same time records user interaction operations such as data filtering and sorting through an operation stack to achieve real-time feedback and dynamic adjustment, significantly improving the timeliness of end-side large model evaluation.
[0058] Multi-dimensional data quality assessment and intelligent decision support: The intelligent screening system integrates quality indicators such as data integrity, consistency, and diversity, and the dynamic sampling technique verifies the contribution of different data subsets to model performance through data ablation testing. For example, when verifying the logical reasoning ability of the end-side model, the system can directionally extract high-complexity samples (such as mathematical proof questions or legal provision analysis cases), and combine indicators such as the model training accuracy and resource consumption of the AI scoring system to generate targeted improvement strategies (such as increasing training data in specific fields or adjusting the model quantization bit width).
[0059] Dynamic adaptation to end-side task scenarios and resource limitations: The system dynamically adjusts the following sampling strategies according to the hardware characteristics of the end-side device (such as NPU computing power and memory capacity): Low computing power scenario: Adopt lightweight random sampling and preferentially extract low-dimensional feature data subsets; High-precision requirement scenario: Enable stratified sampling to ensure the coverage density of key data categories (such as abnormal transaction records in financial risk control); Real-time requirement scenario: Quickly capture changes in data distribution through probability proportional sampling (PPS) to support the rapid iteration of the end-side model in a dynamic environment (such as road condition recognition in a vehicle-mounted system).
[0060] Furthermore, the end-side dataset processing system includes a dataset division module, an error analysis module, and a data integration module. The dataset division module performs dynamic stratified sampling, the error analysis module analyzes the evaluation results to obtain error results, and the data integration module integrates the error results to obtain the integrated evaluation results, so that the cloud AI scoring system can analyze the integrated evaluation results in combination with historical evaluation results to generate a scoring analysis result.
[0061] It should be noted that the embodiments of this specification have the following beneficial effects through the above content: Dynamic Stratified Sampling Improves the Adaptability of Evaluation Scenarios: The dataset partitioning module adopts dynamic stratified sampling technology. By real-time sensing the hardware status of edge devices (such as NPU computing power fluctuations, memory occupancy) and task type characteristics (such as text generation, multimodal reasoning), the sampling strategy is dynamically adjusted. For example: Resource-sensitive scenarios: Automatically switch to lightweight stratified sampling in low-power mode, and preferentially extract low-dimensional feature subsets to adapt to the computing power limitations of edge devices; High-precision demand scenarios: For tasks such as financial risk control and medical diagnosis, directionally extract high-complexity samples (such as abnormal transaction records, medical imaging data) to ensure the coverage density of key data categories; Multimodal scenarios: Capture changes in cross-modal data distribution through stratified probability proportional sampling (PPS) to support the dynamic verification of image-text joint reasoning tasks.
[0062] This mechanism breaks through the problem of scenario rigidity in traditional static sampling, enabling the evaluation data to be deeply coupled with the real-time status and task requirements of edge devices, and greatly improving the coverage rate.
[0063] Error Propagation Modeling and Real-time Fault Tolerance Mechanism: The error analysis module uses multi-level error tracing technology to identify anomalies such as quantization error accumulation and hardware resource fluctuations during the inference process of edge models, and constructs an error propagation path model. For example: Quantization error tracking: Monitor the quantization parameter offset based on the KL divergence threshold, and predict the impact of errors on model stability combined with historical evaluation data; Hardware anomaly fusing: When detecting a sudden drop in NPU computing power or an excessive memory fragmentation rate, trigger an emergency weight configuration switch (such as reverting from int4 to int8 quantization) to ensure the stability of the evaluation process; Cross-modal error correlation analysis: For vision-text joint tasks, locate the alignment errors between image encoders and language models, and optimize the attention mechanism of visual encoders such as SigLip.
[0064] This module provides an interpretable error map for the dynamic tuning of edge models, significantly reducing the misjudgment rate in abnormal scenarios.
[0065] Multi-source Data Fusion and Federated Learning Collaboration: The data integration module uses heterogeneous data alignment technology to perform cross-modal fusion of error analysis results with edge hardware logs (such as energy consumption curves, instruction set compatibility reports) and task semantic labels (such as legal compliance verification, code generation quality). Its core values include: Unified evaluation index system: Integrate three dimensions of functionality (semantic understanding), energy efficiency ratio (power consumption per unit precision), and hardware adaptability (chip instruction set compatibility) to build an evaluation framework for end-cloud collaboration; Enhanced Privacy and Security: By adopting the encrypted aggregation technology of federated learning, desensitization and feature extraction of error data are completed locally, and only the structured analysis results (such as error type distribution, confidence interval) are uploaded to the cloud; Dynamic Knowledge Base Construction: Based on historical error data, pruning suggestions (such as weight pruning of the Attention layer) and a quantization parameter optimization strategy library are generated to support online incremental learning of the edge-side model.
[0066] Furthermore, a cross-modal alignment module can be integrated into the cloud evaluation framework so that the cross-modal alignment module constructs a multi-modal semantic consistency evaluation index through feature space mapping technology, where the multi-modal semantic consistency evaluation index is used to quantify the semantic deviation value of the cross-modal output result, and the multi-modal at least includes images, speech, and text.
[0067] It should be noted that the above content can be implemented through the following specific implementation plans: I. Cross-modal Semantic Space Construction and Alignment Module Integration 1. Multi-modal Feature Space Mapping Based on the Align-Anything framework, using the CLIP-style contrastive learning algorithm, map the semantic representations of images, speech, and text to a unified high-dimensional space. For example: Visual modality: Extract more than 200 visual feature points such as image texture and color through ResNet-152; Text / speech modality: Use the BERT or Whisper model to extract semantic vectors, and eliminate modality heterogeneity through adversarial alignment technology.
[0068] Dynamic weight adjustment: Introduce a hierarchical attention mechanism to dynamically allocate the semantic weights of each modality according to task requirements (such as high precision in medical scenarios), and solve the semantic deviation problem of traditional static mapping.
[0069] 2. Modular Integrated Design Decoupled architecture: Refer to the three-module design of DataLoader-Generator-Evaluator in Align-Anything: DataLoader: Load the multi-modal evaluation data set (such as the Align-Anything-400k full-modal data set), and complete data cleaning and format unification; Generator: Support multiple inference backends such as Deepspeed (multi-modal understanding) and Accelerate (multi-modal generation), and adapt to different model structures; Evaluator: Quantify cross-modal output deviation based on semantic consistency evaluation metrics (such as KL divergence, cosine similarity).
[0070] II. Design of Multimodal Semantic Consistency Evaluation Metrics 1. Construction of Deviation Quantification Model Cross-modal Ablation Test: By constructing a visual-voice-text joint ablation matrix, locate the source of semantic deviation (such as the misalignment rate between voice commands and visual understanding).
[0071] Dynamic Threshold Management: Set deviation tolerance thresholds according to task types (such as KL divergence < 0.3% required in medical diagnosis scenarios), and trigger an error correction mechanism when the threshold is exceeded.
[0072] 2. Refined Evaluation Dimensions Functional Metrics: Instruction Following Accuracy: Evaluate the response consistency of cross-modal instructions (such as the semantic matching degree between text and image in the task of "take a picture of a document and read it aloud"); Intention Understanding Coherence: Verify the synchronization between voice commands and touch screen operations (such as the timing matching between the command of "zoom in on the photo" and the image zooming action).
[0073] Credibility Metrics: Through contrastive learning, train a cross-modal consistency discriminator to suppress multimodal hallucinations (such as contradictions between text generation and image content).
[0074] III. Edge-Cloud Collaborative Resource Optimization and Dynamic Scheduling 1. Lightweight Inference Engine Mixed Precision Pipeline: Adopt a mixed computing strategy of FP16 (visual encoding), INT8 (speech processing), and INT4 (text generation) to reduce memory occupancy by more than 40%; Spatio-Temporal Interleaved Processing: Extract key frames (every 5 frames) + differential frame caching from the video stream, and optimize the voice and visual response latency to the 50ms level.
[0075] 2. Dynamic Resource Scheduling Hardware-Aware Routing: Automatically switch alignment strategies according to the real-time status of the device (such as NPU computing power, memory occupancy rate) (prioritize voice-text alignment in low-power mode); Elastic Computing Adaptation: The cloud realizes elastic expansion through Alibaba Cloud GN7 instances, supporting more than 200 concurrent evaluation tasks.
[0076] IV. Closed-Loop Optimization and Security Compliance Mechanisms 1. Federated Learning-Driven Iteration Parameter Encrypted Aggregation: Adopt homomorphic encryption technology to aggregate sensitive layer quantization strategies of multiple devices (such as scaling factors, truncation thresholds), and update the global model library; Incremental pruning: Generate dynamic pruning suggestions based on historical error distribution (such as adjusting the weight pruning rate of the Attention layer) to adapt to different edge chip architectures.
[0077] 2. Privacy and security architecture Localized feature desensitization: The original image / voice data completes feature extraction at the edge side, and only structured semantic vectors (such as 128-dimensional visual embeddings) are uploaded. Emergency fuse mechanism: When detecting anomalies such as semantic mutations in financial transaction instructions, immediately terminate the data stream and activate the local security protocol.
[0078] V. Deep adaptation to industry scenarios 1. Medical diagnosis scenario Directly monitor the decoder layer of the medical image segmentation model, and supplement the authoritative knowledge base data through the RAG technology to ensure sub-pixel accuracy in tumor edge detection.
[0079] 2. Smart home scenario Optimize the temperature control curve for multi-modal parallel processing, balance real-time performance and energy consumption, and support complex instructions such as "adjust the air conditioner temperature and voice broadcast".
[0080] 3. Financial risk control scenario Implement hierarchical calibration for the LSTM time series prediction layer, and dynamic dequantization to ensure the trading response accuracy at the microsecond level, with the semantic mutation detection error < 0.1%.
[0081] It should be noted that the embodiments of this specification have the following beneficial effects through the above content: Breakthrough in multi-modal semantic consistency modeling ability: The cross-modal alignment module projects the semantic representations of images, voices, and texts into a unified high-dimensional space through feature space mapping technology to achieve joint evaluation of cross-modal outputs. For example: Dynamic feature space calibration: Based on the improved SigLIP architecture, use a hierarchical attention mechanism to dynamically adjust the semantic weights of different modalities to solve the semantic gap problem caused by modality heterogeneity in traditional methods.
[0082] Quantization bias threshold management: Introduce a semantic bias value calculation model driven by KL divergence. When the multi-modal output deviation exceeds the preset threshold (such as < 0.3% required in the medical scenario), automatically trigger an error correction mechanism to ensure high reliability of instruction response.
[0083] Dynamic evaluation optimization with edge-side resource awareness: Combining the hardware characteristics of smartphones, the module achieves coordinated balance of computing power - accuracy - latency. For example: Lightweight Alignment Engine: Adopts a mixed-precision pipeline (visual encoding FP16, speech processing INT8, text generation INT4), reducing memory occupancy by 58% and adapting to the computing power limitations of mobile NPU; Dynamic Routing Selection: Based on the real-time status of the device (such as remaining battery power, memory occupancy rate), automatically switches cross-modal alignment strategies (low-power mode prioritizes speech-text alignment); Spatio-temporal Interleaved Processing: For video streams, uses key frame extraction + differential frame caching technology to synchronously optimize the voice command and visual response delay to the 50ms level.
[0084] Multi-modal Hallucination Suppression and Trust Enhancement: Aiming at the strong interaction requirements in the smartphone scenario, the module integrates a dual anti-hallucination mechanism. For example: Visual Anchor Verification: When generating a text response, it forcibly verifies the semantic correlation between visual features and voice commands. If the confidence level is lower than the threshold (e.g., >95% is required in the medical diagnosis scenario), it automatically calls the Retrieval-Augmented Generation (RAG) technology to supplement authoritative knowledge base data; Cross-modal Fact Checking: Trains a cross-modal consistency discriminator through Contrastive Learning, reducing the misrecognition rate.
[0085] Industry Scenario Deep Adaptation Ability: The module supports the customization of multi-dimensional evaluation metrics to meet the complex application requirements of smartphones. For example: Functional Metrics: Visual-voice command response accuracy (such as in the "Take a picture of a document and read the content" task); Cross-modal Intent Understanding Coherence (such as the synchronization between the voice command "Zoom in on this photo" and touch screen operations); Energy Efficiency Ratio Metrics: Power consumption per unit semantic accuracy (milliwatts / accuracy percentage point), temperature control curve during multi-modal parallel processing.
[0086] Hardware Adaptation Metrics: Instruction set compatibility of heterogeneous computing units (CPU / GPU / NPU).
[0087] Furthermore, the cloud AI scoring system collects the resource parameters of the end-side device and constructs a dynamic quantization sensitivity matrix; Identifies vulnerable sensitive layers through a pre-established quantization error propagation model; Based on the dependency relationships between the sensitive layers, constructs an error accumulation prediction model, and predicts an error result through the error accumulation prediction model; Monitors the output features of the sensitive layer based on the error result: When the error result is greater than a preset error value, triggers corresponding hierarchical calibration, and the hierarchical calibration includes local anti-quantization repair of the parameters of the sensitive layer at the end side and federated calibration update of the quantization strategy of the sensitive layer in the cloud.
[0088] It should be noted that the above content can be implemented through the following specific implementation plans: I. Construction of Dynamic Quantization Sensitivity Matrix 1. Collection and Modeling of Edge-Side Resource Parameters Multi-dimensional parameter monitoring: Real-time collection of hardware state parameters such as NPU computing power fluctuations, memory fragmentation rate, and remaining battery power of edge-side devices, and construction of a dynamic quantization sensitivity matrix in combination with the characteristics of the model structure (such as weight distribution and activation value range).
[0089] Hierarchical sensitivity mapping: Based on the improved SZZ algorithm and KD-Tree clustering technology, map device resource parameters to quantization sensitivity levels (such as high-sensitivity layer: Attention module; low-sensitivity layer: pooling layer) to achieve hardware-task co-perception.
[0090] 2. Dynamic Identification of Sensitive Layers Error propagation path analysis: Through Monte Carlo sampling and graph neural networks, quantify the non-linear impact of errors in each layer on the model output (such as the exponential amplification effect of convolutional layer errors on subsequent pooling layers), and generate a dependency graph.
[0091] Hardware perception priority ranking: Combine the real-time load of the device (such as giving priority to detecting computationally intensive layers when the NPU utilization rate > 80%), and dynamically adjust the sensitive layer identification order to reduce resource waste caused by mixed precision.
[0092] II. Error Accumulation Prediction and Hierarchical Fault Tolerance Mechanism 1. Cross-Layer Error Correlation Modeling Temporal error propagation simulation: Use sliding window technology and incremental learning algorithms to capture the dynamic changes of errors between sensitive layers over inference time (such as error accumulation drift caused by residual connections), and construct an error accumulation prediction model based on KL divergence.
[0093] Dynamic threshold management: Automatically adjust the error warning threshold according to task scenario requirements (medical diagnosis requires KL divergence < 0.3%; smart home can be relaxed to 1.5%) and device status (low battery mode) to achieve a flexible balance between accuracy and energy efficiency.
[0094] 2. Edge-Cloud Collaborative Calibration Strategy Edge-side local dequantization repair: When the error exceeds the threshold, only perform FP16 precision recovery on sensitive layers (such as the Attention layer) to avoid a sharp increase in computing power caused by full-model fallback, and reduce memory occupancy by more than 40%.
[0095] Cloud Federal Calibration Update: Aggregate quantization strategies (such as scaling factors, truncation thresholds) of multi-device sensitive layers through homomorphic encryption, and use Bayesian optimization to dynamically update the global quantization parameter library, reducing the communication data volume by 90%.
[0096] III. Lightweight and Real-time Optimization 1. Hybrid Precision Pipeline Design Hierarchical Quantization Strategy: Retain FP16 precision for highly sensitive layers (such as the bounding box regression layer), and compress non-critical layers (such as fully connected layers) to INT8 / INT4, reducing memory occupancy by 58%.
[0097] Spatial-temporal Interleaved Processing: For video streams, adopt key frame extraction (every 5 frames) + differential frame caching technology, and synchronously optimize the voice command and visual response delay to the 50ms level.
[0098] 2. Resource-aware Dynamic Scheduling Elastic Computing Adaptation: Automatically switch the quantization mode based on the real-time device status (memory occupancy rate > 80%) (low-power mode prioritizes voice-text alignment), and support policy synchronization in weak network environments.
[0099] IV. Security Compliance and Industry Adaptation 1. Privacy Protection Mechanism Localized Feature Desensitization: Extract structured semantic vectors (such as 128-dimensional visual embeddings) from the original image / voice data on the edge side, and only upload the encrypted quantization parameters.
[0100] Emergency Fuse Mechanism: When detecting a semantic mutation in financial transaction instructions, immediately terminate the data stream and activate the local security protocol, meeting the requirements of regulations such as GDPR.
[0101] 2. Deep Adaptation to Industry Scenarios Medical Diagnosis Scenario: Directly monitor the decoder layer of the medical image segmentation model, and supplement authoritative knowledge base data through RAG technology to ensure sub-pixel accuracy in tumor edge detection.
[0102] Financial Risk Control Scenario: Implement hierarchical calibration for the LSTM time series prediction layer, and dynamically dequantize to ensure microsecond-level trading response accuracy, with the semantic mutation detection error < 0.1%.
[0103] It should be noted that through the above content, the embodiments of this specification have the following beneficial effects: Precise Location of Quantization Sensitive Layers and Dynamic Resource Adaptation: By constructing a dynamic quantization sensitivity matrix, the system can real-time perceive the hardware status of edge devices (such as NPU computing power fluctuations, memory fragmentation rate) and model structure characteristics (such as weight distribution, activation value range), and dynamically identify neural network layers sensitive to quantization errors. For example: Hardware-aware sensitivity assessment: Combined with the computing power and memory limitations of the device, it prioritizes detection of layers with high computational intensity (such as the Attention layer) or layers with low-precision quantization that are prone to mismatch (such as the bounding box regression layer), reducing hardware efficiency loss caused by mixed precision.
[0104] Error propagation path modeling: Based on the quantized error propagation model, the dependencies between sensitive layers (such as residual connections and cross-layer attention) are analyzed to predict the nonlinear impact of error accumulation on model output, thus avoiding the systematic bias caused by traditional single-layer analysis.
[0105] Error accumulation prediction and hierarchical fault tolerance mechanism: The error accumulation prediction model enhances the stability of end-side reasoning through the following mechanisms: Cross-layer error correlation analysis: Construct a dependency graph based on a graph neural network to capture the error propagation effect between sensitive layers (such as the exponential amplification effect of the quantization error of the convolutional layer on the accuracy of the subsequent pooling layer).
[0106] Dynamic threshold management: Dynamically adjust the preset error threshold according to the task type (such as medical diagnosis requires high precision) and hardware status (such as low power mode) to achieve a flexible balance between accuracy and efficiency.
[0107] End-cloud collaborative fault tolerance: When the error exceeds the threshold, it triggers local dequantization repair on the end (such as restoring the key weights of FP16 precision) and cloud-side federated calibration (aggregating the error distribution of multiple devices to update the quantization strategy library), forming a closed-loop optimization ecosystem.
[0108] Lightweight and real-time optimization: The system significantly reduces resource consumption through a hierarchical calibration strategy, such as: Local repair on the client side: Dequantization is performed only on sensitive layers (such as restoring the FP16 precision of the Attention layer) to avoid a surge in computing power caused by full model rollback, reducing memory usage by more than 40%.
[0109] Incremental federated learning: The cloud only aggregates the quantitative policy parameters of the sensitive layer (such as scaling factor and truncation threshold), reducing the amount of communication data by 90% and supporting policy synchronization in weak network environments.
[0110] Multi-scenario adaptive capabilities, such as the following scenarios: Financial risk control scenario: For transaction fraud detection models, the quantization error accumulation of the LSTM time series prediction layer is monitored in a targeted manner, and dynamic dequantization is triggered to ensure microsecond-level response accuracy.
[0111] Medical imaging scenario: Implement hierarchical calibration on the decoder layer of the medical image segmentation model to ensure sub-pixel accuracy requirements for tumor edge detection.
[0112] Autonomous driving scenario: Based on the sensor data stream, the monitoring frequency of the sensitive layer of the target detection model is adjusted in real time to balance real-time performance and energy consumption.
[0113] It should be noted that large models have been widely applied in different fields such as finance, education, and law. However, the increase in model parameters and training datasets has brought more and more uncertainties and new capabilities, which not only pose challenges to model training but also constitute potential risks to the legality of model answers. Therefore, it is necessary to carefully evaluate the capabilities of the large model during the training process to ensure its legality and effectiveness. Thus, the evaluation of model performance is very important for improving large models (Large Language Models (LLMs) and Large Vision-Language Models (LVLMs)). The rapid development of large models requires a lightweight and easy-to-use framework for rapid deployment and evaluation. Common means of evaluating large models include: 1. Opencompass (https: / / github.com / open-compass / opencompass): Deeply and comprehensively evaluate large language models using over a hundred selected evaluation sets in eight core ability dimensions, namely language understanding, knowledge accuracy, logical reasoning, creative generation, math problem-solving, code writing, long text processing, and agent interaction.
[0114] 2. UltraEval (https: / / github.com / OpenBMB / UltraEval): UltraEval is an open-source framework for evaluating the capabilities of large models, providing a lightweight and easy-to-use evaluation system that supports the evaluation of the capabilities of mainstream large models.
[0115] 3. VLMEvalKit (https: / / github.com / open-compass / VLMEvalKit): Based on the VLMEvalKit open-source evaluation framework, evaluate the comprehensive performance of LVLMs on dozens of open-source evaluation sets in a pure generation manner.
[0116] However, the above evaluation frameworks are all used for cloud model evaluation. When using these evaluation tools on the edge side (such as mobile phones or in-vehicle devices), there are many limitations or even they cannot be used (such as the inability to install the Python environment on the edge side). Models deployed on the edge side cannot directly use these frameworks for evaluation. At the same time, there is also a lack of optimization solutions for the performance of edge-side models. Previous evaluation frameworks often only give an evaluation result without analyzing the result.
[0117] Figure 2 Provides a schematic diagram of the structure of an edge-side evaluation platform. This evaluation platform mainly consists of two parts, namely the cloud module and the edge-side module.
[0118] 1. Initialization and Deployment of Cloud Devices Cloud Devices: Such as personal computers, host servers, etc., which can run the cloud evaluation framework.
[0119] The cloud evaluation framework generates evaluation data (such as input text, parameter configuration), and transmits this data to the edge device through a communication path (such as HTTP).
[0120] 2. Edge Device Starts Service Edge Device: Starts the model service through the model_api module and receives tasks sent by the cloud evaluation framework.
[0121] After startup, the edge model service waits for specific evaluation data from the cloud.
[0122] 3. Transmission and Processing of Evaluation Data The cloud device processes and transmits the evaluation data through the cloud evaluation framework, and sends the data required for evaluation (such as model input, task configuration) to the edge model service.
[0123] After receiving the data, the edge model service converts it into an input format acceptable to the model, and then calls the locally running model for inference.
[0124] 4. Edge Model Service Returns Results After the edge device completes model inference, it returns the generated model output results to the cloud device.
[0125] The return path is implemented through a communication protocol, such as an HTTP request response.
[0126] 5. Cloud Device Collects and Processes Results After receiving the model output from the edge device, the cloud device parses, records, and evaluates the results.
[0127] These evaluation results can be further used to analyze the performance of the edge model, supporting tuning and improvement.
[0128] 6. Cloud AI Scoring System Processes Results Obtained by Cloud Evaluation Framework Deploy a cloud large model for cloud evaluation effect processing.
[0129] Based on the historical results of cloud evaluation, summarize the improvement or decline of this evaluation.
[0130] Give the existing problems of the model according to the change of evaluation effect.
[0131] For the edge model service: Start the model online API (Application Programming Interface) service through model_api; To access the edge model service using this method, the cloud device and the edge device need to be on the same local area network, or the edge device needs to have a public IP address.
[0132] Before using the service, an environment that can run Python (such as Termux (https: / / termux.dev / en / )) needs to be installed on the edge device because the httpserver (edge model service) needs to run in a Python environment.
[0133] The present invention proposes a model_api module. Running this module can convert the model service on the edge device into an API service, and the cloud evaluation framework can use the model through this API service.
[0134] Figure 3 A schematic diagram of the interaction structure between the client and the edge model service is provided, specifically: 1. Interaction between the client and the httpserver Sending a request: The client sends an HTTP request (usually a POST request) containing input data to the httpserver.
[0135] Returning an answer: The httpserver returns the processing result to the client to complete the response to the request.
[0136] 2. Functions of the httpserver The httpserver is an intermediate bridge responsible for processing client requests and communicating with the edge model.
[0137] Receiving a request: Accepting the request sent by the client and parsing the data.
[0138] Accessing the model: Passing the parsed data to the edge model for inference.
[0139] Returning an answer: Obtaining the inference result of the edge model, encapsulating it, and returning it to the client.
[0140] 3. Functions of the edge model Receiving data: The edge model receives the input data (such as the model input text) from the httpserver.
[0141] Performing inference: The model processes the input and generates an answer (inference result).
[0142] Returning the result: Passing the inference result back to the httpserver.
[0143] For the cloud-side AI scoring system: The Cloud-side AI Scoring System is an intelligent evaluation system based on cloud-based large models, aiming to comprehensively analyze and score the evaluation results of edge-side models. Through the computing and analysis capabilities of cloud-based large models, it automatically generates evaluation reports and provides targeted optimization suggestions to help developers identify problems in the models and improve the performance and reliability of edge-side models.
[0144] The Cloud-side AI Scoring System not only quantitatively scores the evaluation results, but also provides a scientific basis for model optimization through multi-dimensional analysis and comparison, making the entire edge-side model evaluation process more intelligent and closed-loop.
[0145] For the scoring system: 1. Core Logic: Multi-dimensional Dynamic Weight Fusion Scoring Inputs: Output results of edge-side models (text, classification probabilities, generated content, etc.); Task type labels (classification, generation, inference, etc.); Historical evaluation records (including scores, optimization suggestions, and effect feedback).
[0146] Dynamic Weight Allocation: Weight Generation Model: Use a two-layer LSTM network. Input the task type and real-time metrics (such as BLEU, inference time, error rate), and output the weights for each dimension: w i = Softmax(LSTM(Task Encoding ⊕ Metrics Vector)); Symbol Explanation: w i : The weight of the i-th dimension; Task Encoding: Encoding vector of the task type; Metrics Vector: Real-time metrics vector (such as BLEU, inference time, error rate); ⊕: Vector concatenation operation.
[0147] Weight Correction Mechanism: According to the historical optimization effect feedback, update the LSTM parameters through backpropagation to ensure that the weight allocation is positively correlated with the actual improvement of the model.
[0148] Comprehensive Scoring Formula: ; : Credibility correction coefficient (default 0.2); : Standardization, normalizing the original metric Mi to a unified dimension, such as Min-Max normalization; : Original metric value, that is, the original value of the i-th evaluation dimension; : Consistency between the scores of the small dataset on the edge side and the complete dataset on the cloud side; : A subset extracted from the complete dataset through intelligent sampling, specifically designed for edge devices; : The complete evaluation dataset on the cloud side, containing samples of all categories, used to comprehensively evaluate the model performance.
[0149] 2. Performance evaluation: Inference performance analysis: Analysis of the inference time of the edge-side model under different input scales.
[0150] Evaluation of CPU, GPU, and memory occupancy to judge the adaptability of the edge-side model.
[0151] Error and bias detection: Identify syntax errors, factual errors, and reasoning loopholes. Evaluate whether the model has biases.
[0152] 3. Identification of potential problems: Error classification: Analyze common error types in the model output, such as syntax errors, factual errors, and reasoning loopholes.
[0153] Bias analysis: Detect whether the model has unfair results.
[0154] Intelligent optimization suggestions: The large model on the cloud side combines the evaluation results and scores to generate specific optimization directions and improvement suggestions: 1. Model optimization suggestions: Structure optimization: Recommend methods such as pruning, quantization, and knowledge distillation to reduce the model size and computational requirements.
[0155] Training optimization: Suggest introducing new datasets or readjusting training parameters to improve the generalization ability of the model.
[0156] Parameter adjustment: Propose specific parameter adjustment schemes (such as learning rate, optimizer, etc.).
[0157] 2. Task-level optimization suggestions: For generative tasks: Improve the fluency of the output text or reduce repeated generation.
[0158] For classification tasks: Introduce hard example data augmentation or use more labeled data.
[0159] 3. Device-level optimization suggestions: For the hardware limitations on the edge side, recommend more efficient hardware acceleration libraries or frameworks (such as ONNX Runtime, TensorRT).
[0160] For device resource bottlenecks, optimize memory occupancy or inference time.
[0161] 4. Potential problem fixing: Hint the reasons for the error types in the output, such as lack of domain-specific knowledge or model input noise.
[0162] Provide targeted fixing solutions, such as improving the preprocessing steps and using data debiasing strategies.
[0163] Visual report: The system finally generates a visual report containing detailed evaluation results and suggestions: Graphically display the model performance (inference time, accuracy) and various scores.
[0164] Intuitive display of error analysis, such as heatmaps, highlighting error locations, etc.
[0165] Hierarchical display of optimization suggestions, sorted from high to low priority.
[0166] Edge-side dataset processing system: I. System design overview To narrow the gap between cloud and edge computing capabilities, especially the problem that the time required for edge devices to run the complete cloud evaluation set is too long, we designed an efficient edge-side dataset processing system. This system uses intelligent screening and dynamic sampling techniques to extract a representative small-scale subset from the cloud evaluation dataset to improve the evaluation efficiency while ensuring the statistical reliability of the results. In addition, this system integrates the evaluation results of the cloud and the edge, and uses an AI scoring system for in-depth analysis to optimize the overall evaluation process.
[0167] II. Core functional modules 2.1 Dataset partitioning module 1. Dynamic stratified sampling: Active sampling based on model uncertainty: For classification tasks, use entropy to measure the model prediction uncertainty: ; : Input sample.
[0168] C: Total number of classes.
[0169] p(c∣x): Model's predicted probability of the class.
[0170] For generation tasks, use Perplexity as the uncertainty metric.
[0171] Active sampling strategy: Select high-value samples based on model uncertainty to make the dataset more representative.
[0172] Dynamically adjust the sampling ratio in combination with historical evaluation data to reduce data redundancy.
[0173] Stratified sampling ratio optimization: The sampling quantity of each category is proportional to its uncertainty: ; : Sampling quantity of the category. N: Total sampling quantity. : Sample set of the category. : All sample sets of the i-th category. K: Total number of categories.
[0174] 2. Small dataset credibility assessment: Consistency index: Calculate the scoring difference rate between the small dataset and the complete dataset: ; : Score of the small dataset. : Score of the complete dataset.
[0175] Credibility grading: Consistency ≥ 90%: High credibility (green); 80% ≤ Consistency < 90%: Medium credibility (yellow); Consistency < 80%: Low credibility (red), resampling is required.
[0176] Statistical significance test: Use Bootstrap resampling (1000 times) to calculate the confidence interval: ; : Mean score of the small dataset. : Standard deviation of the small dataset score. n: Number of samples in the small dataset. If the score of the complete dataset falls within this interval, the small dataset is considered credible.
[0177] 2.2 Data storage and version management 1. Hierarchical storage mechanism: Adopt a hierarchical storage structure, store the complete evaluation set in the cloud and the evaluation set on the edge independently, and attach version information and sampling rules.
[0178] The data supports JSON and CSV formats, and establish indexes to optimize query efficiency.
[0179] 2. Version management and data evolution: Record the rules and versions of each data sampling, and provide a traceable version comparison function.
[0180] Support the analysis of data distribution changes between different versions to evaluate the long-term stability of the sampling strategy.
[0181] 3. Data Compression Optimization: Reduce storage and transmission costs through techniques such as feature screening, data dimensionality reduction, and redundant sample removal.
[0182] 2.3 Error Analysis and Data Integration 1. Comparative Analysis of Evaluation Results: Calculate key metrics for the edge side and the complete evaluation set, including: Classification tasks: Accuracy, Recall, F1-score.
[0183] Generation tasks: BLEU, ROUGE, text similarity.
[0184] Task-level analysis: such as response time and stability of inference tasks.
[0185] 2. Error Calculation: Absolute Error (AE): ; Relative Error (RE): ; Category Distribution Deviation: Calculate the statistical difference in category distribution using KL divergence.
[0186] 3. Evaluation Result Integration: Calculate the edge-side data first to quickly screen for abnormal model performance.
[0187] Run the complete cloud evaluation set to generate a comprehensive performance report.
[0188] Compare the edge-side and complete evaluation data, mark key performance deviations, and provide optimization suggestions based on the AI scoring system.
[0189] III. Example of Dataset Partition 3.1 Data Sampling Process 1. Input: Complete cloud evaluation set.
[0190] 2. Processing Steps: Statistically analyze the category distribution and calculate category weights. Combine historical evaluation data and category skewness to optimize the dynamic sampling strategy. Select representative samples based on the uncertainty quantification results. Generate the edge-side evaluation set and perform consistency checks to ensure evaluation stability.
[0191] 3. Output: Optimized edge-side evaluation set.
[0192] 3.2 Example Partition Suppose the class distribution of the cloud dataset is as follows: Class A: 100 items; Class B: 50 items; Class C: 150 items.
[0193] Edge-side evaluation set division (based on dynamic stratified sampling): Class A: Randomly select 15 items (adjust the weight according to the level of uncertainty); Class B: Randomly select 10 items; Class C: Randomly select 25 items; Use the consistency index to evaluate the quality of the edge-side subset. If it is lower than the threshold, resample.
[0194] IV. Data Integration and Analysis Process 4.1 Comparison of Edge-side and Complete Evaluation Results 1. Edge-side evaluation set evaluation: Calculate the edge-side evaluation data first to quickly detect potential problems.
[0195] 2. Complete evaluation set evaluation: Run the complete dataset to obtain the final model evaluation results.
[0196] 4.2 Analysis of Result Differences 1. Calculate the evaluation deviation: Calculate the deviation of key indicators between the edge-side and the complete evaluation set.
[0197] Mark the classes or tasks that may affect the evaluation consistency.
[0198] 2. Comprehensive evaluation report: Combine the AI scoring system to generate a complete evaluation comparison report. Provide optimization suggestions based on data distribution and error analysis.
[0199] 1. Edge-side model evaluation framework: A lightweight and easy-to-deploy edge-side model evaluation framework is proposed, which solves the problem that existing evaluation tools cannot run on edge-side devices.
[0200] 2. Collaborative evaluation mechanism between the cloud and the edge-side: Design a collaborative evaluation process between the cloud and the edge-side to achieve data transmission and result feedback through communication protocols such as HTTP or ADB.
[0201] 3. Intelligent scoring system based on large models: A multi-dimensional dynamic weight fusion scoring method is proposed, which combines the computing power of cloud large models to generate comprehensive evaluation reports and optimization suggestions.
[0202] 4. Edge-side dataset processing system: Design a dynamic stratified sampling technique to extract a representative small-scale subset from the cloud evaluation dataset for edge-side evaluation.
[0203] Through intelligent sampling and statistical optimization, this system enables edge devices to efficiently complete model evaluation under limited computing power, while ensuring the representativeness and statistical stability of the dataset. The in-depth analysis based on AI further enhances the scientific nature and optimization ability of the evaluation process.
[0204] Figure 4 A structural schematic diagram of an edge large model evaluation device based on artificial intelligence provided for one or more embodiments of this specification. The device is applied to an evaluation platform, which includes a cloud module and an edge module. The cloud module includes a cloud device, a cloud evaluation framework, and a cloud AI scoring system. The edge module includes an edge device and an edge model service, including: an evaluation data generation unit 401, an output result generation unit 402, an output result analysis unit 403, and an evaluation result analysis unit 404.
[0205] The evaluation data generation unit 401. The cloud device runs the cloud evaluation framework, generates evaluation data through the cloud evaluation framework, and transmits the evaluation data to the edge device through the cloud evaluation framework. The evaluation data includes evaluation input data and evaluation tasks. The output result generation unit 402. The edge device receives the evaluation data through the edge model service, invokes the edge large model running locally for inference to obtain a model output result, and transmits the model output result to the cloud device through the edge model service. The output result analysis unit 403. The cloud device analyzes the model output result to obtain the evaluation result of the edge large model. The evaluation result analysis unit 404. The cloud AI scoring system analyzes the evaluation result in combination with historical evaluation results, generates a scoring analysis result, and generates an improvement strategy adapted to the edge large model based on the scoring analysis result.
[0206] Figure 5A structural schematic diagram of an end-side large model evaluation device based on artificial intelligence provided for one or more embodiments of this specification. The device is applied to an evaluation platform, which includes a cloud module and an end-side module. The cloud module includes a cloud device, a cloud evaluation framework, and a cloud AI scoring system. The end-side module includes an end-side device and an end-side model service, including: at least one processor; and a memory communicatively connected to the at least one processor. Wherein, the memory stores instructions executable by the at least one processor. When the instructions are executed by the at least one processor, the at least one processor is enabled to: The cloud device runs the cloud evaluation framework, generates evaluation data through the cloud evaluation framework, and transmits the evaluation data to the end-side device through the cloud evaluation framework. The evaluation data includes evaluation input data and evaluation tasks. The end-side device receives the evaluation data through the end-side model service, invokes the locally running end-side large model for inference, obtains a model output result, and transmits the model output result to the cloud device through the end-side model service. The cloud device analyzes the model output result to obtain an evaluation result of the end-side large model. The cloud AI scoring system analyzes the evaluation result in combination with historical evaluation results, generates a scoring analysis result, and generates an improvement strategy adapted to the end-side large model based on the scoring analysis result.
[0207] A non-volatile computer storage medium provided for one or more embodiments of this specification, applied to an evaluation platform, which includes a cloud module and an end-side module. The cloud module includes a cloud device, a cloud evaluation framework, and a cloud AI scoring system. The end-side module includes an end-side device and an end-side model service, and stores computer-executable instructions that can be realized when executed by a computer: The cloud device runs the cloud evaluation framework, generates evaluation data through the cloud evaluation framework, and transmits the evaluation data to the end-side device through the cloud evaluation framework. The evaluation data includes evaluation input data and evaluation tasks. The end-side device receives the evaluation data through the end-side model service, invokes the locally running end-side large model for inference, obtains a model output result, and transmits the model output result to the cloud device through the end-side model service. The cloud device analyzes the model output result to obtain an evaluation result of the end-side large model. The cloud AI scoring system analyzes the evaluation result in combination with historical evaluation results, generates a scoring analysis result, and generates an improvement strategy adapted to the end-side large model based on the scoring analysis result.
[0208] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the embodiments of the apparatus, device, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.
[0209] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the apparatus embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.
[0210] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0211] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network device and method can be implemented in other ways. For example, the apparatus / network device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0212] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0213] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above units can be implemented in the form of hardware or software.
[0214] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0215] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. An artificial intelligence-based large-scale model evaluation method on the end, characterized in that: The method is applied to an evaluation platform, the evaluation platform includes a cloud module and a terminal module, the cloud module includes a cloud device, a cloud evaluation framework and a cloud AI scoring system, the terminal module includes a terminal device and a terminal model service, and the method includes: The cloud device runs the cloud evaluation framework, generates evaluation data through the cloud evaluation framework, and transmits the evaluation data to the terminal device through the cloud evaluation framework, wherein the evaluation data includes evaluation input data and evaluation tasks; The end-side device receives the evaluation data through the end-side model service, calls the locally running end-side large model to perform inference, obtains a model output result, and transmits the model output result to the cloud device through the end-side model service; The cloud device analyzes the model output result to obtain the evaluation result of the terminal-side large model; The cloud-based AI scoring system analyzes the evaluation results in combination with historical evaluation results to generate a scoring analysis result, and generates an improvement strategy adapted to the end-side large model based on the scoring analysis result.
2. The method according to claim 1, characterized in that The method further comprises: The evaluation platform converts the end-side model service into an API service through a preset API module, so that the cloud-based evaluation framework calls the end-side large model for reasoning through the end-side model service.
3. The method according to claim 1, characterized in that The evaluation results include real-time indicators and task types of the large model on the client side, and the scoring analysis results are weights of each dimension; the cloud AI scoring system analyzes the evaluation results in combination with historical evaluation results to generate scoring analysis results, including: The cloud-based AI scoring system uses a two-layer LSTM network, inputs the real-time indicators and the task types, and outputs the weights of each dimension.
4. The method according to claim 1, characterized in that The cloud-based evaluation framework includes an end-side data set processing system, which uses intelligent screening and dynamic sampling technology to extract representative subsets from the data set of the cloud-based evaluation framework and send the subsets to the AI scoring system for analysis.
5. The method according to claim 4, characterized in that The end-side data set processing system includes a data set division module, an error analysis module and a data integration module. The data set division module uses dynamic stratified sampling. The error analysis module analyzes the evaluation results to obtain error results. The data integration module integrates the error results to obtain integrated evaluation results, so that the cloud-based AI scoring system can analyze the integrated evaluation results in combination with historical evaluation results to generate scoring analysis results.
6. The method according to claim 1, characterized in that The method further comprises: A cross-modal alignment module is integrated into the cloud-based evaluation framework so that the cross-modal alignment module constructs a multimodal semantic consistency evaluation index through feature space mapping technology, wherein the multimodal semantic consistency evaluation index is used to quantify the semantic deviation value of the cross-modal output result, and the multimodality includes at least image, voice and text.
7. The method according to claim 1, characterized in that The method further comprises: The cloud AI scoring system collects resource parameters of the terminal device and constructs a dynamic quantitative sensitivity matrix; Identify the sensitive layers that are susceptible to quantization error propagation through a pre-established quantization error propagation model; Based on the dependency relationship between the sensitive layers, an error accumulation prediction model is constructed, and an error result is predicted by the error accumulation prediction model; The output characteristics of the sensitive layer are monitored based on the error results: When the error result is greater than a preset error value, a corresponding hierarchical calibration is triggered, wherein the hierarchical calibration includes local inverse quantization on the end side to repair the parameters of the sensitive layer and federal calibration on the cloud side to update the quantization strategy of the sensitive layer.
8. An artificial intelligence-based large-scale model evaluation device on the end, characterized in that: The device is applied to an evaluation platform, the evaluation platform includes a cloud module and a terminal module, the cloud module includes a cloud device, a cloud evaluation framework and a cloud AI scoring system, the terminal module includes a terminal device and a terminal model service, including: An evaluation data generating unit, wherein the cloud device runs the cloud evaluation framework, generates evaluation data through the cloud evaluation framework, and transmits the evaluation data to the terminal device through the cloud evaluation framework, wherein the evaluation data includes evaluation input data and evaluation tasks; An output result generating unit, wherein the end-side device receives the evaluation data through the end-side model service, calls the locally running end-side large model to perform reasoning, obtains a model output result, and transmits the model output result to the cloud device through the end-side model service; An output result analysis unit, wherein the cloud device analyzes the model output result to obtain an evaluation result of the large model on the terminal side; An evaluation result analysis unit, wherein the cloud-based AI scoring system analyzes the evaluation results in combination with historical evaluation results to generate a scoring analysis result, and generates an improvement strategy adapted to the end-side large model based on the scoring analysis result.
9. An artificial intelligence-based large-scale model evaluation device on the end, characterized in that: The device is applied to an evaluation platform, the evaluation platform includes a cloud module and a terminal module, the cloud module includes a cloud device, a cloud evaluation framework and a cloud AI scoring system, the terminal module includes a terminal device and a terminal model service, including: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: The cloud device runs the cloud evaluation framework, generates evaluation data through the cloud evaluation framework, and transmits the evaluation data to the terminal device through the cloud evaluation framework, wherein the evaluation data includes evaluation input data and evaluation tasks; The end-side device receives the evaluation data through the end-side model service, calls the locally running end-side large model to perform inference, obtains a model output result, and transmits the model output result to the cloud device through the end-side model service; The cloud device analyzes the model output result to obtain the evaluation result of the terminal-side large model; The cloud-based AI scoring system analyzes the evaluation results in combination with historical evaluation results to generate a scoring analysis result, and generates an improvement strategy adapted to the end-side large model based on the scoring analysis result.
10. A non-volatile computer storage medium, characterized in that: Applied to an evaluation platform, the evaluation platform includes a cloud module and a terminal module, the cloud module includes a cloud device, a cloud evaluation framework and a cloud AI scoring system, the terminal module includes a terminal device and a terminal model service, and stores computer executable instructions, which can achieve the following when executed by a computer: The cloud device runs the cloud evaluation framework, generates evaluation data through the cloud evaluation framework, and transmits the evaluation data to the terminal device through the cloud evaluation framework, wherein the evaluation data includes evaluation input data and evaluation tasks; The end-side device receives the evaluation data through the end-side model service, calls the locally running end-side large model to perform inference, obtains a model output result, and transmits the model output result to the cloud device through the end-side model service; The cloud device analyzes the model output result to obtain the evaluation result of the terminal-side large model; The cloud-based AI scoring system analyzes the evaluation results in combination with historical evaluation results to generate a scoring analysis result, and generates an improvement strategy adapted to the end-side large model based on the scoring analysis result.
Citation Information
Patent Citations
Large model-based data processing method and server
CN116383026A
Overall step, model calling and data set loading method for large model evaluation
CN118551191A
End-cloud collaboration-based large model ecological online evolution learning method and application
CN119721173A
Data processing method based on large model, and server
WO2024251168A1
Terminal-cloud computing power collaborative scheduling method, related system, and device
WO2025082323A1
Cited By
Mechanical arm motion control method based on multi-agent cooperation
CN120620234A
A robotic arm motion control method based on multi-agent cooperation
CN120620234B
Large language model security assessment method and device and electronic equipment
CN120805149A
Adaptive quantization language interaction method and related equipment
CN121214929A