Model lightweight processing method capable of being used for high-real-time scene and related system
By dynamically determining the model framework and tools, and based on the resources and requirements of edge devices, lightweight modeling in high real-time scenarios is achieved, solving the problem of poor model lightweighting effect in existing technologies and meeting the needs of real-time data processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-09
- Publication Date
- 2026-03-10
AI Technical Summary
Existing lightweight modeling methods generally perform poorly in high real-time scenarios and cannot meet the needs of real-time data processing.
By receiving model configuration requests from edge devices, and based on their hardware and software resource parameters and data processing requirements, the model framework and tools are dynamically determined, the model is lightweighted, and a lightweight model framework that meets the requirements is generated.
It achieves lightweight deep modeling in high real-time scenarios, meeting the real-time data processing needs of edge devices, while optimizing computing resources and storage space.
Smart Images

Figure CN121638362A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a model lightweight processing method and related system applicable to high real-time scenarios. BACKGROUND
[0002] With the development of artificial intelligence and the wide application of deep learning technology in various fields, people have increasingly realized that the performance and effect of a model are important, but factors such as the size and computational complexity of the model are also increasingly concerned, and model lightweight is increasingly favored. Model lightweight aims to reduce the demand for computing resources, reduce the storage space occupation of the model, and speed up the inference rate on the basis of maintaining the original performance of the model as much as possible, so that the lightweight model can be more suitable for resource-constrained devices (such as mobile terminals, embedded devices, etc.) and application scenarios with high real-time requirements (such as real-time image classification, real-time target detection, etc.).
[0003] At present, model lightweight is often mechanical, that is, fixed model lightweight parameters are generally set, which leads to general model lightweight effect and cannot meet the real-time data processing requirements. Therefore, how to deeply realize model lightweight to meet the real-time data processing requirements. SUMMARY
[0004] The embodiments of the present application provide a model lightweight processing method and related system applicable to high real-time scenarios, which can dynamically determine corresponding model lightweight processing parameters based on the actual data processing requirements of real-time high real-time scenarios and the characteristics of the local device, so that model lightweight can be deeply realized to meet the real-time data processing requirements.
[0005] In a first aspect, the embodiments of the present application provide a model lightweight processing method applicable to high real-time scenarios. The method is applied to a server device in a model lightweight processing system applicable to high real-time scenarios. The system also includes n edge devices, where n is a positive integer. The method includes: receiving a model configuration request of a first edge device, the model configuration request including first hardware resource parameters, first software resource parameters of the first edge device, and first data processing requirement parameters for a target high real-time scenario; the first data processing requirement parameters including a first data processing rate, a first function type, and a first model precision; the first edge device being one of the n edge devices; determining a first model framework and a first model tool according to the first data processing requirement parameters; determining first model lightweight processing parameters according to the first hardware resource parameters, the first software resource parameters, and the first model framework; According to the first model lightweight processing parameter, the first model framework is processed to obtain a second model framework; the second model framework and the first model tool are used to configure a corresponding model service for the first edge device.
[0006] In a second aspect, the embodiments of the present application provide a model lightweight processing system applicable to a high real-time scenario, the model lightweight processing system applicable to the high real-time scenario comprising a server device and n edge devices, n being a positive integer; the model lightweight processing system applicable to the high real-time scenario further comprising a receiving unit, a determining unit and a processing unit, wherein, The receiving unit is configured to receive a model configuration request of a first edge device, the model configuration request comprising first hardware resource parameters, first software resource parameters and first data processing requirement parameters of the first edge device for a target high real-time scenario; the first data processing requirement parameters comprising a first data processing rate, a first function type and a first model precision; the first edge device being one of the n edge devices; The determining unit is configured to determine a first model framework and a first model tool according to the first data processing requirement parameters; The processing unit is configured to determine first model lightweight processing parameters according to the first hardware resource parameters, the first software resource parameters and the first model framework; and process the first model framework according to the first model lightweight processing parameters to obtain a second model framework; the second model framework and the first model tool being used to configure a corresponding model service for the first edge device.
[0007] The embodiments of the present application have the following beneficial effects: It can be seen that the model lightweight processing method and related system for high real-time scenarios described in the embodiments of the application, wherein the method is applied to a server device in a model lightweight processing system for high real-time scenarios, and the system further includes: n edge devices, n is a positive integer, first, receiving a model configuration request of a first edge device, the model configuration request including first hardware resource parameters, first software resource parameters and first data processing requirement parameters for a target high real-time scenario; the first data processing requirement parameters include a first data processing rate, a first function type and a first model accuracy; the first edge device is one of the n edge devices, determining a first model framework and a first model tool according to the first data processing requirement parameters, the first data processing requirement parameters representing the model configuration requirements of the first edge device, and the corresponding model framework and model tool corresponding to the model framework can be configured based on the model configuration requirements of the first edge device to ensure that the first model framework meets the model configuration requirements of the first edge device, then determining first model lightweight processing parameters according to the first hardware resource parameters, the first software resource parameters and the first model framework, processing the first model framework according to the first model lightweight processing parameters to obtain a second model framework; the second model framework and the first model tool are used for the first edge device to configure a corresponding model service, the first hardware resource parameters can be used to represent the hardware performance of the first edge device and the first software resource parameters can be used to represent the software performance of the first edge device, that is, the model characteristics of the first model framework and the hardware performance and software performance of the first edge device can be used to dynamically determine the corresponding model lightweight processing parameters, so that the model lightweight can be deeply realized to meet the real-time data processing requirements. BRIEF DESCRIPTION OF DRAWINGS
[0008] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0009] Figure 1 is an architecture schematic diagram of a model lightweight processing system for high real-time scenarios provided by the embodiments of the present application; Figure 2 is a flow schematic diagram of a model lightweight processing method for high real-time scenarios provided by the embodiments of the present application; Figure 3 is a functional unit composition block diagram of a model lightweight processing system for high real-time scenarios provided by the embodiments of the present application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of this application. Detailed Implementation
[0010] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0011] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0012] It should be understood that the term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this document indicates that the preceding and following related objects are in an "or" relationship. In the embodiments of this application, "multiple" refers to two or more.
[0013] In the embodiments of this application, "at least one item" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. "One or more" refers to one or more items, while "multiple" refers to two or more items. For example, "at least one item" of a, b, or c can represent the following seven cases: a, b, c; a and b; a and c; b and c; a, b, and c. Each of a, b, and c can be an element or a set containing one or more elements.
[0014] In this application, the term "connection" refers to various connection methods, such as direct connection or indirect connection, to achieve communication between devices. This application does not impose any limitations on this.
[0015] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0016] The edge devices described in the embodiments of this application may include at least one of the following: edge server, edge device, mobile terminal, smartphone, embedded device, smart switch cabinet, tablet computer, handheld computer, laptop computer, video matrix, monitoring platform, mobile internet device (MID) or wearable device, etc. The above are merely examples and not exhaustive, and include but are not limited to the above devices.
[0017] The server-side device described in the embodiments of this application may include a server, for example, a cloud server.
[0018] The model deployment tool described in this application can be understood as a software tool used to transfer a trained machine learning or deep learning model, for example, from a development environment to a real-world operating environment (such as a server, edge device, or mobile device). Its main function is to enable the model to provide inference services efficiently and stably.
[0019] The model conversion tool described in this application can be understood as software used to convert models trained on different deep learning frameworks (e.g., PyTorch, TensorFlow, etc.) into a format that can run efficiently on specific hardware or platforms. It primarily addresses the compatibility issues of model deployment across frameworks and devices.
[0020] The model training tools described in the embodiments of this application can be used to implement model training, such as TensorFlow and PyTorch.
[0021] The structural visualization tools described in this application can perform structuring of models to enable human-computer interaction. Examples of such tools include Netron and TensorBoard.
[0022] The experimental tracking tools described in this application can be used to track and monitor edge models. For example, experimental tracking tools such as Wandb.
[0023] The model lightweighting tools described in the embodiments of this application can be used to implement model lightweighting processes. For example, a model lightweighting tool such as the quantization tool PPQ.
[0024] Please see Figure 1 , Figure 1 This is a schematic diagram of the architecture of a lightweight model processing system that can be used in high real-time scenarios, provided by an embodiment of this application. The lightweight model processing system that can be used in high real-time scenarios includes: an edge device and n edge devices, where n is a positive integer, and communication connections between the edge device and the n edge devices.
[0025] Among them, the server-side equipment in the lightweight model processing system, which can be used in high real-time scenarios, can be used to implement the following functions: The system receives a model configuration request from a first edge device. The model configuration request includes first hardware resource parameters, first software resource parameters, and first data processing requirement parameters for the target high real-time scenario of the first edge device. The first data processing requirement parameters include a first data processing rate, a first function type, and a first model accuracy. The first edge device is one of the n edge devices. The first model framework and the first model tool are determined based on the first data processing requirement parameters; The first model lightweighting processing parameters are determined based on the first hardware resource parameters and the first software resource parameters. The first model framework is processed according to the first model lightweighting processing parameters to obtain the second model framework; the second model framework and the first model tool are used to configure the corresponding model services on the first edge device.
[0026] As can be seen, the lightweight model processing system for high real-time scenarios described in this application includes a server device and n edge devices, where n is a positive integer. First, the server device receives a model configuration request from a first edge device. The model configuration request includes first hardware resource parameters, first software resource parameters, and first data processing requirement parameters for the target high real-time scenario. The first data processing requirement parameters include a first data processing rate, a first function type, and a first model accuracy. The first edge device is one of the n edge devices. Based on the first data processing requirement parameters, a first model framework and a first model tool are determined. The first data processing requirement parameters characterize the model configuration requirements of the first edge device. Based on the model configuration requirements of the first edge device, the corresponding model framework and tool can be configured. A corresponding model tool is provided to ensure that the first model framework meets the model configuration requirements of the first edge device. Then, the first model lightweighting processing parameters are determined based on the first hardware resource parameters, the first software resource parameters, and the first model framework. The first model framework is then processed according to the first model lightweighting processing parameters to obtain the second model framework. The second model framework and the first model tool are used to configure the corresponding model services on the first edge device. The first hardware resource parameters can be used to characterize the hardware performance of the first edge device, and the first software resource parameters can be used to characterize the software performance of the first edge device. That is, the corresponding model lightweighting processing parameters can be dynamically determined based on the model characteristics of the first model framework and the hardware and software performance of the first edge device. In this way, deep model lightweighting can be achieved to meet the real-time data processing requirements.
[0027] Furthermore, based on this lightweight model processing system applicable to high real-time scenarios, the server-side equipment can perform batch and targeted lightweight model configuration for edge devices to meet the real-time data processing needs of each edge device.
[0028] Please see Figure 2 , Figure 2 This is a flowchart illustrating a lightweight model processing method for high real-time scenarios provided in this application embodiment, applicable to, for example... Figure 1 The model lightweight processing system described above, applicable to high real-time scenarios, is a server-side device within the system. The system further includes n edge devices, where n is a positive integer. The model lightweight processing method for high real-time scenarios may include the following steps: S201: Receive a model configuration request from a first edge device. The model configuration request includes first hardware resource parameters, first software resource parameters, and first data processing requirement parameters for the target high real-time scenario of the first edge device. The first data processing requirement parameters include a first data processing rate, a first function type, and a first model accuracy. The first edge device is one of the n edge devices.
[0029] Among them, the target high real-time scenarios can be preset or defaulted to by the system. For example, the target high real-time scenarios can be image classification, fault diagnosis, anomaly detection, target detection, traffic violation recognition, etc., without limitation.
[0030] The server-side device can communicate with n edge devices. In the specific implementation, taking the first edge device as an example, the first edge device is any one of the n edge devices. The first edge device can actively send a model configuration request to the server-side device. The model configuration request includes the first hardware resource parameters, the first software resource parameters, and the first data processing requirement parameters for the target high real-time scenario of the first edge device.
[0031] The first hardware resource parameter can be used to characterize the hardware performance of the first edge device. The first hardware resource parameter may include at least one of the following: processor type, memory size, storage capacity, computing power, etc., which are not limited here.
[0032] The first software resource parameter can be used to characterize the software performance of the first edge device. The first software resource parameter may include at least one of the following: operating system, runtime library, CPU utilization, response time, number of concurrent users, resource utilization, etc., without limitation.
[0033] The first data processing requirement parameter can be used to characterize the model configuration requirements of the first edge device. This parameter includes a first data processing rate, a first function type, and a first model accuracy. The first data processing rate reflects the model's real-time requirements; the first function type reflects the model's functionalities, such as target detection, image classification, fault diagnosis, partial discharge type recognition, etc.; and the first model accuracy reflects the model's precision requirements. The first data processing rate can be understood as the input data rate of the model, or it can also be understood as the data transmission rate.
[0034] S202: Determine the first model framework and the first model tool based on the first data processing requirement parameters.
[0035] Among them, since the first data processing requirement parameter characterizes the model configuration requirements of the first edge device, such as what functions the model should implement, what the model recognition accuracy is, and the model's real-time capability, the corresponding model framework and the model tool corresponding to the model framework can be configured based on the model configuration requirements of the first edge device to ensure that the first model framework meets the model configuration requirements of the first edge device, and the model deployment and model conversion can be completed based on the first model tool corresponding to the first model framework, so that the model of the first edge device can be successfully implemented.
[0036] The model framework can be preset or set by the system default. For example, the model framework can be a deep learning model, a large model, a neural network model, etc., without any restrictions.
[0037] S203: Determine the first model lightweighting processing parameters based on the first hardware resource parameters, the first software resource parameters, and the first model framework.
[0038] The first model lightweighting processing parameter may include a model lightweighting strategy, which may include a model lightweighting algorithm and algorithm control parameters for the model lightweighting algorithm. The algorithm control parameters for the model lightweighting algorithm are used to control the model lightweighting effect, such as the memory ratio, the computational resource ratio, and the speed ratio of model lightweighting.
[0039] The model lightweighting algorithm may include at least one of the following: model quantization algorithm, model pruning algorithm, model distillation algorithm, etc., without limitation.
[0040] The model quantization algorithm may include at least one of the following: static quantization algorithm, dynamic quantization algorithm, quantization-aware training algorithm, etc., without limitation.
[0041] Among them, the model quantization algorithm mainly utilizes the distribution pattern of data within a certain range and uses fewer bits to represent parameters through a reasonable mapping relationship. While reducing storage requirements, the hardware platform can calculate low-precision data faster and reduce computing resource consumption.
[0042] In deep learning, neural network models often have a huge number of parameters and complex structures. For example, some advanced image recognition models and natural language processing models may contain tens of millions or even hundreds of millions of parameters. The core idea of model pruning is to identify and remove the relatively unimportant parts of the model that contribute little to the final output (such as neurons and connections). This reduces the model size, lowers computational resource consumption, and speeds up inference, making the model easier to deploy on resource-constrained devices (such as mobile devices and embedded devices).
[0043] Among them, the knowledge distillation algorithm aims to transfer the knowledge learned by a complex, high-performance large "teacher model" to a relatively simple, smaller "student model", so that the student model can imitate the behavior of the teacher model as much as possible with fewer parameters and a simpler structure, thereby achieving similar performance.
[0044] In specific implementation, since the first hardware resource parameter can be used to characterize the hardware performance of the first edge device and the first software resource parameter can be used to characterize the software performance of the first edge device, the corresponding model lightweighting processing parameters can be dynamically determined based on the model characteristics of the first model framework and the hardware and software performance of the first edge device. In this way, model lightweighting can be deeply realized to meet the real-time data processing requirements.
[0045] S204: The first model framework is processed according to the first model lightweighting processing parameters to obtain a second model framework; the second model framework and the first model tool are used to configure the corresponding model services on the first edge device.
[0046] In specific implementation, the first model framework can be lightweighted according to the first model lightweighting processing parameters to obtain the second model framework. The second model framework and the first model tool are used to configure the corresponding model service on the first edge device. Then, the second model framework and the first model tool are distributed to the first edge device. The first edge device can then use the first model tool to configure the second model framework. That is, a lightweight model that meets the actual data processing requirements of real-time and high real-time scenarios and the characteristics of the device itself can be configured on the first edge device. In this way, the corresponding model lightweighting processing parameters can be dynamically determined based on the actual data processing requirements of real-time and high real-time scenarios and the characteristics of the device itself. Thus, model lightweighting can be deeply implemented to meet the needs of real-time data processing.
[0047] For example, in this embodiment of the application, the specific process of model adaptation is as follows: The server-side equipment can clearly define the hardware resources of the edge devices, such as processor type, memory size, storage capacity, and computing power, as well as the software environment, such as operating system and runtime libraries. Addressing the computing power and storage limitations of the edge devices, model quantization and / or model pruning and / or knowledge distillation are used to optimize the model structure. Finally, based on the hardware and software environment of the edge devices, a suitable model framework and tools are selected.
[0048] Of course, further, the adapted model (second model framework) can be evaluated on edge devices using test datasets or data from real-world scenarios to measure metrics such as accuracy, recall, and F1 score, as well as performance metrics such as inference time, memory usage, and power consumption. Based on the evaluation results, the model can be further optimized by adjusting its parameters and optimizing the algorithm to improve its performance and efficiency.
[0049] For example, in this embodiment of the application, the specific process of model deployment is as follows: The server-side device can build the corresponding development and runtime environment according to the operating system and hardware platform of the edge device, use the corresponding model conversion tool to convert the original model file into the target format, and perform necessary optimizations and adjustments. It can also select a suitable deployment tool for the edge device to ensure that the model can run stably on the edge device. The model is then encapsulated into a callable service, and through a clearly defined interface, other applications can easily interact with the model. Finally, the converted model file and related dependency libraries are deployed to the edge device, and the model service is started.
[0050] Furthermore, the deployed model can be tested in real-world scenarios to verify whether its functionality and performance meet requirements. Additionally, a monitoring system can be established to monitor performance metrics, model running status, and resource usage on edge devices in real time. This allows for timely identification and resolution of problems based on real-time monitoring data, ensuring model capabilities. Simultaneously, the model and deployment can be further optimized based on actual operational conditions to improve system stability and performance.
[0051] As can be seen, the model lightweighting processing method described in this application embodiment, applicable to high real-time scenarios, is applied to a server-side device in a model lightweighting processing system applicable to high real-time scenarios. This system further includes n edge devices, where n is a positive integer. First, upon receiving a model configuration request from a first edge device, the model configuration request includes first hardware resource parameters, first software resource parameters, and first data processing requirement parameters for the target high real-time scenario. The first data processing requirement parameters include a first data processing rate, a first function type, and a first model accuracy. The first edge device is one of the n edge devices. Based on the first data processing requirement parameters, a first model framework and a first model tool are determined. The first data processing requirement parameters characterize the model configuration requirements of the first edge device. The model configuration requirements of the first edge device can be determined based on the model configuration requirements of the first edge device. The process involves configuring a corresponding model framework and its corresponding model tools to ensure that the first model framework meets the model configuration requirements of the first edge device. Then, based on the first hardware resource parameters, the first software resource parameters, and the first model framework, first model lightweighting parameters are determined. The first model framework is then processed according to these parameters to obtain a second model framework. The second model framework and the first model tools are used to configure corresponding model services on the first edge device. The first hardware resource parameters can be used to characterize the hardware performance of the first edge device, and the first software resource parameters can be used to characterize the software performance of the first edge device. In other words, the corresponding model lightweighting parameters can be dynamically determined based on the model characteristics of the first model framework and the hardware and software performance of the first edge device. This allows for deep model lightweighting to meet real-time data processing requirements.
[0052] Optionally, the step of determining the first model lightweighting parameters based on the first hardware resource parameters and the first software resource parameters can be implemented in the following manner: The first performance evaluation parameter is determined based on the first hardware resource parameter; Determine the second performance evaluation parameters based on the first software resource parameters; Determine the target performance evaluation parameters based on the first performance evaluation parameters and the second performance evaluation parameters; Determine the reference performance evaluation parameters corresponding to the first data processing rate and the first model framework; Determine a first deviation value between the target performance evaluation parameter and the reference performance evaluation parameter; The first model lightweighting parameters are determined based on the first deviation value, the first performance evaluation parameter, and the second performance evaluation parameter.
[0053] In a specific implementation, a pre-stored mapping relationship between preset hardware resource parameters and performance evaluation parameters can be used to determine the first performance evaluation parameter corresponding to the first hardware resource parameter. Alternatively, a pre-stored mapping relationship between preset software resource parameters and performance evaluation parameters can be used to determine the second performance evaluation parameter corresponding to the first software resource parameter.
[0054] Next, we can obtain the first weight for hardware and the second weight for software. The sum of the first weight and the second weight is 1. We can then perform a weighted operation based on the first performance evaluation parameter and the second performance evaluation parameter to obtain the target performance evaluation parameter. The target performance evaluation parameter = first weight × first performance evaluation parameter + second weight × second performance evaluation parameter.
[0055] In specific implementation, since different model frameworks have different requirements for data processing rate and performance, a pre-stored correspondence table between preset model frameworks and mapping relationships can be used to represent the mapping relationship between data processing rate and performance evaluation parameters. That is, the first mapping relationship corresponding to the first model framework can be determined based on the correspondence table, and the reference performance evaluation parameters corresponding to the first data processing rate can be determined based on the first mapping relationship. In other words, the corresponding performance evaluation parameters can be evaluated based on different model frameworks and data processing rates.
[0056] Furthermore, the first deviation value is used to characterize the degree of difference between the actual performance and the required performance. That is, to determine the first deviation value between the target performance evaluation parameter and the reference performance evaluation parameter, the first deviation value = (target performance evaluation parameter - reference performance evaluation parameter) / reference performance evaluation parameter. Finally, the first model lightweighting processing parameters can be determined based on the first deviation value, the first performance evaluation parameter, and the second performance evaluation parameter. The first deviation value not only corresponds to the degree of lightweighting processing but also to the direction of lightweighting processing. The first performance evaluation parameter and the second performance evaluation parameter reflect the actual hardware and software performance of the first edge device. Thus, the first model framework can be dynamically quantified in a purposeful and directional manner by combining the actual hardware and software performance of the first edge device. In this way, the corresponding model lightweighting processing parameters can be dynamically determined based on the actual data processing requirements of high real-time scenarios and the characteristics of the device itself. In this way, model lightweighting can be deeply realized to meet the real-time data processing requirements.
[0057] In practice, model lightweighting can be implemented when the target performance evaluation parameter is less than the reference performance evaluation parameter. In this case, it can compensate for the insufficient performance of edge devices by deeply lightweighting the model to achieve deep model lightweighting, so as to meet the real-time data processing requirements and not reduce the model performance or make the model performance controllable. Of course, model lightweighting can also be implemented when the target performance evaluation parameter is greater than or equal to the reference performance evaluation parameter, which can reduce the computing resource requirements, reduce the model's storage space occupation, and accelerate the inference speed.
[0058] Optionally, the above steps, in which the first model lightweighting parameters are determined based on the first deviation value, the first performance evaluation parameter, and the second performance evaluation parameter, can be implemented in the following manner: A first model lightweighting strategy set is determined based on the first performance evaluation parameter; the first model lightweighting strategy set includes at least one model lightweighting strategy; each model lightweighting strategy corresponds to a target second performance evaluation parameter; The target model lightweighting strategy is determined from the first model lightweighting strategy set based on the second performance evaluation parameter and a target second performance evaluation parameter corresponding to each model lightweighting strategy. The first model lightweighting processing parameters corresponding to the target model lightweighting strategy are determined based on the first deviation value.
[0059] In specific implementations, due to different hardware configurations, corresponding model lightweighting strategy sets can be configured for different hardware to ensure that the model lightweighting effect is deeply adapted to the actual hardware configuration. For example, a pre-stored mapping relationship between preset hardware performance evaluation parameters and model lightweighting strategy sets can be stored. The model lightweighting strategy set can include at least one model lightweighting strategy, and each model lightweighting strategy corresponds to a software performance evaluation parameter. That is, the first model lightweighting strategy set corresponding to the first performance evaluation parameter can be determined based on the mapping relationship. Correspondingly, the first model lightweighting strategy set can also include at least one model lightweighting strategy, and each model lightweighting strategy can correspond to a target second performance evaluation parameter, which is the software performance evaluation parameter.
[0060] Next, the second performance evaluation parameter can be matched with a target second performance evaluation parameter corresponding to each model lightweighting strategy. Specifically, for example, the absolute value of the difference between the second performance evaluation parameter and the target second performance evaluation parameter corresponding to each model lightweighting strategy can be determined to obtain at least one absolute value. The minimum value among these at least one absolute value can be selected, and the model lightweighting strategy corresponding to the minimum value can be used as the target model lightweighting strategy. Finally, the first model lightweighting processing parameter corresponding to the target model lightweighting strategy can be determined based on the first deviation value. For example, the algorithm control of the model lightweighting algorithm of the preset target model lightweighting strategy can be stored in advance. The algorithm controls the model by pre-storing a mapping relationship between the deviation value and the optimization parameter. Based on this mapping relationship, the first optimization parameter corresponding to the first deviation value can be determined. This first optimization parameter is then used to optimize the algorithm control parameter, resulting in the optimized algorithm control parameter. For example, the optimized algorithm control parameter = algorithm control parameter + first optimization parameter. The first optimization parameter is used to optimize the model lightweighting effect. It can be understood as the quantization parameter that needs to be optimized along with the algorithm control parameter. This can compensate for the insufficient performance of edge devices by performing deep lightweighting of the model to achieve deep model lightweighting and meet the needs of real-time data processing. For example, the target model lightweighting strategy may include at least one model lightweighting stage, and each model lightweighting stage may correspond to a model lightweighting algorithm, such as quantization + pruning + distillation. The first optimization parameter can optimize the algorithm control parameter of one model lightweighting stage, or it can optimize the algorithm control parameter of multiple model lightweighting stages. The first optimization parameter is used to improve the model lightweighting effect by changing the algorithm control parameter. This means that by combining the actual hardware and software performance of the first edge device, the first model framework can be dynamically quantified in a purposeful and directional manner. Then, based on the actual data processing requirements of high real-time scenarios and the characteristics of the device itself, the corresponding model lightweighting processing parameters can be dynamically determined. In this way, model lightweighting can be deeply realized to meet the needs of real-time data processing.
[0061] Optionally, the above steps, determining the first model framework and the first model tool based on the first data processing requirement parameters, can be implemented in the following manner: The first model framework is determined based on the first function type and the first data processing rate; Obtain the first dependency library related to the first model framework; The first model tool of the first model framework is determined based on the first data transmission rate, the first model accuracy, and the first dependency library.
[0062] In specific implementation, the first function type reflects the function to be implemented, and the first data processing rate reflects the real-time requirements. That is, the first model framework can be determined based on the first function type and the first data processing rate to meet the functional and real-time requirements of high real-time scenarios. Since the model framework is known, the first dependency library related to the first model framework can be obtained. The first dependency library refers to the third-party library or framework used in implementing the first model framework. These libraries provide implementations of various model frameworks, which can simplify the development process and improve the efficiency of model configuration.
[0063] Furthermore, the first model tool of the first model framework can be determined based on the first data transmission rate, the first model accuracy, and the first dependency library. That is, the corresponding model tool is determined based on the real-time performance of the model, the accuracy requirements of the model, and the corresponding dependency library, so as to ensure that the first model framework meets the model configuration requirements of the first edge device. The model deployment and model conversion are also completed based on the first model tool corresponding to the first model framework, so that the model of the first edge device can be successfully implemented.
[0064] Optionally, the above steps, determining the first model framework based on the first function type and the first data processing rate, can be implemented in the following manner: Based on the first function type, at least one model framework is determined from the preset model framework library, and each model framework corresponds to a reference data processing rate. The first model frame is determined based on the first data processing rate and a reference data processing rate corresponding to each model frame.
[0065] In practice, a pre-defined model framework library can be stored in the server device. This pre-defined model framework library can include at least one model framework, and each model framework corresponds to a reference data processing rate.
[0066] In specific implementation, based on the functional requirements of the target high real-time scenario, at least one model framework corresponding to the first functional type can be determined from the preset model framework library. Each model framework corresponds to a reference data processing rate. The first data processing rate and the reference data processing rate corresponding to each model framework are then compared. For example, the difference between the reference data processing rate and the first data processing rate can be determined to obtain at least one difference. A difference greater than 0 can be selected, and the minimum value among the differences greater than 0 can be determined. The model framework corresponding to the minimum value is then obtained to obtain the first model framework. Alternatively, for example, the absolute value of the difference between the reference data processing rate and the first data processing rate can be determined to obtain at least one absolute value. The minimum value among them can be selected, and the model framework corresponding to the minimum value is then obtained to obtain the first model framework. In this way, not only can the model framework that meets the functional requirements and real-time requirements of the high real-time scenario be quickly determined, but the performance requirements of the model framework can also be adapted to the performance of the first edge device.
[0067] Optionally, the first model tool includes: a first model deployment tool, a first model conversion tool, and a first model training tool; The first model tool for determining the first model framework based on the first data processing rate, the first model accuracy, and the first dependency library includes: The first model deployment tool and the first model training tool are determined based on the first data processing rate and the first model accuracy. The first model conversion tool is determined based on the first data processing rate and the first dependency library.
[0068] The first model tool may include at least: a first model deployment tool, a first model conversion tool, and a first model training tool. The first model tool may also include at least one of the following: a first structure visualization tool, a first experiment tracking tool, a first model lightweighting tool, etc., without limitation.
[0069] In specific implementation, a first model deployment tool can be adapted to the first data processing rate so that the first model deployment tool can operate according to the set execution path, thereby adapting it to the first data processing rate. For example, a pre-stored mapping relationship between the preset data processing rate and the model deployment tool can be used to determine the first model deployment tool corresponding to the first data processing rate. Alternatively, a first model training tool can be adapted to the first model accuracy so that it can be trained according to the corresponding test data, iteration number, and convergence condition, so that the trained model framework meets the first model accuracy. For example, a pre-stored mapping relationship between the preset model accuracy and the model training tool can be used to determine the first model training tool corresponding to the first model accuracy.
[0070] Next, the first model conversion tool can be determined based on the first data processing rate and the first dependency library. This allows us to obtain the model conversion tool set corresponding to the first dependency library. This model conversion tool set can include at least one model conversion tool, and each model conversion tool corresponds to a data processing rate. Then, we select a data processing rate that matches the first data processing rate from the data processing rates corresponding to these model conversion tools, and obtain its corresponding model conversion tool. In other words, we determine the appropriate model tool based on the model's real-time performance, the model's accuracy requirements, and the corresponding dependency library to ensure that the first model framework meets the model configuration requirements of the first edge device. We also complete the model deployment and model conversion based on the first model tool corresponding to the first model framework, so that the model of the first edge device can be successfully deployed.
[0071] Optionally, after processing the first model framework according to the first model lightweighting parameters to obtain the second model framework, the step can also be implemented in the following manner: Obtain the first calculation result and the second model accuracy of the first edge device, wherein the first calculation result includes a first probability value; Determine the threshold probability corresponding to the second model framework, and obtain the first preset range corresponding to the accuracy of the second model; Determine the difference between the first probability value and the threshold probability; When the difference is within the first preset range, the first input data corresponding to the first calculation result is obtained; Obtain the current model parameters of the second model framework; Based on the current model parameters, the first model framework is configured on the server device to obtain the third model framework; The first input data is processed according to the third model framework to obtain the second probability value; Determine a second deviation value between the second probability value and the first probability value; When the second deviation value is within the second preset range, the first calculation result is updated according to the second probability value to obtain the second calculation result, and the second calculation result is pushed to the first edge device; When the second deviation value is not within the second preset range, the second deviation value, the first probability value, and the first model accuracy are determined to determine the first model training parameters, and the first model training parameters are pushed to the first edge device.
[0072] In practice, the first edge device can be monitored in real time, or the first edge device can monitor itself in real time and then synchronize its monitoring results to the server device.
[0073] The first preset range can be preset or set by system default. For example, the first preset range can be related to the model accuracy. For instance, the mapping relationship between preset model accuracy and the preset range corresponding to the difference can be stored in advance, so that the first preset range corresponding to the second model accuracy can be determined based on the mapping relationship. The second preset range can also be preset or set by system default.
[0074] In specific implementation, the server-side device can obtain the first calculation result and the second model accuracy from the first edge device. The first calculation result includes a first probability value. The probability threshold of the model framework can be preset, serving as a dividing line for judgment. For example, a value greater than the probability threshold can be considered an anomaly, while a value less than or equal to the probability threshold can be considered normal. The probability threshold of the second model framework can then be determined. Next, the difference between the first probability value and the threshold probability is determined: difference = first probability value - threshold probability. If the difference is within a first preset range, it indicates that the approximate value of the first calculation result is questionable and the accuracy of the model result cannot be guaranteed. In this case, the first input data corresponding to the first calculation result and the current model parameters of the second model framework can be obtained. For example, if the threshold probability is 80%, and the actual probability value of the calculated result is greater than 80%, it is usually considered an anomaly. Conversely, if the probability value is less than 80%, it means there is no anomaly. However, in reality, because the difference between the actual probability value and the threshold probability is small, the anomaly is not obvious, or there may be some controversy regarding the anomaly. Furthermore, for example, if the first probability value is 80.01%, then the difference is 80.01% - 80% = 0.01%. If 0.01% is within the first preset range, it means that there is some controversy regarding the anomaly.
[0075] Next, the first model framework is configured on the server device according to the current model parameters to obtain the third model framework. That is, the same model can be configured on the server device based on the model performance of the first edge device. Of course, the corresponding computing resources can also be called based on the target performance evaluation parameters, and the third model framework can be run based on the computing resources.
[0076] Furthermore, the first input data can be processed according to the third model framework to obtain a second probability value. This involves inputting the first input data into the third model framework and then determining a second deviation value between the second and first probability values. The second deviation value characterizes the difference between the model capabilities of the edge device and the required model capabilities. The second deviation value is calculated as (second probability value - first probability value) / first probability value. If the second deviation value falls within a second preset range, it indicates that the model of the first edge device is not abnormal. Since the performance of the server device is often greater than that of the edge device, in cases of dispute, the server device's calculation result can be used. That is, the first calculation result can be updated based on the second probability value to obtain a second calculation result, which is then pushed to the first edge device. In other words, if there is a dispute regarding the calculation result of the first edge device, a model with the corresponding capabilities can be copied to the server device based on the model capabilities of the first edge device, and anomaly detection can be performed on the model based on the corresponding input data. If no anomalies are found, the calculation result of the server device can be used to further dynamically update the calculation result.
[0077] Conversely, if the second deviation value is not within the second preset range, it indicates that the model is abnormal. In this case, the second deviation value, the first probability value, and the first model accuracy can be used to determine the first model training parameters. The first model training parameters are then pushed to the first edge device so that the first edge device can use these first model training parameters to update the model parameters of the second model framework. This is the first probability value anomaly localization. The training workload is quantified based on the second deviation value and the first model accuracy, which helps to ensure model calibration efficiency, accurately locate model anomalies, and accurately quantify the model training workload.
[0078] Optionally, the above steps, determining the second deviation value, the first probability value, and the first model accuracy to determine the first model optimization strategy, can be implemented in the following manner: The first label type is determined based on the first probability value; The first test set is determined based on the first label type and the second deviation value; The training parameters of the first model are determined based on the second deviation value, the first test set, and the first model accuracy.
[0079] In practical implementation, different probability values can correspond to different label types, and different label types correspond to different model functions. A pre-stored mapping relationship between preset label types and test sets can be used. This mapping relationship allows for the determination of a reference test set corresponding to the first label type. Then, the reference test set is filtered based on a second deviation value to obtain a first test set that meets the requirements. For example, the reference test sets can be sorted based on the significance of the first label type. Significance indicates the degree to which a label belongs to a certain category; the higher the significance, the higher the probability of being identified as belonging to the category corresponding to the first label type, and vice versa. For instance, they can be sorted from low to high significance based on the first label type. The second deviation value determines the amount of test set data. For example, the larger the second deviation value, the larger the amount of test data, and vice versa. A pre-stored mapping relationship between preset deviation values and the amount of test data in the test set allows for the determination of the target amount of test data corresponding to the second deviation value. Then, the first test set corresponding to this target amount of test data is selected from the reference test set in order of increasing significance. The amount of test data can be a specific quantity, or it can be a proportion of the total data.
[0080] Next, the training parameters of the first model can be determined based on the second deviation value, the first test set, and the first model accuracy. Specifically, a pre-defined mapping relationship between the deviation value and the number of iterations can be stored. Based on this mapping relationship, the target number of iterations corresponding to the second deviation value can be determined. The second deviation value can also determine whether the propagation direction is forward or backward, and / or the number of iterations in the forward direction and the number of iterations in the backward direction. In other words, the training parameters of the first model can be determined by the target number of iterations, the training direction (forward or backward), the first test set, and the first model accuracy, so that the second model framework is trained with the target number of iterations, the training direction (forward or backward), and the first test set to meet the first model accuracy requirement. This not only locates the anomaly of the first probability value but also quantifies the training workload based on the second deviation value and the first model accuracy, helping to ensure model calibration efficiency, accurately pinpoint the model anomaly location, and precisely quantify the model training workload.
[0081] Figure 3 This is a functional unit block diagram of a lightweight model processing system 300 applicable to high real-time scenarios, as described in this application embodiment. The lightweight model processing system 300 for high real-time scenarios includes a server-side device and n edge devices, where n is a positive integer; the lightweight model processing system 300 further includes: a receiving unit 310, a determining unit 320, and a processing unit 330, wherein... The receiving unit 310 is configured to receive a model configuration request from a first edge device. The model configuration request includes first hardware resource parameters, first software resource parameters, and first data processing requirement parameters for the target high real-time scenario of the first edge device. The first data processing requirement parameters include a first data processing rate, a first function type, and a first model accuracy. The first edge device is one of the n edge devices. The determining unit 320 is used to determine the first model framework and the first model tool according to the first data processing requirement parameters; The processing unit 330 is configured to determine first model lightweighting processing parameters based on the first hardware resource parameters, the first software resource parameters, and the first model framework; and to process the first model framework based on the first model lightweighting processing parameters to obtain a second model framework; the second model framework and the first model tool are used to configure corresponding model services on the first edge device.
[0082] Optionally, in determining the first model lightweighting processing parameters based on the first hardware resource parameters, the first software resource parameters, and the first model framework, the determining unit 320 is specifically used for: The first performance evaluation parameter is determined based on the first hardware resource parameter; Determine the second performance evaluation parameters based on the first software resource parameters; Determine the target performance evaluation parameters based on the first performance evaluation parameters and the second performance evaluation parameters; Determine the reference performance evaluation parameters corresponding to the first data processing rate and the first model framework; Determine a first deviation value between the target performance evaluation parameter and the reference performance evaluation parameter; The first model lightweighting parameters are determined based on the first deviation value, the first performance evaluation parameter, and the second performance evaluation parameter.
[0083] It is understood that the functions of the lightweight model processing system in this embodiment, which can be used in high real-time scenarios, can be specifically implemented according to the methods in the above method embodiments. The specific implementation process can be referred to the relevant descriptions in the above method embodiments, and will not be repeated here.
[0084] Please see Figure 4 , Figure 4This is a schematic diagram of a server-side device provided in an embodiment of this application. The server-side device includes a processor, a memory, a communication interface, and one or more programs, applied to a lightweight model processing system suitable for high real-time scenarios. This lightweight model processing system for high real-time scenarios includes not only the server-side device but also n edge devices, where n is a positive integer. The one or more programs are stored in the memory and configured to be executed by the processor. In this embodiment, the programs include instructions for performing the following steps: The system receives a model configuration request from a first edge device. The model configuration request includes first hardware resource parameters, first software resource parameters, and first data processing requirement parameters for the target high real-time scenario of the first edge device. The first data processing requirement parameters include a first data processing rate, a first function type, and a first model accuracy. The first edge device is one of the n edge devices. The first model framework and the first model tool are determined based on the first data processing requirement parameters; The first model lightweighting parameters are determined based on the first hardware resource parameters, the first software resource parameters, and the first model framework. The first model framework is processed according to the first model lightweighting processing parameters to obtain the second model framework; the second model framework and the first model tool are used to configure the corresponding model services on the first edge device.
[0085] Optionally, in determining the first model lightweighting processing parameters based on the first hardware resource parameters, the first software resource parameters, and the first model framework, the above procedure includes instructions for performing the following steps: The first performance evaluation parameter is determined based on the first hardware resource parameter; Determine the second performance evaluation parameters based on the first software resource parameters; Determine the target performance evaluation parameters based on the first performance evaluation parameters and the second performance evaluation parameters; Determine the reference performance evaluation parameters corresponding to the first data processing rate and the first model framework; Determine a first deviation value between the target performance evaluation parameter and the reference performance evaluation parameter; The first model lightweighting parameters are determined based on the first deviation value, the first performance evaluation parameter, and the second performance evaluation parameter.
[0086] Optionally, in determining the first model lightweighting parameters based on the first deviation value, the first performance evaluation parameter, and the second performance evaluation parameter, the above procedure includes instructions for performing the following steps: A first model lightweighting strategy set is determined based on the first performance evaluation parameter; the first model lightweighting strategy set includes at least one model lightweighting strategy; each model lightweighting strategy corresponds to a target second performance evaluation parameter; The target model lightweighting strategy is determined from the first model lightweighting strategy set based on the second performance evaluation parameter and a target second performance evaluation parameter corresponding to each model lightweighting strategy. The first model lightweighting processing parameters corresponding to the target model lightweighting strategy are determined based on the first deviation value.
[0087] Optionally, in determining the first model framework and the first model tool based on the first data processing requirement parameters, the above procedure includes instructions for performing the following steps: The first model framework is determined based on the first function type and the first data processing rate; Obtain the first dependency library related to the first model framework; The first model tool of the first model framework is determined based on the first data transmission rate, the first model accuracy, and the first dependency library.
[0088] Optionally, in determining the first model framework based on the first function type and the first data processing rate, the above procedure includes instructions for performing the following steps: Based on the first function type, at least one model framework is determined from the preset model framework library, and each model framework corresponds to a reference data processing rate. The first model frame is determined based on the first data processing rate and a reference data processing rate corresponding to each model frame.
[0089] Optionally, the first model tool includes: a first model deployment tool, a first model conversion tool, and a first model training tool; Regarding the first model tool for determining the first model framework based on the first data processing rate, the first model accuracy, and the first dependency library, the above program includes instructions for performing the following steps: The first model deployment tool and the first model training tool are determined based on the first data processing rate and the first model accuracy. The first model conversion tool is determined based on the first data processing rate and the first dependency library.
[0090] Optionally, after processing the first model framework according to the first model lightweighting parameters to obtain the second model framework, the above program further includes instructions for performing the following steps: Obtain the first calculation result and the second model accuracy of the first edge device, wherein the first calculation result includes a first probability value; Determine the threshold probability corresponding to the second model framework, and obtain the first preset range corresponding to the accuracy of the second model; Determine the difference between the first probability value and the threshold probability; When the difference is within the first preset range, the first input data corresponding to the first calculation result is obtained; Obtain the current model parameters of the second model framework; Based on the current model parameters, the first model framework is configured on the server device to obtain the third model framework; The first input data is processed according to the third model framework to obtain the second probability value; Determine a second deviation value between the second probability value and the first probability value; When the second deviation value is within the second preset range, the first calculation result is updated according to the second probability value to obtain the second calculation result, and the second calculation result is pushed to the first edge device; When the second deviation value is not within the second preset range, the second deviation value, the first probability value, and the first model accuracy are determined to determine the first model training parameters, and the first model training parameters are pushed to the first edge device.
[0091] Optionally, in determining the first model optimization strategy based on the second deviation value, the first probability value, and the first model accuracy, the above procedure includes instructions for performing the following steps: The first label type is determined based on the first probability value; The first test set is determined based on the first label type and the second deviation value; The training parameters of the first model are determined based on the second deviation value, the first test set, and the first model accuracy.
[0092] Please see Figure 5 This application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement some or all of the steps of the model lightweight processing method for high real-time scenarios in any of the above embodiments. The computer-readable storage medium is located at a server device.
[0093] This application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps of any of the methods described in the above method embodiments. The computer program product may be a software installation package, and the computer may include an electronic device.
[0094] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0095] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0096] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical or other forms.
[0097] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0098] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0099] If the aforementioned integrated units are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0100] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage device, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0101] The embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A model lightweight processing method for a high real-time scenario, characterized in that, The method is applied to a server device in a model lightweight processing system applicable to a high real-time scene, the system further comprising: n edge devices, n being a positive integer; the method comprising: receiving a model configuration request of a first edge device, the model configuration request comprising first hardware resource parameters, first software resource parameters of the first edge device, and first data processing requirement parameters for a target high real-time scene; the first data processing requirement parameters comprising a first data processing rate, a first function type, and a first model precision; the first edge device being one of the n edge devices; determining a first model framework and a first model tool according to the first data processing requirement parameters; determining first model lightweight processing parameters according to the first hardware resource parameters, the first software resource parameters, and the first model framework; processing the first model framework according to the first model lightweight processing parameters to obtain a second model framework; the second model framework and the first model tool being used for the first edge device to configure a corresponding model service.
2. The model lightweight processing method for high real-time scenarios according to claim 1, characterized in that, The determining of the first model lightweight processing parameters according to the first hardware resource parameters, the first software resource parameters, and the first model framework comprises: determining first performance evaluation parameters according to the first hardware resource parameters; determining second performance evaluation parameters according to the first software resource parameters; determining target performance evaluation parameters according to the first performance evaluation parameters and the second performance evaluation parameters; determining reference performance evaluation parameters corresponding to the first data processing rate and the first model framework; determining a first deviation value between the target performance evaluation parameters and the reference performance evaluation parameters; determining the first model lightweight processing parameters according to the first deviation value, the first performance evaluation parameters, and the second performance evaluation parameters.
3. The model lightweight processing method for high real-time scenarios according to claim 2, characterized in that, The determining of the first model lightweight processing parameters according to the first deviation value, the first performance evaluation parameters, and the second performance evaluation parameters comprises: determining a first model lightweight strategy set according to the first performance evaluation parameters; the first model lightweight strategy set comprising at least one model lightweight strategy; each model lightweight strategy corresponding to a target second performance evaluation parameter; determining a target model lightweight strategy from the first model lightweight strategy set according to the second performance evaluation parameters and each model lightweight strategy corresponding to a target second performance evaluation parameter; determining the first model lightweight processing parameters corresponding to the target model lightweight strategy according to the first deviation value.
4. The model lightweight processing method for high real-time scenarios according to any one of claims 1-3, characterized in that, The determining of the first model framework and the first model tool according to the first data processing requirement parameters comprises: determining the first model framework according to the first function type and the first data processing rate; obtaining a first dependent library related to the first model framework; determining the first model tool of the first model framework according to the first data transmission rate, the first model precision, and the first dependent library.
5. The model lightweight processing method for high real-time scenarios according to claim 4, characterized in that, The determining the first model framework according to the first function type and the first data processing rate comprises: determining at least one model framework from a preset model framework library according to the first function type, each model framework corresponding to a reference data processing rate; determining the first model framework according to the first data processing rate and the reference data processing rate corresponding to each model framework.
6. The model lightweight processing method for high real-time scenarios according to claim 4, characterized in that, The first model tool comprises a first model deployment tool, a first model conversion tool and a first model training tool. The determining the first model framework according to the first data processing rate, the first model precision and the first dependent library comprises: determining the first model deployment tool and the first model training tool according to the first data processing rate and the first model precision; determining the first model conversion tool according to the first data processing rate and the first dependent library.
7. The model lightweight processing method for high real-time scenarios according to any one of claims 1-3, characterized in that, After the processing the first model framework according to the first model lightweight processing parameter to obtain a second model framework, the method further comprises: obtaining a first operation result and a second model precision of the first edge end device, the first operation result comprising a first probability value; determining a threshold probability corresponding to the second model framework and obtaining a first preset range corresponding to the second model precision; determining a difference value between the first probability value and the threshold probability; when the difference value is in the first preset range, obtaining first input data corresponding to the first operation result; obtaining a current model parameter of the second model framework; configuring the first model framework according to the current model parameter in the service end device to obtain a third model framework; performing operation on the first input data according to the third model framework to obtain a second probability value; determining a second deviation degree value between the second probability value and the first probability value; when the second deviation degree value is in a second preset range, updating the first operation result according to the second probability value to obtain a second operation result, and pushing the second operation result to the first edge end device; when the second deviation degree value is not in the second preset range, determining a first model training parameter according to the second deviation degree value, the first probability value and the first model precision, and pushing the first model training parameter to the first edge end device.
8. The model lightweight processing method for high real-time scenarios according to claim 7, characterized in that, The determining the first model optimization strategy according to the second deviation degree value, the first probability value and the first model precision comprises: determining a first label type according to the first probability value; determining a first test set according to the first label type and the second deviation degree value; determining the first model training parameter according to the second deviation degree value, the first test set and the first model precision. 9.A model lightweight processing system for a high real-time scenario, characterized in that, The model lightweight processing system applicable to high real-time scenarios comprises a service end device and n edge end devices, n being a positive integer; the model lightweight processing system applicable to high real-time scenarios further comprises a receiving unit, a determining unit and a processing unit, wherein, The receiving unit is configured to receive a model configuration request of a first edge device, the model configuration request comprising first hardware resource parameters, first software resource parameters and first data processing requirement parameters of the first edge device for a target high real-time scenario; the first data processing requirement parameters comprising a first data processing rate, a first function type and a first model precision; the first edge device being one of the n edge devices; The determining unit is configured to determine a first model framework and a first model tool according to the first data processing requirement parameters; The processing unit is configured to determine first model lightweight processing parameters according to the first hardware resource parameters, the first software resource parameters and the first model framework; and process the first model framework according to the first model lightweight processing parameters to obtain a second model framework; the second model framework and the first model tool being used for the first edge device to configure a corresponding model service.
10. The model lightweight processing system for high real-time scenarios according to claim 9, characterized in that, In the aspect of determining the first model lightweight processing parameters according to the first hardware resource parameters, the first software resource parameters and the first model framework, the determining unit is specifically configured to: determine first performance evaluation parameters according to the first hardware resource parameters; determine second performance evaluation parameters according to the first software resource parameters; determine target performance evaluation parameters according to the first performance evaluation parameters and the second performance evaluation parameters; determine reference performance evaluation parameters corresponding to the first data processing rate and the first model framework; determine a first deviation value between the target performance evaluation parameters and the reference performance evaluation parameters; determine the first model lightweight processing parameters according to the first deviation value, the first performance evaluation parameters and the second performance evaluation parameters.