Artificial intelligence model service method and system based on edge calculation
By performing data feature analysis and model selection on edge computing nodes, combining computing priority and resource management, latency and security issues in traditional cloud computing methods are solved, and efficient and secure edge computing artificial intelligence model services are realized.
Patent Information
- Application Number
- CN202510630339.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-16
AI Technical Summary
Traditional cloud-based AI model service methods have network latency, difficult to meet real-time requirements, and privacy and security issues, and how to effectively utilize edge computing resources to process data efficiently is still a challenge.
By receiving the mobile terminal's artificial intelligence model service request on the edge computing node, performing data feature analysis, dividing the data subset and selecting the matching target artificial intelligence model for processing, determining the calculation priority based on the calculation complexity and resource status, and summarizing the processing model calculation results.
Improve the accuracy and efficiency of data processing, make full use of edge computing resources, reduce network transmission delay, improve response speed, and enhance data security.
Smart Images

Figure CN120144327A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to artificial intelligence technology, and particularly to an artificial intelligence model service method and system based on edge computing. Background Art
[0002] With the rapid development of artificial intelligence technology, more and more mobile applications begin to integrate artificial intelligence functions to provide more intelligent and personalized services. However, due to the limited computing resources and storage space of mobile terminals, it is difficult to run complex artificial intelligence models locally. Therefore, usually, cloud computing is adopted to deploy artificial intelligence models on remote servers, and mobile terminals obtain model services through network requests.
[0003] However, there are some problems with traditional artificial intelligence model service methods based on cloud computing. First, transmitting a large amount of raw data to a remote server for processing will increase network latency and affect the user experience. Second, the centralized cloud computing mode is difficult to meet the requirements of mobile applications for real-time performance and low latency, especially in the case of poor network conditions. Finally, transmitting all data to the cloud for processing may cause privacy and security problems, and users' sensitive information may be leaked or misused.
[0004] To solve these problems, edge computing, as an emerging computing paradigm, is introduced into artificial intelligence model services. Edge computing sinks computing resources and storage resources to the network edge, close to data sources and users, which can quickly process data locally, reduce network transmission latency, and improve service response speed. However, how to effectively utilize edge computing resources, select appropriate artificial intelligence models for different types of data, and how to efficiently process a large amount of data under limited edge computing resources are still key problems to be solved. Summary of the Invention
[0005] Embodiments of the present invention provide an artificial intelligence model service method and system based on edge computing, which can solve the problems in the prior art.
[0006] In the first aspect of the embodiments of the present invention, There is provided an artificial intelligence model service method based on edge computing, including: Receiving an artificial intelligence model service request sent by a mobile terminal, where the artificial intelligence model service request includes raw data to be processed; Performing data feature analysis on the raw data based on an edge computing node, dividing the raw data into multiple data subsets according to the data features of the raw data, and selecting a corresponding target artificial intelligence model from a preset artificial intelligence model library for each data subset, so that each data subset is processed by a target artificial intelligence model matching its data features; The edge computing node determines the computing priority of each target artificial intelligence model according to the computing complexity of each target artificial intelligence model and the computing resource status of the edge computing node, and sequentially calls the computing resources of the edge computing node to execute each target artificial intelligence model according to the computing priority, and generates corresponding model computing results; The edge computing node aggregates and processes the model computing results of each target artificial intelligence model to obtain a processing result for the original data, and returns the processing result to the mobile terminal.
[0007] Dividing the original data into multiple data subsets according to the data characteristics of the original data, and selecting corresponding target artificial intelligence models from a preset artificial intelligence model library based on each data subset, so that each data subset is processed by a target artificial intelligence model matching its data characteristics includes: Extracting the time series features, statistical features and information entropy features in the original data to construct a feature vector; calculating the distance between the data points of the original data and other data points according to the feature vector, and obtaining a local density value according to the relationship between the distance and a preset truncation distance; Performing clustering division on the original data based on the local density value to obtain multiple data subsets; Constructing a matching degree score based on the weighted sum of the feature vectors of multiple data subsets and a preset model feature descriptor; constructing a priority queue for the data subsets, the corresponding model feature descriptors and the matching degree scores; Determining the optimal matching relationship between each data subset and the model by minimizing the allocation cost, and matching each data subset with a model matching its data characteristics based on the optimal matching relationship.
[0008] Calculating the distance between the data points of the original data according to the feature vector, and obtaining a local density value according to the relationship between the distance and a preset truncation distance includes: Calculating the distance between the data points in the original data, where the distance is obtained by weighted summation of a dedicated distance function for each feature dimension; Calculating the structural similarity degree between the data point and its neighboring data points, multiplying the structural similarity degree by an adjustment coefficient and then adding one to obtain a structure-aware density value; Calculating the arithmetic mean of the structure-aware density values based on the adjustment coefficient by arithmetic mean, and dividing the arithmetic mean by the structure-aware density value to obtain a local anomaly factor; Multiply the difference between the local anomaly factor and the preset anomaly threshold by the attenuation coefficient, take the negative of the multiplication result and perform an exponential operation to obtain an attenuation factor, and multiply the attenuation factor by the structure-aware density value to obtain a corrected density value; Perform a weighted combination of the corrected density value and a preset density coefficient to obtain a multi-scale density representation, calculate the product of the smoothing coefficient and the multi-scale density representation, and the product of the smoothing coefficient and the weighted multi-scale density value of neighboring data points, and add the two products to obtain a local density value.
[0009] The edge computing node determines the computing priority of each target artificial intelligence model according to the computing complexity of each target artificial intelligence model and the computing resource status of the edge computing node, and sequentially calls the computing resources of the edge computing node to execute each target artificial intelligence model according to the computing priority, and the generated corresponding model calculation results include: Obtain the computing complexity by performing a weighted sum of the products of the operands and layer weight coefficients in each computing layer of the target artificial intelligence model based on the edge computing node; Construct a resource status vector including the processor occupancy rate, memory occupancy rate, and bandwidth occupancy rate, and determine the resource availability based on the resource status vector according to the minimum value of the ratio of the remaining amount to the maximum capacity of each type of resource; Multiply the normalized result of the computing complexity and the resource availability by the corresponding weight coefficients respectively to obtain a priority score; construct a multi-objective optimization function based on the priority score, and the multi-objective optimization function includes target constraint conditions, resource constraint conditions, and latency constraint conditions; Determine the execution order of the target artificial intelligence models according to the solution result of the multi-objective optimization function, and sequentially call the computing resources of the edge computing node to execute each target artificial intelligence model according to the execution order, and generate corresponding model calculation results, including: Solve the multi-objective optimization function to obtain the execution order, and the solution process of the multi-objective optimization function is restricted by the resource constraint condition that the resource usage does not exceed the available resource capacity, and the latency constraint condition that the model execution time does not exceed the maximum allowable latency; Sequentially call the computing resources of the edge computing node to execute each target artificial intelligence model according to the execution order, and generate corresponding model calculation results.
[0010] The edge computing node aggregates the model calculation results of each target artificial intelligence model to obtain a processing result for the original data, including: Construct a credibility evaluation function for the model results. The credibility evaluation function is obtained based on the weighted combination of the historical accuracy of each target artificial intelligence model, the model execution quality under the current resource state, and the result output time. Apply the credibility evaluation function to the model calculation results of each of the target artificial intelligence models; Perform adaptive weighting on the model calculation results of each of the models based on the credibility evaluation function. The weight coefficient of the adaptive weighting is dynamically adjusted according to the change of the credibility evaluation function. The weight coefficient ensures that the sum of the weights of the model calculation results of each of the models is one through normalization processing; Adopt a sliding window mechanism based on time series correlation for the weighted model calculation results of each of the models, and combine the weighted results within the sliding window to obtain the processing result for the original data.
[0011] In the second aspect of the embodiments of the present invention, Provide an artificial intelligence model service system based on edge computing, including: A first unit for receiving an artificial intelligence model service request sent by a mobile terminal. The artificial intelligence model service request includes the original data to be processed; A second unit for performing data feature analysis on the original data based on an edge computing node, dividing the original data into multiple data subsets according to the data features of the original data, and selecting corresponding target artificial intelligence models from a preset artificial intelligence model library based on each data subset, so that each data subset is processed by a target artificial intelligence model that matches its data features; A third unit for the edge computing node to determine the calculation priority of each target artificial intelligence model according to the calculation complexity of each target artificial intelligence model and the calculation resource state of the edge computing node, and sequentially call the calculation resources of the edge computing node to execute each target artificial intelligence model to generate corresponding model calculation results; A fourth unit for the edge computing node to summarize and process the model calculation results of each of the target artificial intelligence models to obtain the processing result for the original data, and return the processing result to the mobile terminal.
[0012] In the third aspect of the embodiments of the present invention, Provide an electronic device, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0013] In the fourth aspect of the embodiments of the present invention, Provided is a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the foregoing method is implemented.
[0014] The beneficial effects of this application are as follows: 1. By performing feature analysis on the original data and dividing it into multiple data subsets, and selecting a matching target artificial intelligence model for each subset for processing, the accuracy and efficiency of data processing can be improved, and the advantages of different models can be fully utilized to process data with different features.
[0015] 2. By determining the calculation priority according to the computational complexity of the target artificial intelligence model and the resource status of the edge computing node, reasonable allocation and scheduling of computing resources can be achieved, the resource utilization rate of the edge computing node can be improved, and resource waste or over-occupation can be avoided.
[0016] 3. Completing data processing on the edge computing node and returning the result to the mobile terminal reduces the data transmission volume and network latency, improves the response speed, and at the same time protects user privacy and enhances data security. This method gives full play to the advantages of edge computing and provides efficient, secure, and personalized artificial intelligence services for the mobile terminal. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic flowchart of the method for artificial intelligence model service based on edge computing according to an embodiment of the present invention; Figure 2 It is a flowchart for determining the optimal matching relationship based on the original data according to an embodiment of the present invention; Figure 3 It is a flowchart for scheduling artificial intelligence models in an edge computing environment according to an embodiment of the present invention; Figure 4 It is a flowchart for fusing multi-model results based on credibility evaluation according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0019] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0020] Figure 1 This is a schematic flowchart of the artificial intelligence model service method based on edge computing according to an embodiment of the present invention. As Figure 1 shown, the method includes: Receiving an artificial intelligence model service request sent by a mobile terminal, where the artificial intelligence model service request includes raw data to be processed; Based on an edge computing node, performing data feature analysis on the raw data, dividing the raw data into multiple data subsets according to the data features of the raw data, and selecting a corresponding target artificial intelligence model from a preset artificial intelligence model library based on each data subset, so that each data subset is processed by a target artificial intelligence model matching its data features; The edge computing node determines the calculation priority of each target artificial intelligence model according to the calculation complexity of each target artificial intelligence model and the calculation resource status of the edge computing node, and sequentially calls the calculation resources of the edge computing node to execute each target artificial intelligence model according to the calculation priority, generating corresponding model calculation results; The edge computing node performs summary processing on the model calculation results of each target artificial intelligence model to obtain a processing result for the raw data, and returns the processing result to the mobile terminal.
[0021] In an alternative embodiment, dividing the raw data into multiple data subsets according to the data features of the raw data, and selecting a corresponding target artificial intelligence model from a preset artificial intelligence model library based on each data subset, so that each data subset is processed by a target artificial intelligence model matching its data features includes: Extracting the time series feature, statistical feature, and information entropy feature in the raw data to construct a feature vector; calculating the distance between the data points of the raw data and other data points according to the feature vector, and obtaining a local density value according to the relationship between the distance and a preset truncation distance; Based on the local density value, performing clustering division on the raw data to obtain multiple data subsets; Based on the weighted sum of the feature vectors of multiple data subsets and a preset model feature descriptor, constructing a matching degree score; constructing a priority queue with the data subsets, the corresponding model feature descriptors, and the matching degree score; Determining the optimal matching relationship between each data subset and the model by minimizing the allocation cost, and performing matching on each data subset with a model matching its data features based on the optimal matching relationship.
[0022] The original data is divided into multiple data subsets according to the data characteristics of the original data, and the corresponding target artificial intelligence models are selected from a preset artificial intelligence model library based on each data subset, so that each data subset is processed by a target artificial intelligence model that matches its data characteristics.
[0023] Extract the time series features, statistical features, and information entropy features in the original data to construct a feature vector. The time series features include the periodicity, trend, seasonality, etc. of the data, which can be extracted by methods such as autocorrelation function and Fourier transform; the statistical features include features such as mean, variance, skewness, and kurtosis that describe the data distribution; the information entropy feature is used to measure the uncertainty and complexity of the data and is obtained by calculating the Shannon entropy.
[0024] For a set of original data containing 1000 temperature sensor records, extracting its time series features can obtain an obvious daily change pattern with a period of 24 hours and an autocorrelation coefficient of 0.82; the statistical features include a mean of 23.5°C, a standard deviation of 3.2°C, a skewness of -0.18, and a kurtosis of 2.76; the information entropy feature is 4.36. These features are combined into a feature vector [0.82, 24, 23.5, 3.2, -0.18, 2.76, 4.36] as a comprehensive description of the characteristics of this data set.
[0025] According to the feature vector, calculate the distance between the data points of the original data and other data points, and obtain the local density value according to the relationship between the distance and the preset truncation distance. Specifically, select the Euclidean distance, Manhattan distance, Mahalanobis distance, etc. as the measurement standard for the distance between data points. Set the truncation distance threshold, for example, the 10% quantile of the distances of all data point pairs. For each data point, calculate its distance from all other data points, and count the number of points less than the truncation distance, which is the local density value of this point.
[0026] In the above temperature sensor data case, assume that the Euclidean distance is selected as the measurement standard, and the calculated truncation distance is 1.5. For data point 1, 78 of its distances from other points are less than 1.5, so its local density value is 78; for data point 2, 45 of its distances from other points are less than 1.5, so its local density value is 45, and so on to calculate the local density values of all data points.
[0027] Based on the local density values, cluster and partition the original data to obtain multiple data subsets. Use the density peak clustering algorithm to first identify the density peak points (points with relatively high local density values and far distances from other high-density points) as the clustering centers, and then assign the remaining points to the nearest clustering center. To avoid the influence of outliers, density thresholds and distance thresholds can be set to screen the clustering centers.
[0028] In the temperature sensor data case, 3 clustering centers were identified through the density peak clustering algorithm, corresponding to data points with density values of 92, 83, and 76 respectively. After assigning the remaining data points to the nearest clustering center, 3 data subsets were obtained, with sizes of 420, 350, and 230 records respectively. Subset 1 mainly contains daytime temperature data, subset 2 mainly contains nighttime temperature data, and subset 3 mainly contains temperature transition period data.
[0029] The matching degree score is composed of the weighted sum of the feature vectors of multiple data subsets and the preset model feature descriptors. The preset model feature descriptors are descriptions of the data characteristics applicable to the artificial intelligence model, including quantization values in dimensions such as time series, stationarity, periodicity, and volatility. The dot product operation is performed on the feature vectors of the data subsets and the model feature descriptors, and a weight coefficient is introduced to reflect the importance of each feature dimension, resulting in the matching degree score.
[0030] There are three models in the artificial intelligence model library: The LSTM model is suitable for processing strong time series data, and its model feature descriptor is [0.9, 0.5, 0.7, 0.3]; The ARIMA model is suitable for processing stationary time series data, and its model feature descriptor is [0.6, 0.8, 0.4, 0.2]; The Prophet model is suitable for processing seasonal data, and its model feature descriptor is [0.5, 0.6, 0.9, 0.4]. The feature weights are set to [0.4, 0.3, 0.2, 0.1].
[0031] For data subset 1, its feature vector is simplified to [0.85, 0.6, 0.75, 0.3], and the calculated matching degree score with the LSTM model is: 0.85×0.9×0.4 + 0.6×0.5×0.3 + 0.75×0.7×0.2 + 0.3×0.3×0.1 = 0.306 + 0.09 + 0.105 + 0.009 = 0.51; The matching degree score with the ARIMA model is: 0.85×0.6×0.4 + 0.6×0.8×0.3 + 0.75×0.4×0.2 + 0.3×0.2×0.1 = 0.204 + 0.144 + 0.06 + 0.006 = 0.414; The matching degree score with the Prophet model is: 0.85×0.5×0.4 + 0.6×0.6×0.3 + 0.75×0.9×0.2 + 0.3×0.4×0.1 = 0.17 + 0.108 + 0.135 + 0.012 = 0.425.
[0032] Construct a priority queue with data subsets, corresponding model feature descriptors, and matching score. The priority queue is sorted in descending order of the matching score, and the queue elements include data subset identifiers, model identifiers, and matching scores.
[0033] Determine the optimal matching relationship between each data subset and the model by minimizing the assignment cost. Based on the optimal matching relationship, match each data subset with the model that matches its data features. The assignment cost is defined as 1 minus the matching score. Establish a bipartite graph matching problem and solve it using the Hungarian algorithm to obtain the overall optimal matching scheme of data subsets and models.
[0034] The optimal matching scheme obtained by solving with the Hungarian algorithm: Data subset 1 is matched with the LSTM model (cost 0.49), data subset 2 is matched with the ARIMA model (cost 0.52), data subset 3 is matched with the Prophet model (cost 0.47), and the total cost is 1.48, which is the minimum cost among all possible matching schemes.
[0035] According to the optimal matching scheme, hand over 420 daytime temperature data of data subset 1 to the LSTM model for processing, 350 nighttime temperature data of data subset 2 to the ARIMA model for processing, and 230 temperature transition period data of data subset 3 to the Prophet model for processing. This matching relationship makes full use of the strengths of different models. The LSTM model is good at capturing the complex nonlinear relationships in daytime temperature data, the ARIMA model is suitable for dealing with relatively stable nighttime temperature changes, and the Prophet model is good at dealing with data with obvious transition characteristics.
[0036] Figure 2 Flowchart for determining the optimal matching relationship based on the original data in the embodiment of the present invention: This figure shows a complete data processing and model matching flowchart, starting from the raw data input at the top, passing through multiple key processing steps, and finally achieving the optimized matching of data and model. The specific process includes: First, extract features from the raw data, including time series features, statistical features, and information entropy features, and construct feature vectors; then calculate the distance relationships and local density values between data points to evaluate the distribution characteristics of the data through these metrics; perform data clustering based on the obtained local density values to divide similar data into the same subset; then calculate the matching degree between the feature vectors of each data subset and the model features in the artificial intelligence model library, and evaluate using the matching score composed of weighted sums; construct a priority queue between the data subset and the model according to the matching score, which reflects the matching priority relationship between data and model; finally, determine the optimal matching relationship between the data subset and the model by minimizing the allocation cost. The entire process reflects a systematic processing process from data feature extraction to model matching. Through local density calculation and the construction of a priority queue, the intelligent matching of data and model is achieved. The main innovation of this process lies in the organic combination of data feature extraction, density analysis, and model matching, forming a complete intelligent processing framework.
[0037] Existing data processing methods usually use a unified artificial intelligence model to process the entire data set or select models based on simple rules (such as the size of the data volume, data type), without fully considering the matching relationship between the internal characteristics of the data and the applicable scenarios of the model. For example, some methods only select models based on the source type of the data, uniformly using time series models for temperature data, ignoring that temperature data in different periods may have different characteristics; there are also methods that only select the model complexity based on the size of the data volume, choosing deep learning models for large data volumes and statistical learning models for small data volumes. This rough division method is difficult to achieve the optimal matching.
[0038] This application proposes a fine-grained data partitioning and model matching method based on data features. Compared with the prior art, the main improvements include: First, introducing multi-dimensional feature vectors to represent data characteristics, including time series features, statistical features, and information entropy features, to make the description of data features more comprehensive; second, using the density peak clustering algorithm for data partitioning, which can better identify data clusters with irregular shapes compared to traditional methods such as K-means; third, constructing a matching degree evaluation mechanism for data feature and model feature descriptors to obtain accurate matching scores through weighted calculation; fourth, transforming the model allocation problem into a bipartite graph optimal matching problem to obtain a globally optimal matching scheme by minimizing the overall allocation cost. These improvements make data processing more refined, enabling the selection of the most suitable model according to the internal characteristics of the data, thereby improving the model processing effect and resource utilization efficiency.
[0039] In an alternative embodiment, obtaining the local density value according to the relationship between the distance and the preset truncation distance by calculating the distance between the data points of the original data and other data points based on the feature vector includes: Calculating the distance between data points in the original data, where the distance is obtained by weighted summation of dedicated distance functions for each feature dimension; Calculating the structural similarity degree between the data point and its neighboring data points, and adding 1 after multiplying the structural similarity degree by an adjustment coefficient to obtain a structure-aware density value; Calculating the arithmetic mean of the structure-aware density values by arithmetic mean based on the adjustment coefficient, and dividing the arithmetic mean by the structure-aware density value to obtain a local outlier factor; Multiplying the difference between the local outlier factor and the preset outlier threshold by an attenuation coefficient, performing an exponential operation on the negative of the multiplication result to obtain an attenuation factor, and multiplying the attenuation factor by the structure-aware density value to obtain a corrected density value; Performing a weighted combination of the corrected density value and a preset density coefficient to obtain a multi-scale density representation, calculating the product of a smoothing coefficient and the multi-scale density representation, and the product of the smoothing coefficient and the weighted multi-scale density value of neighboring data points, and adding the two products to obtain a local density value.
[0040] Calculate the distance between data points in the original data. The distance here is not a simple Euclidean distance, but is obtained by weighted summation of dedicated distance functions for each feature dimension. Specifically, for each feature dimension, a suitable distance function is selected according to the data characteristics of that dimension. For example, the Euclidean distance can be used for numerical features, and the Hamming distance can be used for categorical features, etc. Then, the distances of each dimension are weighted and summed, and the weights can be set according to the importance of each feature. For example, assuming there are 3 feature dimensions with weights of 0.5, 0.3, and 0.2 respectively, the final distance can be expressed as: 0.5×Feature 1 distance + 0.3×Feature 2 distance + 0.2×Feature 3 distance.
[0041] Calculate the structural similarity degree between each data point and its neighboring data points. The structural similarity degree here reflects the distribution characteristics of data points in the local area. The structural similarity degree can be measured by comparing the distance distributions of data points and their neighboring points. For example, the average distance from a data point to its k nearest neighbor points can be calculated, and then compared with the average distance from each of these k neighboring points to their k nearest neighbors to obtain a similarity score. Then, multiplying this structural similarity by an adjustment coefficient and adding 1 to obtain a structure-aware density value. The adjustment coefficient is used to control the influence degree of the structural similarity on the density and can be adjusted according to the specific application scenario.
[0042] Calculate the arithmetic mean of the structure-aware density values based on the above adjustment coefficient. Specifically, the sum of the structure-aware density values of all data points can be calculated and then divided by the total number of data points to obtain the average value. Divide this arithmetic mean by the structure-aware density value of each data point itself to obtain the local outlier factor. The local outlier factor reflects the degree of abnormality of a data point relative to the whole, and the larger the value, the more abnormal it is.
[0043] Subtract the local outlier factor from the preset outlier threshold, and multiply the difference by a decay coefficient. This decay coefficient is used to control the impact of the degree of abnormality on the final density. Take the negative of the multiplication result and perform an exponential operation to obtain the decay factor. The role of the decay factor is to adjust the density of the outlier points. Multiply the decay factor by the previously obtained structure-aware density value to obtain the corrected density value. The purpose of this step is to reduce the density value of the outlier points so that they can be more easily identified in subsequent clustering or outlier detection.
[0044] Perform a weighted combination of the corrected density value and the preset density coefficient to obtain a multi-scale density representation. Here, multi-scale means considering the density information in neighborhoods of different ranges. Multiple density coefficients can be set, corresponding to neighborhoods of different sizes respectively, and then a weighted combination is performed. For example, three scales of near neighbors, medium neighbors, and far neighbors can be set, with weights of 0.5, 0.3, and 0.2 respectively. Then calculate the product of a smoothing coefficient and this multi-scale density representation, as well as the product of the smoothing coefficient and the weighted multi-scale density value of neighboring data points. The role of the smoothing coefficient is to locally smooth the density and reduce the influence of noise. Finally, add these two products together to obtain the final local density value.
[0045] Illustrate this process through a specific data case. Suppose there is a two-dimensional data set containing 1000 sample points, and each sample point has two features x and y. First, calculate the distance between sample points. Here, it is assumed that the weights of x and y are 0.6 and 0.4 respectively. For sample points A(1, 2) and B(4, 6), the distance between them can be calculated as: 0.6×|4 - 1| + 0.4×|6 - 2| = 3.4.
[0046] Calculate the structure similarity. Suppose 10 nearest neighbors are selected. For sample point A, calculate the average distance from it to its 10 nearest neighbors as 2.5, and the average value of the average distances from these 10 neighboring points to their 10 nearest neighbors is 3.0. Then the structure similarity of A can be expressed as 2.5 / 3.0 = 0.833. Suppose the adjustment coefficient is 0.5, then the structure-aware density value of A is 1 + 0.5×0.833 = 1.4165.
[0047] Calculate the arithmetic mean of the structure-aware density values of all sample points, which is assumed to be 1.5. Then the local outlier factor of A is 1.5 / 1.4165 = 1.059. Assume that the preset outlier threshold is 1.2 and the attenuation coefficient is 0.1, then the attenuation factor of A is exp(-0.1×(1.2 - 1.059)) = 0.986. Multiply this attenuation factor by the structure-aware density value of A to obtain the corrected density value 1.4165×0.986 = 1.397.
[0048] Assume that three density coefficients 0.5, 0.3, and 0.2 are set, corresponding to the densities of 10, 20, and 30 nearest neighbors respectively. The densities of A at these three scales are calculated to be 1.397, 1.425, and 1.410 respectively. Then the multi-scale density of A is expressed as 0.5×1.397 + 0.3×1.425 + 0.2×1.410 = 1.408. Assume that the smoothing coefficient is 0.8, and the average multi-scale density of the 10 nearest neighbors of A is 1.420, then the final local density value of A is 0.8×1.408 + 0.2×1.420 = 1.4104.
[0049] In an alternative implementation, the edge computing node determines the computing priority of each target artificial intelligence model according to the computing complexity of each target artificial intelligence model and the computing resource status of the edge computing node, and sequentially calls the computing resources of the edge computing node to execute each target artificial intelligence model according to the computing priority, generating corresponding model computing results, including: Obtain the computing complexity by performing weighted summation on the product of the number of operands and the layer weight coefficient in each computing layer of the target artificial intelligence model based on the edge computing node; Construct a resource status vector including the processor occupancy rate, memory occupancy rate, and bandwidth occupancy rate, and determine the resource availability based on the minimum value of the ratio of the remaining amount to the maximum capacity of various resources according to the resource status vector; Multiply the normalized result of the computing complexity and the resource availability by the corresponding weight coefficients respectively to obtain the priority score; construct a multi-objective optimization function based on the priority score, and the multi-objective optimization function includes objective constraint conditions, resource constraint conditions, and latency constraint conditions; Determine the execution order of the target artificial intelligence models according to the solution result of the multi-objective optimization function, and sequentially call the computing resources of the edge computing node according to the execution order to execute each target artificial intelligence model, generating corresponding model computing results.
[0050] When the edge computing node evaluates the computational complexity of the target artificial intelligence model, it performs a weighted sum based on the product of the number of operations and the layer weight coefficient of each computational layer in the target artificial intelligence model. Specifically, for a convolutional neural network model, the model structure can be parsed into multiple computational layers, including convolutional layers, pooling layers, fully connected layers, etc. For each layer, count the number of its operations. For example, the number of operations in a convolutional layer can be calculated from parameters such as the size of the input feature map, the size of the convolutional kernel, and the number of convolutional kernels; the number of operations in a pooling layer can be calculated from parameters such as the size of the input feature map and the size of the pooling window; the number of operations in a fully connected layer can be calculated from the product of the number of input neurons and the number of output neurons.
[0051] For a target artificial intelligence model with 3 convolutional layers and 2 fully connected layers, the number of operations in each layer is 10 million, 8 million, 5 million, 2 million, and 1 million respectively, and the corresponding layer weight coefficients are 0.3, 0.25, 0.2, 0.15, and 0.1 respectively.
[0052] Then the computational complexity of this model is: 10 million × 0.3 + 8 million × 0.25 + 5 million × 0.2 + 2 million × 0.15 + 1 million × 0.1 = 3 million + 2 million + 1 million + 300,000 + 100,000 = 6.4 million.
[0053] The edge computing node constructs a resource status vector including the processor occupancy rate, memory occupancy rate, and bandwidth occupancy rate, and determines the resource availability based on the resource status vector according to the minimum value of the ratio of the remaining amount to the maximum capacity of various resources. Specifically, regularly collect the CPU usage rate, memory usage rate, and network bandwidth usage rate of the edge computing node to form a resource status vector. For example, the resource status vector at a certain moment is [70%, 60%, 50%], indicating that the CPU usage rate is 70%, the memory usage rate is 60%, and the bandwidth usage rate is 50%.
[0054] When calculating the resource availability based on the resource status vector, first calculate the ratio of the remaining amount to the maximum capacity of various resources, that is, [1 - 70%, 1 - 60%, 1 - 50%] = [30%, 40%, 50%], and then take the minimum value among them as the resource availability. In this example, it is 30%.
[0055] The edge computing node first normalizes the computational complexity and maps it to the interval [0, 1] to obtain the normalized computational complexity. For example, assume that the maximum computational complexity of all models in the system is 10 million and the minimum is 1 million. Then, the normalized value of a model with a computational complexity of 6.4 million is (6.4 - 1) / (10 - 1)=0.6. Then, the normalized computational complexity and the resource availability are multiplied by their corresponding weight coefficients respectively and summed to obtain the priority score. Let the weight coefficient of the normalized computational complexity be 0.6 and the weight coefficient of the resource availability be 0.4. For the target artificial intelligence model with a normalized computational complexity of 0.6, under the condition that the resource availability is 30%, its priority score is: 0.6×0.6 + 30%×0.4 = 0.36+0.12 = 0.48.
[0056] A multi-objective optimization function is constructed based on the priority score. The multi-objective optimization function includes objective constraint conditions, resource constraint conditions, and latency constraint conditions. The objective constraint conditions specify that all target artificial intelligence models that need to be executed must be scheduled; the resource constraint conditions specify that all scheduled target artificial intelligence models do not exceed the resource capacity limit of the edge computing node during execution; the latency constraint conditions specify that the execution time of all scheduled target artificial intelligence models does not exceed a preset latency threshold.
[0057] For the three target artificial intelligence models A, B, and C, their priority scores are 3.8412 million, 2.9635 million, and 1.5278 million respectively. The objective constraint conditions of the multi-objective optimization function are that all models must be executed; the resource constraint conditions are that the CPU usage rate does not exceed 95%, the memory usage rate does not exceed 90%, and the bandwidth usage rate does not exceed 85%; the latency constraint conditions are that the total model execution time does not exceed 200 milliseconds.
[0058] The execution order of the target artificial intelligence models is determined according to the solution result of the multi-objective optimization function, and the computing resources of the edge computing node are called in sequence according to the execution order to execute each target artificial intelligence model, generating the corresponding model calculation results.
[0059] The edge computing node can use the greedy algorithm or the dynamic programming algorithm to solve the multi-objective optimization function. Taking the greedy algorithm as an example, the models are arranged for execution in descending order of the priority score. For the above three models, the execution order is A→B→C.
[0060] The edge computing node monitors the resource usage in real time to ensure that the resource usage does not exceed the constraint conditions. For example, after model A is executed, the CPU usage rate rises to 85%, the memory usage rate rises to 75%, and the bandwidth usage rate rises to 65%; after model B is executed, the CPU usage rate rises to 92%, the memory usage rate rises to 82%, and the bandwidth usage rate rises to 75%; before model C is executed, it is judged whether the resource usage will exceed the constraint conditions after C is executed. If it will not exceed, model C is executed; if it will exceed, the execution of model C is postponed until the resources are released and then executed.
[0061] The edge computing node records the execution time of each model to ensure that the total execution time does not exceed the latency constraint conditions. For example, the execution time of model A is 80 milliseconds, the execution time of model B is 60 milliseconds, the expected execution time of model C is 50 milliseconds, and the total execution time is 190 milliseconds, which does not exceed the threshold of 200 milliseconds. Therefore, it can be executed in the planned order.
[0062] The edge computing node collects the calculation results of each model, such as the object detection result of model A, the image classification result of model B, and the speech recognition result of model C, and sends these results to the requester or proceeds to the next step of processing.
[0063] Figure 3 The following is the flow chart of artificial intelligence model scheduling in the edge computing environment of the embodiment of the present invention: This figure shows a complete model scheduling optimization process, which includes two parallel initial input branches: the left branch starts from the analysis of the computational complexity of the target artificial intelligence model, and obtains an accurate computational complexity evaluation value by calculating the product of the number of operations of each layer of the model and the weight coefficient of the corresponding layer and performing weighted summation; the right branch starts from the resource status of the edge computing node, and calculates the current resource availability by constructing a resource status vector that includes multiple dimensions such as processor utilization, memory occupancy, and network bandwidth. The results of these two branches converge in the middle of the process. The system obtains a comprehensive priority score by weighted combination of the normalized result of the computational complexity and the resource availability. Based on this score, the system further constructs a multi-objective optimization function, which simultaneously considers the target constraint conditions (such as model accuracy requirements), resource constraint conditions (such as computing resource limitations), and latency constraint conditions (such as execution time limitations). By solving this multi-objective optimization function, the execution order of the target artificial intelligence model is finally determined, and the computing resources are allocated accordingly to execute the corresponding model. The entire process shows a complete technical route from model complexity evaluation to resource status perception and then to optimized scheduling, ensuring the efficiency and reliability of model scheduling in the edge computing environment through multi-dimensional constraint conditions.
[0064] Existing methods for scheduling artificial intelligence models on edge computing nodes are mainly based on the principles of fixed priority or first-come-first-served, failing to fully consider the computational complexity of the models and the resource status of the nodes, resulting in low utilization efficiency of computing resources and inability to meet the requirements of multi-model parallel processing. For example, some existing technologies adopt a polling scheduling algorithm to execute each model in a fixed order, or a random scheduling algorithm to randomly select a model for execution. These methods cannot dynamically adjust the execution order according to the actual situation.
[0065] This application proposes a dynamic scheduling method based on computational complexity and resource status. By accurately evaluating the computational complexity of the model and real-time monitoring the resource status, a multi-objective optimization function is constructed to guide the scheduling decision. Compared with the existing technologies, the improvements of this application are as follows: First, a computational complexity evaluation method based on the number of operations and layer weights of each computational layer of the model is introduced, improving the accuracy of complexity evaluation; Second, a resource availability calculation method based on a multi-dimensional resource status vector is proposed, considering various resource factors such as processors, memory, and bandwidth; Third, an optimization function containing various constraint conditions is designed to achieve the optimal scheduling under the premise of meeting resource and latency constraints. These improvements enable edge computing nodes to utilize computing resources more efficiently, improve the processing efficiency of artificial intelligence models, and meet real-time requirements.
[0066] In an optional implementation manner, determining the execution order of the target artificial intelligence models according to the solution result of the multi-objective optimization function, and sequentially calling the computing resources of the edge computing nodes according to the execution order to execute each of the target artificial intelligence models, generating corresponding model calculation results includes: Solving the multi-objective optimization function to obtain the execution order, and the solution process of the multi-objective optimization function is restricted by a resource constraint condition that the resource usage does not exceed the available resource capacity, and a latency constraint condition that the model execution time does not exceed the maximum allowable latency; Sequentially call the computing resources of the edge computing nodes according to the execution order to execute each target artificial intelligence model, and generate corresponding model calculation results.
[0067] Obtain multiple target artificial intelligence models to be executed and their related parameter information, including the resource requirements and execution time of each model. Then, construct a multi-objective optimization function, and the objectives of this function include minimizing the total execution time and maximizing the resource utilization rate, etc.
[0068] Solve this multi-objective optimization function to determine the optimal execution order of the target artificial intelligence models. During the solution process, two main constraint conditions need to be considered: resource constraint and latency constraint. The resource constraint requires that the total amount of computing resources used by each model does not exceed the available resource capacity of the edge node; the latency constraint requires that the total execution time of the model does not exceed the preset maximum allowable latency.
[0069] Randomly generate a certain number of initial solutions, where each solution represents a possible model execution order. Then, new solutions are generated through crossover and mutation operations, and the roulette wheel selection method is used to select excellent individuals to enter the next generation. When evaluating the fitness of the solutions, two objectives, namely the total execution time and resource utilization rate, are considered simultaneously, and a penalty factor is introduced to handle the solutions that do not meet the constraint conditions. After multiple generations of iteration, the Pareto optimal solution set is finally obtained, and a compromise solution is selected from it as the final model execution order.
[0070] Suppose there are three artificial intelligence models A, B, and C to be executed, with their resource requirements being 2, 3, and 4 units respectively, and the execution times being 10, 15, and 20 ms respectively. The available resources of the edge node are 5 units, and the maximum allowable time delay is 40 ms. By solving the multi-objective optimization function, one possible optimal execution order is B - A - C.
[0071] Send an execution request to the edge node, including the relevant information of model B. After receiving the request, the edge node allocates 3 units of computing resources to execute model B. After model B finishes execution, it releases the occupied resources and returns the calculation result to the system.
[0072] The edge node allocates 2 units of resources to execute model A, and after completion, it also returns the result and releases the resources.
[0073] Since model C requires 4 units of resources, which exceeds the current available resource amount, the system needs to wait for the previous models to finish execution and release resources before it can execute. When the resources are sufficient, the edge node allocates 4 units of resources to execute model C and returns the result after completion.
[0074] Monitor the actual execution time and resource usage of each model in real time. If it is found that the execution time of a certain model significantly deviates from the expectation, or the resource usage amount exceeds the preset value, the system will timely adjust the execution plan of the subsequent models. For example, if the actual execution time of model B reaches 18 ms, exceeding the expected 15 ms, the system may decide to skip executing model C to ensure that the total execution time does not exceed the maximum allowable time delay of 40 ms.
[0075] To further optimize the execution efficiency, the system can also execute the models in a pipeline manner. In the above example, when model B has been executed to a certain extent, if there are sufficient remaining resources, the system can start the execution of model A simultaneously, thereby reducing the overall execution time.
[0076] After all target artificial intelligence models have been executed, the system integrates and analyzes the calculation results of each model. For example, if these three models are respectively used for image recognition, speech processing, and natural language understanding, the system can combine their output results to generate a multi-modal intelligent analysis report.
[0077] Evaluate and summarize the execution process. Evaluation metrics include the actual total execution time, resource utilization rate, execution effectiveness of each model, etc. This information will be used to optimize future model execution strategies, such as adjusting the model priorities, updating resource requirement estimates, etc.
[0078] Efficiently execute multiple artificial intelligence models in an edge computing environment, while meeting latency and resource constraints, and achieving optimization of execution time and resource utilization. This method has strong adaptability and scalability, and can dynamically adjust the execution strategy according to the actual situation, and is applicable to various complex edge computing scenarios.
[0079] In an alternative embodiment, the edge computing node aggregates the model calculation results of each of the target artificial intelligence models, and the processing results for the original data obtained include: Construct a credibility evaluation function for the model results, where the credibility evaluation function is obtained based on a weighted combination of the historical accuracy of each target artificial intelligence model, the model execution quality under the current resource state, and the result output time, and apply the credibility evaluation function to the model calculation results of each of the target artificial intelligence models; Perform adaptive weighting on the model calculation results of each of the models based on the credibility evaluation function, where the weight coefficients of the adaptive weighting are dynamically adjusted according to the change of the credibility evaluation function, and the weight coefficients are normalized to ensure that the sum of the weights of the model calculation results of each of the models is one; Use a sliding window mechanism based on temporal correlation for the weighted model calculation results of each of the models, and combine the weighted results within the sliding window to obtain the processing results for the original data.
[0080] The edge computing node receives the calculation results of multiple target artificial intelligence models for the original data. These models can be of different types, such as image recognition models, speech recognition models, etc., which respectively process the original data and output results.
[0081] Build a credibility evaluation function for the model results. This function comprehensively considers three key factors: the historical accuracy of each target artificial intelligence model, the model execution quality under the current resource state, and the result output time. Specifically, the historical accuracy can be obtained by statistically calculating the matching degree between the prediction results and the actual results of each model over a long period; the model execution quality can be evaluated by monitoring the usage of current computing resources (such as CPU utilization, memory occupancy, etc.); the result output time directly measures the time required for the model to generate results. These three factors are combined through weighted combination to obtain the final credibility evaluation value.
[0082] Suppose there are three models A, B, and C, with historical accuracies of 0.9, 0.85, and 0.8 respectively, execution quality scores of 0.95, 0.9, and 0.85 under the current resource state, and result output times (after normalization) of 0.8, 0.9, and 0.7. We can assign weights to these three factors, such as 0.5, 0.3, and 0.2. Then the credibility evaluation value of model A is: 0.9×0.5 + 0.95×0.3 + 0.8×0.2 = 0.895. Similarly, the credibility evaluation values of B and C can be calculated.
[0083] Apply the credibility evaluation function to the calculation results of each target artificial intelligence model. This step will assign a credibility score to the result of each model, reflecting the reliability of the result.
[0084] Based on the results of the credibility evaluation function, perform adaptive weighting on the calculation results of each model. The key here is that the weight coefficient will be dynamically adjusted according to the changes in the credibility evaluation function. When specifically implemented, the credibility evaluation value of each model can be divided by the sum of the credibility evaluation values of all models to obtain the weight of the result of this model. This method naturally ensures that the sum of all weights is 1.
[0085] Suppose the credibility evaluation values of models A, B, and C are 0.895, 0.875, and 0.785 respectively. Then their weights can be calculated as follows: A: 0.895 / (0.895 + 0.875 + 0.785) ≈ 0.35; B: 0.875 / (0.895 + 0.875 + 0.785) ≈ 0.34; C: 0.785 / (0.895 + 0.875 + 0.785) ≈ 0.31; This adaptive weighting mechanism can dynamically adjust the influence of the model in the final result according to the real-time performance of the model, improving the flexibility and robustness of the system.
[0086] Adopt a sliding window mechanism based on temporal correlation to combine the weighted model calculation results to obtain the final processing result for the original data. The sliding window mechanism takes into account the temporal continuity of the data and is particularly suitable for processing time series data.
[0087] A fixed-size time window can be set, such as 5 minutes. Within this window, collect the weighted results of all models. As time progresses, the window slides forward continuously, always maintaining the latest 5-minute data. For the data within the window, various combination methods can be adopted, such as simple average, weighted average, median, etc. The specific choice depends on the requirements of the application scenario.
[0088] Suppose within a certain 5-minute window, three models made multiple predictions for the same target, and the weighted results are as follows: A: [0.7, 0.72, 0.68, 0.71, 0.69]; B: [0.65, 0.67, 0.66, 0.68, 0.64]; C: [0.62, 0.63, 0.61, 0.64, 0.62]; Using the simple average method, the average prediction value of each model within this window can be obtained: A: 0.7, B: 0.66, C: 0.624; Considering the previously calculated model weights (A: 0.35, B: 0.34, C: 0.31), the final weighted combination result can be obtained: 0.7×0.35 + 0.66×0.34 + 0.624×0.31 ≈ 0.663; This value is the final result of the system's processing of the original data within this time window.
[0089] It should be noted that the size of the sliding window can be adjusted according to the actual application requirements. A larger window can provide more stable results, but may reduce the system's response speed to sudden events; a smaller window can reflect data changes faster, but may introduce more noise. In practical applications, the most suitable window size can be determined through experiments.
[0090] An update frequency can also be set, such as updating the results once per second or per minute. Each time an update occurs, the credibility of the model is recalculated, the weights are adjusted, and new processing results are generated based on the latest sliding window data. This can ensure that the system always uses the latest and most relevant information to generate results.
[0091] The edge computing node can effectively aggregate and process the calculation results of multiple target artificial intelligence models to obtain high-quality processing results for the original data. This method not only considers the historical performance and current state of each model, but also captures the temporal characteristics of the data through a sliding window mechanism, thus providing a flexible and reliable data processing solution.
[0092] Figure 4 The flowchart of multi-model result fusion based on credibility evaluation in the embodiment of the present invention is as follows: This figure details a complete fusion process from the model calculation results to the final processing results. The process starts with the input of the calculation results of each target artificial intelligence model, and then a credibility evaluation function based on three key dimensions is constructed: historical accuracy (reflecting the long-term performance of the model), model execution quality (reflecting the processing ability under the current resource state), and result output time (characterizing the processing efficiency). These three dimensions form a credibility evaluation function through weighted combination. On this basis, the system adaptively weights the model calculation results, where the weight coefficients will be dynamically adjusted according to the credibility evaluation results, and through normalization, it is ensured that the sum of all weights is one, guaranteeing the scientificity and rationality of the weighting process. Finally, the system adopts a sliding window mechanism based on temporal correlation to combine the weighted results within the sliding window, and finally generates the processing results for the original data. This fusion scheme not only considers the historical performance of the model, but also takes into account the current execution state, and at the same time ensures the timeliness and continuity of the results through the consideration of temporal correlation, forming a complete multi-model result fusion technical solution.
[0093] In the second aspect of the embodiment of the present invention, An artificial intelligence model service system based on edge computing is provided, including: The first unit is used to receive an artificial intelligence model service request sent by a mobile terminal, and the artificial intelligence model service request includes the original data to be processed; The second unit is used to perform data feature analysis on the original data based on an edge computing node, divide the original data into multiple data subsets according to the data characteristics of the original data, and select corresponding target artificial intelligence models from a preset artificial intelligence model library based on each data subset, so that each data subset is processed by a target artificial intelligence model matching its data characteristics; The third unit is used for the edge computing node to determine the calculation priority of each target artificial intelligence model according to the calculation complexity of each target artificial intelligence model and the calculation resource state of the edge computing node, and sequentially call the calculation resources of the edge computing node to execute each target artificial intelligence model according to the calculation priority, and generate corresponding model calculation results; The fourth unit is used for the edge computing node to summarize the model calculation results of each of the target artificial intelligence models, obtain a processing result for the original data, and return the processing result to the mobile terminal.
[0094] In the third aspect of the embodiments of the present invention, A kind of electronic device is provided, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0095] In the fourth aspect of the embodiments of the present invention, A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0096] The present invention can be a method, device, system and / or computer program product. The computer program product can include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.
[0097] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An artificial intelligence model service method based on edge computing, characterized in that: include: Receiving an artificial intelligence model service request sent by a mobile terminal, wherein the artificial intelligence model service request includes raw data to be processed; Performing data feature analysis on the original data based on the edge computing node, dividing the original data into multiple data subsets according to the data features of the original data, and selecting a corresponding target artificial intelligence model from a preset artificial intelligence model library based on each of the data subsets, so that each of the data subsets is processed using a target artificial intelligence model that matches its data features; The edge computing node determines the computing priority of each target artificial intelligence model according to the computing complexity of each target artificial intelligence model and the computing resource status of the edge computing node, and sequentially calls the computing resources of the edge computing node to execute each target artificial intelligence model according to the computing priority to generate a corresponding model calculation result; The edge computing node aggregates and processes the model calculation results of each of the target artificial intelligence models to obtain processing results for the original data, and returns the processing results to the mobile terminal.
2. The method according to claim 1, characterized in that Dividing the original data into a plurality of data subsets according to the data features of the original data, and selecting a corresponding target artificial intelligence model from a preset artificial intelligence model library based on each of the data subsets, so that each of the data subsets is processed using a target artificial intelligence model that matches its data features, includes: Extracting time series features, statistical features and information entropy features from the original data to construct feature vectors; calculating the distance between a data point of the original data and other data points according to the feature vectors, and obtaining a local density value according to the relationship between the distance and a preset cutoff distance; Clustering the original data based on the local density value to obtain multiple data subsets; A matching score is formed based on a weighted sum of feature vectors of a plurality of said data subsets and preset model feature descriptors; a priority queue is constructed by combining said data subsets with corresponding model feature descriptors and said matching scores; The optimal matching relationship between each of the data subsets and the model is determined by minimizing the allocation cost, and each of the data subsets is matched with the model that matches its data features based on the optimal matching relationship.
3. The method according to claim 2, characterized in that By calculating the distance between the data point of the original data and other data points according to the feature vector, obtaining the local density value according to the relationship between the distance and the preset cutoff distance includes: Calculating the distance between data points in the original data, wherein the distance is obtained by weighted summation of a dedicated distance function for each feature dimension; Calculating the structural similarity between the data point and its neighboring data points, multiplying the structural similarity by the adjustment coefficient and adding 1 to obtain a structural perception density value; Calculating an arithmetic mean of the structure-perceived density value by arithmetic averaging based on the adjustment coefficient, and dividing the arithmetic mean by the structure-perceived density value to obtain a local abnormality factor; The difference between the local abnormality factor and the preset abnormality threshold is multiplied by the attenuation coefficient, the multiplication result is negated and then subjected to exponential operation to obtain the attenuation factor, and the attenuation factor is multiplied by the structural perception density value to obtain the corrected density value; The modified density value and the preset density coefficient are weightedly combined to obtain a multi-scale density representation, the product of the smoothing coefficient and the multi-scale density representation, and the product of the smoothing coefficient and the weighted multi-scale density value of the adjacent data points are calculated, and the two products are added to obtain a local density value.
4. The method according to claim 1, characterized in that The edge computing node determines the computing priority of each target artificial intelligence model according to the computing complexity of each target artificial intelligence model and the computing resource status of the edge computing node, and sequentially calls the computing resources of the edge computing node to execute each target artificial intelligence model according to the computing priority, and generates corresponding model calculation results including: Based on the edge computing node, a weighted sum is performed on the product of the number of operations of each computing layer in the target artificial intelligence model and the layer weight coefficient to obtain the computational complexity; Constructing a resource state vector including processor occupancy, memory occupancy and bandwidth occupancy, and determining resource availability based on the minimum value of the ratio of the remaining amount to the maximum capacity of each type of resource according to the resource state vector; The result of normalizing the computational complexity and the resource availability are respectively multiplied by corresponding weight coefficients to obtain a priority score; a multi-objective optimization function is constructed based on the priority score, wherein the multi-objective optimization function includes an objective constraint condition, a resource constraint condition, and a delay constraint condition; The execution order of the target artificial intelligence models is determined according to the solution results of the multi-objective optimization function, and the computing resources of the edge computing node are called to execute each of the target artificial intelligence models in sequence according to the execution order to generate corresponding model calculation results.
5. The method according to claim 4, characterized in that Determining the execution order of the target artificial intelligence models according to the solution results of the multi-objective optimization function, calling the computing resources of the edge computing node to execute each of the target artificial intelligence models in sequence according to the execution order, and generating corresponding model calculation results includes: Solving the multi-objective optimization function to obtain an execution order, wherein the solving process of the multi-objective optimization function is subject to a resource constraint condition that resource usage does not exceed available resource capacity, and a delay constraint condition that model execution time does not exceed a maximum allowable delay; The computing resources of the edge computing nodes are called in sequence in the execution order to execute each target artificial intelligence model and generate corresponding model calculation results.
6. The method according to claim 1, characterized in that The edge computing node aggregates the model calculation results of each of the target artificial intelligence models to obtain processing results for the original data, including: Construct a credibility evaluation function for the model results, the credibility evaluation function is obtained based on a weighted combination of the historical accuracy of each target artificial intelligence model, the model execution quality under the current resource state, and the result output time, and apply the credibility evaluation function to the model calculation results of each of the target artificial intelligence models; Based on the credibility evaluation function, each of the model calculation results is adaptively weighted, and the weight coefficient of the adaptive weighting is dynamically adjusted as the credibility evaluation function changes, and the weight coefficient is normalized to ensure that the sum of the weights of the model calculation results is one; The weighted calculation results of each model are processed using a sliding window mechanism based on time series correlation, and the weighted results in the sliding window are combined to obtain the processing results for the original data.
7. An artificial intelligence model service system based on edge computing, used to implement the method described in any one of claims 1 to 6, characterized in that: include: A first unit is configured to receive an artificial intelligence model service request sent by a mobile terminal, wherein the artificial intelligence model service request includes raw data to be processed; A second unit is used to perform data feature analysis on the original data based on the edge computing node, divide the original data into multiple data subsets according to the data features of the original data, and select a corresponding target artificial intelligence model from a preset artificial intelligence model library based on each data subset, so that each data subset is processed by a target artificial intelligence model that matches its data features; The third unit is used for the edge computing node to determine the computing priority of each target artificial intelligence model according to the computing complexity of each target artificial intelligence model and the computing resource status of the edge computing node, and to call the computing resources of the edge computing node in turn according to the computing priority to execute each target artificial intelligence model to generate a corresponding model calculation result; The fourth unit is used for the edge computing node to aggregate and process the model calculation results of each of the target artificial intelligence models, obtain the processing results for the original data, and return the processing results to the mobile terminal.
8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Electric energy meter with monitoring data analysis function
CN118378110A
Station area intelligent fusion terminal data processing system based on edge calculation
CN119440800A
Cloud data processing system based on artificial intelligence algorithm
CN119473645A
Cited By
Data processing method and related equipment
CN120434293A
Industrial equipment group intelligent operation and maintenance and energy consumption optimization method based on edge computing
CN121742380A
Industrial device group intelligent operation and maintenance and energy consumption optimization method based on edge computing
CN121742380B