Artificial Intelligence Model Service Method and System Based on Edge Computing
Through edge computing nodes, the data is characterized and matched with model processing, which solves the latency and privacy issues in cloud computing, and realizes efficient and secure personalized artificial intelligence services.
Patent Information
- Application Number
- CN202510630339.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The artificial intelligence model service method of traditional cloud computing has problems such as high network latency, difficulty in meeting real-time requirements and privacy security. How to effectively use edge computing resources to process different types of data is still a key issue.
The original data is characterized by edge computing nodes, divided into multiple data subsets, and the target artificial intelligence model matching the data characteristics is selected for processing, and priority is determined based on the calculation complexity and resource status, computing resources are reasonably allocated, and processing results are generated and summarized.
It improves the accuracy and efficiency of data processing, reduces network latency, protects user privacy, enhances data security, and provides mobile terminals with efficient, secure and personalized artificial intelligence services.
Smart Images

Figure CN120144327B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to artificial intelligence technology, and particularly to an artificial intelligence model service method and system based on edge computing. Background Art
[0002] With the rapid development of artificial intelligence technology, more and more mobile applications have started to integrate artificial intelligence functions to provide more intelligent and personalized services. However, due to the limited computing resources and storage space of mobile terminals, it is difficult to run complex artificial intelligence models locally. Therefore, usually, the cloud computing method is adopted to deploy the artificial intelligence model on a remote server, and the mobile terminal obtains the model service through a network request.
[0003] However, the traditional artificial intelligence model service method based on cloud computing has some problems. First, transmitting a large amount of raw data to a remote server for processing will increase network latency and affect the user experience. Second, the centralized cloud computing mode is difficult to meet the requirements of mobile applications for real-time performance and low latency, especially in the case of poor network conditions. Finally, transmitting all data to the cloud for processing may cause privacy and security problems, and the sensitive information of users may be leaked or misused.
[0004] To solve these problems, edge computing, as an emerging computing paradigm, has been introduced into artificial intelligence model services. Edge computing sinks computing resources and storage resources to the network edge, close to data sources and users, and can quickly process data locally, reduce network transmission latency, and improve service response speed. However, how to effectively utilize edge computing resources, select appropriate artificial intelligence models for different types of data, and how to efficiently process a large amount of data under limited edge computing resources are still key problems to be solved. Summary of the Invention
[0005] Embodiments of the present invention provide an artificial intelligence model service method and system based on edge computing, which can solve the problems in the prior art.
[0006] In the first aspect of the embodiments of the present invention,
[0007] An artificial intelligence model service method based on edge computing is provided, including:
[0008] Receiving an artificial intelligence model service request sent by a mobile terminal, where the artificial intelligence model service request includes raw data to be processed;
[0009] Performing data feature analysis on the original data based on an edge computing node, dividing the original data into multiple data subsets according to the data features of the original data, and selecting a corresponding target artificial intelligence model from a preset artificial intelligence model library for each data subset, so that each data subset is processed by a target artificial intelligence model matching its data features;
[0010] The edge computing node determines the calculation priority of each target artificial intelligence model according to the calculation complexity of each target artificial intelligence model and the calculation resource status of the edge computing node, and sequentially calls the calculation resources of the edge computing node to execute each target artificial intelligence model according to the calculation priority to generate corresponding model calculation results;
[0011] The edge computing node aggregates and processes the model calculation results of each target artificial intelligence model to obtain a processing result for the original data, and returns the processing result to the mobile terminal.
[0012] Dividing the original data into multiple data subsets according to the data features of the original data, and selecting a corresponding target artificial intelligence model from a preset artificial intelligence model library for each data subset, so that each data subset is processed by a target artificial intelligence model matching its data features includes:
[0013] Extracting the time series features, statistical features, and information entropy features in the original data to construct a feature vector; obtaining a local density value according to the relationship between the distance calculated between the data points of the original data and other data points through the feature vector and a preset truncation distance;
[0014] Performing clustering division on the original data based on the local density value to obtain multiple data subsets;
[0015] Constructing a matching degree score based on the weighted sum of the feature vectors of multiple data subsets and a preset model feature descriptor; constructing a priority queue for the data subsets, the corresponding model feature descriptors, and the matching degree score;
[0016] Determining the optimal matching relationship between each data subset and the model by minimizing the allocation cost, and matching each data subset with a model matching its data features based on the optimal matching relationship.
[0017] Obtaining a local density value according to the relationship between the distance calculated between the data points of the original data and other data points through the feature vector and a preset truncation distance includes:
[0018] Calculate the distance between data points in the original data, where the distance is obtained by weighted summation of dedicated distance functions for each feature dimension;
[0019] Calculate the structural similarity degree between the data point and its neighboring data points, multiply the structural similarity degree by an adjustment coefficient and then add one to obtain a structure-aware density value;
[0020] Based on the adjustment coefficient, calculate the arithmetic mean of the structure-aware density values by arithmetic mean, divide the arithmetic mean by the structure-aware density value to obtain a local outlier factor;
[0021] Multiply the difference between the local outlier factor and a preset outlier threshold by an attenuation coefficient, take the negative of the multiplication result and perform an exponential operation to obtain an attenuation factor, multiply the attenuation factor by the structure-aware density value to obtain a corrected density value;
[0022] Perform a weighted combination of the corrected density value and a preset density coefficient to obtain a multi-scale density representation, calculate the product of a smoothing coefficient and the multi-scale density representation, and the product of the smoothing coefficient and the weighted multi-scale density value of neighboring data points, and add the two products to obtain a local density value.
[0023] The edge computing node determines the computing priority of each target artificial intelligence model according to the computing complexity of each target artificial intelligence model and the computing resource status of the edge computing node, and sequentially calls the computing resources of the edge computing node to execute each target artificial intelligence model according to the computing priority, and the generated corresponding model computing results include:
[0024] Based on the edge computing node, perform a weighted summation of the products of the operands and layer weight coefficients in each computing layer of the target artificial intelligence model to obtain the computing complexity;
[0025] Construct a resource status vector including processor occupancy rate, memory occupancy rate and bandwidth occupancy rate, and determine the resource availability based on the resource status vector according to the minimum value of the ratio of the remaining amount to the maximum capacity of various resources;
[0026] Multiply the normalized result of the computing complexity and the resource availability by corresponding weight coefficients respectively to obtain a priority score; construct a multi-objective optimization function based on the priority score, and the multi-objective optimization function includes objective constraint conditions, resource constraint conditions and delay constraint conditions;
[0027] Determine the execution order of the target artificial intelligence models according to the solution result of the multi-objective optimization function, and sequentially call the computing resources of the edge computing node to execute each target artificial intelligence model according to the execution order, and generate corresponding model computing results, including:
[0028] Solve the multi-objective optimization function to obtain the execution order. The solution process of the multi-objective optimization function is restricted by the resource constraint condition that the resource usage does not exceed the available resource capacity, and the latency constraint condition that the model execution time does not exceed the maximum allowable latency;
[0029] Call the computing resources of the edge computing nodes in sequence according to the execution order to execute each target artificial intelligence model, and generate corresponding model calculation results.
[0030] The edge computing node aggregates and processes the model calculation results of each target artificial intelligence model to obtain the processing results for the original data, including:
[0031] Construct a credibility evaluation function for the model results. The credibility evaluation function is obtained based on the weighted combination of the historical accuracy of each target artificial intelligence model, the model execution quality under the current resource state, and the result output time. Apply the credibility evaluation function to the model calculation results of each target artificial intelligence model;
[0032] Perform adaptive weighting on the model calculation results based on the credibility evaluation function. The weight coefficient of the adaptive weighting is dynamically adjusted with the change of the credibility evaluation function. The weight coefficient is normalized to ensure that the sum of the weights of the model calculation results is one;
[0033] Use a sliding window mechanism based on temporal correlation for the weighted model calculation results, and combine the weighted results within the sliding window to obtain the processing results for the original data.
[0034] In the second aspect of the embodiments of the present invention,
[0035] Provide an artificial intelligence model service system based on edge computing, including:
[0036] The first unit is used to receive an artificial intelligence model service request sent by a mobile terminal. The artificial intelligence model service request includes the original data to be processed;
[0037] The second unit is used to perform data feature analysis on the original data based on an edge computing node, divide the original data into multiple data subsets according to the data features of the original data, and select corresponding target artificial intelligence models from a preset artificial intelligence model library based on each data subset, so that each data subset is processed by a target artificial intelligence model that matches its data features;
[0038] A third unit, configured to determine, according to the computing complexity of each target artificial intelligence model and the computing resource status of the edge computing node, the computing priority of each target artificial intelligence model, and sequentially call the computing resources of the edge computing node to execute each target artificial intelligence model, so as to generate corresponding model computing results;
[0039] A fourth unit, configured to aggregate and process the model computing results of each target artificial intelligence model to obtain a processing result for the original data, and return the processing result to the mobile terminal.
[0040] In a third aspect of the embodiments of the present invention,
[0041] There is provided an electronic device, including:
[0042] A processor;
[0043] A memory for storing instructions executable by the processor;
[0044] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0045] In a fourth aspect of the embodiments of the present invention,
[0046] There is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0047] The beneficial effects of this application are as follows:
[0048] 1. By performing feature analysis on the original data and dividing it into multiple data subsets, and selecting a matching target artificial intelligence model for each subset for processing, the accuracy and efficiency of data processing can be improved, and the advantages of different models can be fully utilized to process data with different features.
[0049] 2. By determining the computing priority according to the computing complexity of the target artificial intelligence model and the resource status of the edge computing node, reasonable allocation and scheduling of computing resources can be achieved, the resource utilization rate of the edge computing node can be improved, and resource waste or over-occupation can be avoided.
[0050] 3. Completing data processing on the edge computing node and returning the results to the mobile terminal reduces the amount of data transmission and network latency, improves the response speed, and at the same time protects user privacy and enhances data security. This method gives full play to the advantages of edge computing and provides efficient, secure, and personalized artificial intelligence services for mobile terminals. Description of the Drawings
[0051] Figure 1 Flow chart of the artificial intelligence model service method based on edge computing according to an embodiment of the present invention;
[0052] Figure 2 Flow chart for determining the optimal matching relationship based on the original data according to an embodiment of the present invention;
[0053] Figure 3 Flow chart of artificial intelligence model scheduling in the edge computing environment according to an embodiment of the present invention;
[0054] Figure 4 Flow chart of multi-model result fusion based on credibility evaluation according to an embodiment of the present invention. Detailed implementation manners
[0055] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only some of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0056] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments may be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0057] Figure 1 Flow chart of the artificial intelligence model service method based on edge computing according to an embodiment of the present invention, as Figure 1 shown, the method includes:
[0058] Receiving an artificial intelligence model service request sent by a mobile terminal, where the artificial intelligence model service request includes original data to be processed;
[0059] Performing data feature analysis on the original data based on an edge computing node, dividing the original data into multiple data subsets according to the data features of the original data, and selecting a corresponding target artificial intelligence model from a preset artificial intelligence model library for each data subset, so that each data subset is processed by a target artificial intelligence model matching its data features;
[0060] The edge computing node determines the calculation priority of each target artificial intelligence model according to the calculation complexity of each target artificial intelligence model and the calculation resource status of the edge computing node, and sequentially calls the calculation resources of the edge computing node to execute each target artificial intelligence model according to the calculation priority, generating corresponding model calculation results;
[0061] The edge computing nodes aggregate the model calculation results of each of the target artificial intelligence models to obtain a processing result for the original data, and return the processing result to the mobile terminal.
[0062] In an alternative embodiment, dividing the original data into a plurality of data subsets according to the data characteristics of the original data, and selecting a corresponding target artificial intelligence model from a preset artificial intelligence model library based on each data subset, such that each data subset is processed by a target artificial intelligence model matching its data characteristics includes:
[0063] Extracting the time series features, statistical features, and information entropy features in the original data to construct a feature vector; calculating the distance between the data points of the original data and other data points according to the feature vector, and obtaining a local density value according to the relationship between the distance and a preset truncation distance;
[0064] Performing clustering division on the original data based on the local density value to obtain a plurality of data subsets;
[0065] Constructing a matching degree score based on the weighted sum of the feature vectors of the plurality of data subsets and a preset model feature descriptor; constructing a priority queue with the data subsets, the corresponding model feature descriptors, and the matching degree score;
[0066] Determining the optimal matching relationship between each data subset and a model by minimizing the allocation cost, and matching each data subset with a model matching its data characteristics based on the optimal matching relationship.
[0067] Dividing the original data into a plurality of data subsets according to the data characteristics of the original data, and selecting a corresponding target artificial intelligence model from a preset artificial intelligence model library based on each data subset, such that each data subset is processed by a target artificial intelligence model matching its data characteristics.
[0068] Extracting the time series features, statistical features, and information entropy features in the original data to construct a feature vector. The time series features include the periodicity, trend, seasonality, etc. of the data, and can be extracted by methods such as autocorrelation function, Fourier transform, etc.; the statistical features include the features describing the data distribution such as mean, variance, skewness, kurtosis, etc.; the information entropy features are used to measure the uncertainty and complexity of the data, and are obtained by calculating the Shannon entropy.
[0069] For a set of 1000 original temperature sensor records, extracting their temporal features yields an obvious daily variation pattern with a period of 24 hours and an autocorrelation coefficient of 0.82. The statistical features include a mean of 23.5°C, a standard deviation of 3.2°C, a skewness of -0.18, and a kurtosis of 2.76. The information entropy feature is 4.36. These features are combined into a feature vector [0.82, 24, 23.5, 3.2, -0.18, 2.76, 4.36] as a comprehensive description characterizing the properties of this dataset.
[0070] Based on the feature vector, calculate the distances between the data points of the original data and other data points, and obtain the local density values according to the relationship between the distances and a preset truncation distance. Specifically, select the Euclidean distance, Manhattan distance, Mahalanobis distance, etc. as the metric for the distances between data points. Set a truncation distance threshold, such as the 10% quantile of the distances for all data point pairs. For each data point, calculate its distances to all other data points, and count the number of points with distances less than the truncation distance, which is the local density value of this point.
[0071] In the above temperature sensor data case, assume that the Euclidean distance is selected as the metric, and the calculated truncation distance is 1.5. For data point 1, 78 of its distances to other points are less than 1.5, so its local density value is 78; for data point 2, 45 of its distances to other points are less than 1.5, so its local density value is 45, and so on to calculate the local density values of all data points.
[0072] Based on the local density values, perform clustering on the original data to obtain multiple data subsets. Use the density peak clustering algorithm. First, identify the density peak points (points with relatively high local density values and far distances from other high-density points) as the clustering centers, and then assign the remaining points to the nearest clustering center. To avoid the influence of outliers, density thresholds and distance thresholds can be set to screen the clustering centers.
[0073] In the temperature sensor data case, 3 clustering centers are identified through the density peak clustering algorithm, corresponding to the data points with density values of 92, 83, and 76 respectively. After assigning the remaining data points to the nearest clustering center, 3 data subsets are obtained, with sizes of 420, 350, and 230 records respectively. Subset 1 mainly contains daytime temperature data, subset 2 mainly contains nighttime temperature data, and subset 3 mainly contains temperature transition period data.
[0074] The matching degree score is composed of the weighted sum of the feature vectors of multiple data subsets and the preset model feature descriptors. The preset model feature descriptors describe the characteristics of the data applicable to the artificial intelligence model, including quantization values in dimensions such as time series, stationarity, periodicity, and volatility. The dot product operation is performed on the feature vectors of the data subsets and the model feature descriptors, and a weight coefficient is introduced to reflect the importance of each feature dimension to obtain the matching degree score.
[0075] There are three models in the artificial intelligence model library: The LSTM model is suitable for processing strongly time-series data, and its model feature descriptor is [0.9, 0.5, 0.7, 0.3]; The ARIMA model is suitable for processing stationary time-series data, and its model feature descriptor is [0.6, 0.8, 0.4, 0.2]; The Prophet model is suitable for processing seasonal data, and its model feature descriptor is [0.5, 0.6, 0.9, 0.4]. The feature weights are set as [0.4, 0.3, 0.2, 0.1].
[0076] For data subset 1, its feature vector is simplified to [0.85, 0.6, 0.75, 0.3]. The calculated matching degree score with the LSTM model is:
[0077] 0.85×0.9×0.4 + 0.6×0.5×0.3 + 0.75×0.7×0.2 + 0.3×0.3×0.1 = 0.306 + 0.09 + 0.105 + 0.009 = 0.51;
[0078] The matching degree score with the ARIMA model is:
[0079] 0.85×0.6×0.4 + 0.6×0.8×0.3 + 0.75×0.4×0.2 + 0.3×0.2×0.1 = 0.204 + 0.144 + 0.06 + 0.006 = 0.414;
[0080] The matching degree score with the Prophet model is:
[0081] 0.85×0.5×0.4 + 0.6×0.6×0.3 + 0.75×0.9×0.2 + 0.3×0.4×0.1 = 0.17 + 0.108 + 0.135 + 0.012 = 0.425.
[0082] Construct a priority queue with the data subset, the corresponding model feature descriptors, and the matching degree scores. The priority queue is sorted in descending order of the matching degree scores, and the queue elements include the data subset identifier, the model identifier, and the matching degree score.
[0083] Determine the optimal matching relationship between each data subset and the model by minimizing the allocation cost, and match each data subset with the model that matches its data characteristics based on the optimal matching relationship. The allocation cost is defined as 1 minus the matching degree score. Establish a bipartite graph matching problem and solve it using the Hungarian algorithm to obtain the overall optimal matching scheme of data subsets and models.
[0084] The optimal matching scheme is obtained by solving with the Hungarian algorithm: Data subset 1 is matched with the LSTM model (cost 0.49), data subset 2 is matched with the ARIMA model (cost 0.52), data subset 3 is matched with the Prophet model (cost 0.47), and the total cost is 1.48, which is the minimum cost among all possible matching schemes.
[0085] According to the optimal matching scheme, 420 daytime temperature data of data subset 1 are handed over to the LSTM model for processing, 350 nighttime temperature data of data subset 2 are handed over to the ARIMA model for processing, and 230 temperature transition period data of data subset 3 are handed over to the Prophet model for processing. This matching relationship makes full use of the strengths of different models. The LSTM model is good at capturing the complex non-linear relationships in daytime temperature data, the ARIMA model is suitable for dealing with relatively stable nighttime temperature changes, and the Prophet model is good at dealing with data with obvious transition characteristics.
[0086] Figure 2 Flowchart for determining the optimal matching relationship based on the original data in the embodiments of the present invention:
[0087] This figure shows a complete flowchart of data processing and model matching. Starting from the input of the original data at the top, it goes through multiple key processing steps and finally realizes the optimal matching of data and models. The specific process includes: First, extract features from the original data, including time series features, statistical features, and information entropy features, and construct feature vectors; then calculate the distance relationship and local density values between data points to evaluate the distribution characteristics of the data through these metrics; perform data clustering based on the obtained local density values to divide similar data into the same subset; then calculate the matching degree between the feature vectors of each data subset and the model features in the artificial intelligence model library, and evaluate it using the matching score composed of weighted sums; construct a priority queue between the data subset and the model according to the matching score, which reflects the matching priority relationship between the data and the model; finally, determine the optimal matching relationship between the data subset and the model by minimizing the allocation cost. The entire process reflects a systematic processing process from data feature extraction to model matching. Through local density calculation and the construction of a priority queue, the intelligent matching of data and models is realized. The main innovation point of this process is the organic combination of data feature extraction, density analysis, and model matching, forming a complete intelligent processing framework.
[0088] Existing data processing methods usually use a unified artificial intelligence model to process the entire dataset, or select a model based on simple rules (such as the size of the data volume, data type), without fully considering the matching relationship between the internal characteristics of the data and the applicable scenarios of the model. For example, some methods only select a model based on the source type of the data, uniformly use a time series model for temperature data, ignoring that temperature data at different times may have different characteristics; there are also methods that only select the model complexity based on the size of the data volume, select a deep learning model when the data volume is large, and select a statistical learning model when the data volume is small. This rough division method is difficult to achieve the optimal match.
[0089] This application proposes a fine-grained data partitioning and model matching method based on data characteristics. Compared with the prior art, the main improvements include: First, introducing a multi-dimensional feature vector to characterize data characteristics, including time series characteristics, statistical characteristics, and information entropy characteristics, making the description of data characteristics more comprehensive; Second, using the density peak clustering algorithm for data partitioning, which can better identify data clusters with irregular shapes compared with traditional methods such as K-means; Third, constructing a matching degree evaluation mechanism for data characteristics and model feature descriptors, and obtaining an accurate matching score through weighted calculation; Fourth, transforming the model allocation problem into a bipartite graph optimal matching problem, and obtaining a globally optimal matching scheme by minimizing the overall allocation cost. These improvements make data processing more refined, can select the most suitable model according to the internal characteristics of the data, and thus improve the model processing effect and resource utilization efficiency.
[0090] In an optional implementation manner, obtaining a local density value according to the relationship between the distance and a preset truncation distance by calculating the distance between the data points of the original data and other data points according to the feature vector includes:
[0091] Calculate the distance between the data points in the original data, where the distance is obtained by weighted summation of a dedicated distance function for each feature dimension;
[0092] Calculate the structural similarity degree between the data point and its neighboring data points, and multiply the structural similarity degree by an adjustment coefficient and then add one to obtain a structure-aware density value;
[0093] Based on the adjustment coefficient, calculate the arithmetic mean of the structure-aware density values by arithmetic mean, and divide the arithmetic mean by the structure-aware density value to obtain a local anomaly factor;
[0094] Multiply the difference between the local anomaly factor and a preset anomaly threshold by an attenuation coefficient, take the negative of the multiplication result and perform an exponential operation to obtain an attenuation factor, and multiply the attenuation factor by the structure-aware density value to obtain a corrected density value;
[0095] The corrected density value is weighted and combined with a preset density coefficient to obtain a multi-scale density representation. Calculate the product of the smoothing coefficient and the multi-scale density representation, as well as the product of the smoothing coefficient and the weighted multi-scale density value of neighboring data points. Add the two products to obtain the local density value.
[0096] Calculate the distances between data points in the original data. The distance here is not a simple Euclidean distance, but is obtained by weighted summation of dedicated distance functions for each feature dimension. Specifically, for each feature dimension, an appropriate distance function is selected according to the data characteristics of that dimension. For example, the Euclidean distance can be used for numerical features, and the Hamming distance can be used for categorical features, etc. Then, the distances of each dimension are weighted and summed, and the weights can be set according to the importance of each feature. For example, assuming there are 3 feature dimensions with weights of 0.5, 0.3, and 0.2 respectively, the final distance can be expressed as: 0.5×distance of feature 1 + 0.3×distance of feature 2 + 0.2×distance of feature 3.
[0097] Calculate the structural similarity degree between each data point and its neighboring data points. The structural similarity degree here reflects the distribution characteristics of data points in the local area. The structural similarity degree can be measured by comparing the distance distributions of data points and their neighboring points. For example, the average distance from a data point to its k nearest neighbor points can be calculated, and then compared with the average distance from each of these k neighboring points to their k nearest neighbors to obtain a similarity score. Then, multiply this structural similarity by an adjustment coefficient and add 1 to obtain the structure-aware density value. The adjustment coefficient is used to control the influence degree of the structural similarity on the density and can be adjusted according to the specific application scenario.
[0098] Calculate the arithmetic mean of the structure-aware density values based on the above adjustment coefficient through arithmetic mean. Specifically, the structure-aware density values of all data points can be summed and then divided by the total number of data points to obtain the average value. Divide this arithmetic mean by the structure-aware density value of each data point itself to obtain the local outlier factor. The local outlier factor reflects the degree of abnormality of a data point relative to the whole, and the larger the value, the more abnormal it is.
[0099] Subtract the local outlier factor from the preset outlier threshold, and multiply the difference by a decay coefficient. This decay coefficient is used to control the influence of the degree of abnormality on the final density. Take the negative of the multiplication result and perform an exponential operation to obtain the decay factor. The role of the decay factor is to adjust the density of outlier points. Multiply the decay factor by the previously obtained structure-aware density value to obtain the corrected density value. The purpose of this step is to reduce the density value of outlier points so that they are more easily recognized in subsequent clustering or outlier detection.
[0100] The corrected density values are weighted and combined with preset density coefficients to obtain a multi-scale density representation. Here, multi-scale means considering density information in neighborhoods of different ranges. Multiple density coefficients can be set, corresponding to neighborhoods of different sizes respectively, and then weighted combination is performed. For example, three scales of near neighbors, medium neighbors, and far neighbors can be set, with weights of 0.5, 0.3, and 0.2 respectively. Then, calculate the product of a smoothing coefficient and this multi-scale density representation, as well as the product of the smoothing coefficient and the weighted multi-scale density value of neighboring data points. The role of the smoothing coefficient is to locally smooth the density and reduce the influence of noise. Finally, add these two products to obtain the final local density value.
[0101] Illustrate this process through specific data cases. Suppose there is a two-dimensional data set containing 1000 sample points, and each sample point has two features x and y. First, calculate the distances between sample points. Here, assume the weights of x and y are 0.6 and 0.4 respectively. For sample points A(1, 2) and B(4, 6), the distance between them can be calculated as: 0.6×|4 - 1| + 0.4×|6 - 2| = 3.4.
[0102] Calculate the structural similarity. Suppose 10 nearest neighbors are selected. For sample point A, calculate the average distance to its 10 nearest neighbors as 2.5, and the average of the average distances of these 10 neighboring points to their 10 nearest neighbors is 3.0. Then the structural similarity of A can be expressed as 2.5 / 3.0 = 0.833. Assume the adjustment coefficient is 0.5, then the structure-aware density value of A is 1 + 0.5×0.833 = 1.4165.
[0103] Calculate the arithmetic mean of the structure-aware density values of all sample points, assumed to be 1.5. Then the local outlier factor of A is 1.5 / 1.4165 = 1.059. Assume the preset outlier threshold is 1.2 and the attenuation coefficient is 0.1, then the attenuation factor of A is exp(-0.1×(1.2 - 1.059)) = 0.986. Multiply this attenuation factor by the structure-aware density value of A to obtain the corrected density value 1.4165×0.986 = 1.397.
[0104] Suppose three density coefficients 0.5, 0.3, 0.2 are set, corresponding to the densities of 10, 20, 30 nearest neighbors respectively. The densities of A at these three scales are calculated as 1.397, 1.425, 1.410 respectively. Then the multi-scale density representation of A is 0.5×1.397 + 0.3×1.425 + 0.2×1.410 = 1.408. Assume the smoothing coefficient is 0.8, and the average multi-scale density of A's 10 nearest neighbors is 1.420, then the final local density value of A is 0.8×1.408 + 0.2×1.420 = 1.4104.
[0105] In an alternative embodiment, the edge computing node determines the computing priority of each target artificial intelligence model according to the computing complexity of each target artificial intelligence model and the computing resource status of the edge computing node, and sequentially calls the computing resources of the edge computing node to execute each target artificial intelligence model according to the computing priority, and generates corresponding model computing results, including:
[0106] Based on the edge computing node, the weighted sum of the products of the operands and layer weight coefficients in each computing layer of the target artificial intelligence model is used to obtain the computing complexity;
[0107] Construct a resource status vector including processor occupancy, memory occupancy, and bandwidth occupancy, and determine the resource availability based on the minimum value of the ratio of the remaining amount to the maximum capacity of each type of resource according to the resource status vector;
[0108] Multiply the normalized result of the computing complexity and the resource availability by the corresponding weight coefficients respectively to obtain a priority score; construct a multi-objective optimization function based on the priority score, and the multi-objective optimization function includes target constraint conditions, resource constraint conditions, and latency constraint conditions;
[0109] Determine the execution order of the target artificial intelligence model according to the solution result of the multi-objective optimization function, and sequentially call the computing resources of the edge computing node to execute each target artificial intelligence model according to the execution order, and generate corresponding model computing results.
[0110] When the edge computing node evaluates the computing complexity of the target artificial intelligence model, the weighted sum of the products of the operands and layer weight coefficients in each computing layer of the target artificial intelligence model is used. Specifically, for a convolutional neural network model, the model structure can be parsed into multiple computing layers, including convolutional layers, pooling layers, fully connected layers, etc. For each layer, count its operation quantity. For example, the operation quantity of the convolutional layer can be calculated through parameters such as the input feature map size, convolutional kernel size, and number of convolutional kernels; the operation quantity of the pooling layer can be calculated through parameters such as the input feature map size and pooling window size; the operation quantity of the fully connected layer can be calculated through the product of the number of input neurons and the number of output neurons.
[0111] For a target artificial intelligence model including 3 convolutional layers and 2 fully connected layers, the operation quantities of each layer are 10 million, 8 million, 5 million, 2 million, and 1 million respectively, and the corresponding layer weight coefficients are 0.3, 0.25, 0.2, 0.15, and 0.1 respectively.
[0112] Then the computing complexity of this model is:
[0113] 10 million × 0.3 + 8 million × 0.25 + 5 million × 0.2 + 2 million × 0.15 + 1 million × 0.1 = 3 million + 2 million + 1 million + 0.3 million + 0.1 million = 6.4 million.
[0114] The edge computing node constructs a resource status vector including the processor occupancy rate, memory occupancy rate, and bandwidth occupancy rate, and determines the resource availability based on the resource status vector according to the minimum value of the ratio of the remaining amount to the maximum capacity of various resources. Specifically, the CPU usage rate, memory usage rate, and network bandwidth usage rate of the edge computing node are periodically collected to form a resource status vector. For example, the resource status vector at a certain moment is [70%, 60%, 50%], indicating that the CPU usage rate is 70%, the memory usage rate is 60%, and the bandwidth usage rate is 50%.
[0115] When calculating the resource availability based on the resource status vector, first calculate the ratio of the remaining amount to the maximum capacity of various resources, that is, [1 - 70%, 1 - 60%, 1 - 50%] = [30%, 40%, 50%], and then take the minimum value among them as the resource availability, which is 30% in this example.
[0116] The edge computing node first normalizes the computational complexity and maps it to the interval [0, 1] to obtain the normalized computational complexity. For example, assuming that the maximum value of the computational complexity of all models in the system is 10 million and the minimum value is 1 million, then the normalized value of the model with a computational complexity of 6.4 million is (6.4 - 1) / (10 - 1)=0.6. Then, the normalized computational complexity and the resource availability are multiplied by the corresponding weight coefficients respectively and summed to obtain the priority score. Let the weight coefficient of the normalized computational complexity be 0.6 and the weight coefficient of the resource availability be 0.4. For the target artificial intelligence model with a normalized computational complexity of 0.6, under the condition that the resource availability is 30%, its priority score is: 0.6×0.6 + 30%×0.4 = 0.36 + 0.12 = 0.48.
[0117] Based on the priority score, a multi-objective optimization function is constructed. The multi-objective optimization function includes objective constraint conditions, resource constraint conditions, and latency constraint conditions. The objective constraint conditions specify that all target artificial intelligence models that need to be executed must be scheduled; the resource constraint conditions specify that all scheduled target artificial intelligence models do not exceed the resource capacity limit of the edge computing node during execution; the latency constraint conditions specify that the execution time of all scheduled target artificial intelligence models does not exceed the preset latency threshold.
[0118] For three target artificial intelligence models A, B, and C, their priority scores are 3.8412 million, 2.9635 million, and 1.5278 million respectively. The objective constraint conditions of the multi-objective optimization function are that all models must be executed; the resource constraint conditions are that the CPU usage rate does not exceed 95%, the memory usage rate does not exceed 90%, and the bandwidth usage rate does not exceed 85%; the latency constraint condition is that the total model execution time does not exceed 200 milliseconds.
[0119] Determine the execution order of the target artificial intelligence models according to the solution results of the multi-objective optimization function, and call the computing resources of the edge computing nodes in the execution order to execute each target artificial intelligence model in turn to generate corresponding model computing results.
[0120] The edge computing node can use the greedy algorithm or the dynamic programming algorithm to solve the multi-objective optimization function. Taking the greedy algorithm as an example, arrange the model execution in the order of decreasing priority score. For the above three models, the execution order is A→B→C.
[0121] The edge computing node monitors the resource usage in real time to ensure that the resource usage does not exceed the constraint conditions. For example, after model A is executed, the CPU usage rate rises to 85%, the memory usage rate rises to 75%, and the bandwidth usage rate rises to 65%; after model B is executed, the CPU usage rate rises to 92%, the memory usage rate rises to 82%, and the bandwidth usage rate rises to 75%; before model C is executed, judge whether the resource usage will exceed the constraint conditions after executing C. If it will not exceed, then execute model C; if it will exceed, then suspend the execution of model C and wait for the resources to be released before executing.
[0122] The edge computing node records the execution time of each model to ensure that the total execution time does not exceed the latency constraint conditions. For example, the execution time of model A is 80 milliseconds, the execution time of model B is 60 milliseconds, the expected execution time of model C is 50 milliseconds, and the total execution time is 190 milliseconds, which does not exceed the threshold of 200 milliseconds. Therefore, it can be executed in the planned order.
[0123] The edge computing node collects the computing results of each model, such as the object detection result of model A, the image classification result of model B, and the speech recognition result of model C, and sends these results to the requester or proceeds to the next step of processing.
[0124] Figure 3 The following is the flowchart of the artificial intelligence model scheduling in the edge computing environment of the embodiment of the present invention:
[0125] The figure shows a complete model scheduling optimization process, which includes two parallel initial input branches: the left branch starts from the computational complexity analysis of the target artificial intelligence model, and obtains an accurate computational complexity evaluation value by calculating the product of the number of operations of each layer of the model and the corresponding layer weight coefficient and performing weighted summation; the right branch starts from the resource status of the edge computing node, and calculates the current resource availability by constructing a multi-dimensional resource status vector including processor utilization, memory occupancy, network bandwidth, etc. The results of these two branches converge in the middle of the process. The system obtains a comprehensive priority score by weighted combination of the normalized result of the computational complexity and the resource availability. Based on this score, the system further constructs a multi-objective optimization function, which simultaneously considers target constraint conditions (such as model accuracy requirements), resource constraint conditions (such as computational resource limitations), and latency constraint conditions (such as execution time limitations). By solving this multi-objective optimization function, the execution order of the target artificial intelligence model is finally determined, and the computational resources are allocated accordingly to execute the corresponding model. The entire process shows a complete technical route from model complexity evaluation to resource status perception and then to optimized scheduling. By considering multi-dimensional constraint conditions, the efficiency and reliability of model scheduling in the edge computing environment are ensured.
[0126] Existing methods for scheduling artificial intelligence models on edge computing nodes mainly rely on the principles of fixed priority or first-come-first-served, and fail to fully consider the computational complexity of the models and the resource status of the nodes, resulting in low utilization efficiency of computational resources and inability to meet the requirements of multi-model parallel processing. For example, some existing technologies adopt a polling scheduling algorithm to execute each model in a fixed order, or a random scheduling algorithm to randomly select a model for execution. These methods cannot dynamically adjust the execution order according to the actual situation.
[0127] This application proposes a dynamic scheduling method based on computational complexity and resource status. By accurately evaluating the model computational complexity and real-time monitoring the resource status, a multi-objective optimization function is constructed to guide the scheduling decision. Compared with the existing technologies, the improvements of this application are as follows: First, a computational complexity evaluation method based on the number of operations of each computational layer of the model and the layer weights is introduced, which improves the accuracy of complexity evaluation; second, a resource availability calculation method based on a multi-dimensional resource status vector is proposed, considering various resource factors such as processors, memory, and bandwidth; third, an optimization function including various constraint conditions is designed to achieve the optimal scheduling under the premise of meeting resource and latency constraints. These improvements enable the edge computing node to more efficiently utilize computational resources, improve the processing efficiency of artificial intelligence models, and meet the real-time requirements.
[0128] In an alternative embodiment, determining the execution order of the target artificial intelligence models according to the solution result of the multi-objective optimization function, and calling the computing resources of the edge computing nodes in sequence according to the execution order to execute each of the target artificial intelligence models, generating corresponding model calculation results includes:
[0129] Solving the multi-objective optimization function to obtain the execution order, where the solution process of the multi-objective optimization function is restricted by a resource constraint condition that the resource usage does not exceed the available resource capacity, and a latency constraint condition that the model execution time does not exceed the maximum allowable latency;
[0130] Sequentially calling the computing resources of the edge computing nodes according to the execution order to execute each target artificial intelligence model, and generating corresponding model calculation results.
[0131] Obtain multiple target artificial intelligence models to be executed and their related parameter information, including the resource requirements, execution time, etc. of each model. Then, construct a multi-objective optimization function, and the objectives of this function include minimizing the total execution time, maximizing the resource utilization rate, etc.
[0132] Solve this multi-objective optimization function to determine the optimal execution order of the target artificial intelligence models. During the solution process, two main constraint conditions need to be considered: resource constraint and latency constraint. The resource constraint requires that the total amount of computing resources used by each model does not exceed the available resource capacity of the edge node; the latency constraint requires that the total execution time of the model does not exceed the preset maximum allowable latency.
[0133] Randomly generate a certain number of initial solutions, and each solution represents a possible model execution order. Then, generate new solutions through crossover and mutation operations, and use the roulette wheel selection method to select excellent individuals to enter the next generation. When evaluating the fitness of the solutions, consider both the total execution time and the resource utilization rate as objectives, and introduce a penalty factor to handle the solutions that do not meet the constraint conditions. After multiple generations of iteration, finally obtain the Pareto optimal solution set, and select a compromise solution from it as the final model execution order.
[0134] Suppose there are three artificial intelligence models A, B, and C to be executed, their resource requirements are 2, 3, and 4 units respectively, and their execution times are 10, 15, and 20 ms respectively. The available resources of the edge node are 5 units, and the maximum allowable latency is 40 ms. By solving the multi-objective optimization function, a possible optimal execution order may be B - A - C.
[0135] Send an execution request to the edge node, including the relevant information of model B. After receiving the request, the edge node allocates 3 units of computing resources to execute model B. After model B is executed, the occupied resources are released, and the calculation result is returned to the system.
[0136] The edge node allocates 2 units of resources to execute Model A. After completion, it also returns the result and releases the resources.
[0137] Since Model C requires 4 units of resources, which exceeds the current available resource amount, the system needs to wait for the previous model to finish execution and release resources before it can execute. When resources are sufficient, the edge node allocates 4 units of resources to execute Model C and returns the result after completion.
[0138] Monitor the actual execution time and resource usage of each model in real time. If it is found that the execution time of a certain model significantly deviates from the expectation, or the resource usage exceeds the preset value, the system will timely adjust the execution plan of the subsequent models. For example, if the actual execution time of Model B reaches 18 ms, exceeding the expected 15 ms, the system may decide to skip executing Model C to ensure that the total execution time does not exceed the maximum allowable delay of 40 ms.
[0139] To further optimize the execution efficiency, the system can also execute the models in a pipelined manner. In the above example, when Model B has been executed to a certain extent, if there are sufficient remaining resources, the system can start the execution of Model A simultaneously, thereby reducing the overall execution time.
[0140] After all the target artificial intelligence models have been executed, the system integrates and analyzes the calculation results of each model collected. For example, if these three models are respectively used for image recognition, speech processing, and natural language understanding, the system can combine their output results to generate a multi-modal intelligent analysis report.
[0141] Evaluate and summarize the current execution process. The evaluation metrics include the actual total execution time, resource utilization rate, the execution effects of each model, etc. This information will be used to optimize future model execution strategies, such as adjusting the priorities of the models, updating the resource requirement estimates, etc.
[0142] Efficiently execute multiple artificial intelligence models in the edge computing environment, while meeting the delay and resource constraints, to achieve the optimization of execution time and resource utilization. This method has strong adaptability and scalability, and can dynamically adjust the execution strategy according to the actual situation, and is applicable to various complex edge computing scenarios.
[0143] In an alternative implementation, the edge computing node aggregates and processes the model calculation results of each of the target artificial intelligence models to obtain the processing result for the original data, including:
[0144] Construct a credibility evaluation function for the model results. The credibility evaluation function is obtained based on the weighted combination of the historical accuracy of each target artificial intelligence model, the model execution quality under the current resource state, and the result output time. Apply the credibility evaluation function to the model calculation results of each of the target artificial intelligence models;
[0145] Perform adaptive weighting on the model calculation results of each based on the credibility evaluation function. The weight coefficient of the adaptive weighting is dynamically adjusted according to the change of the credibility evaluation function. The weight coefficient ensures that the sum of the weights of the model calculation results of each is one through normalization processing;
[0146] Use a sliding window mechanism based on temporal correlation for the weighted model calculation results of each. Combine the weighted results within the sliding window to obtain the processing result for the original data.
[0147] The edge computing node receives the calculation results of multiple target artificial intelligence models for the original data. These models can be of different types, such as image recognition models, speech recognition models, etc. They respectively process the original data and output results.
[0148] Construct a credibility evaluation function for the model results. This function comprehensively considers three key factors: the historical accuracy of each target artificial intelligence model, the model execution quality under the current resource state, and the result output time. Specifically, the historical accuracy can be obtained by long-term statistics of the matching degree between the prediction results and the actual results of each model; the model execution quality can be evaluated by monitoring the usage of current computing resources (such as CPU utilization, memory occupancy, etc.); the result output time directly measures the time required for the model to generate results. These three factors are combined through weighting to obtain the final credibility evaluation value.
[0149] Suppose there are three models A, B, and C, with their historical accuracies being 0.9, 0.85, and 0.8 respectively, the execution quality scores under the current resource state being 0.95, 0.9, and 0.85, and the result output times (after normalization) being 0.8, 0.9, and 0.7. We can assign weights to these three factors, such as 0.5, 0.3, and 0.2. Then the credibility evaluation value of model A is: 0.9×0.5 + 0.95×0.3 + 0.8×0.2 = 0.895. Similarly, the credibility evaluation values of B and C can be calculated.
[0150] Apply the credibility evaluation function to the calculation results of each target artificial intelligence model. This step assigns a credibility score to the result of each model, reflecting the reliability of the result.
[0151] Based on the results of the credibility evaluation function, adaptively weight the calculation results of each model. The key here is that the weight coefficient will be dynamically adjusted with the change of the credibility evaluation function. When implementing specifically, the credibility evaluation value of each model can be divided by the sum of the credibility evaluation values of all models to obtain the weight of the result of this model. This method naturally ensures that the sum of all weights is 1.
[0152] Suppose the credibility evaluation values of models A, B, and C are 0.895, 0.875, and 0.785 respectively. Then their weights can be calculated as follows:
[0153] A: 0.895 / (0.895 + 0.875 + 0.785) ≈ 0.35;
[0154] B: 0.875 / (0.895 + 0.875 + 0.785) ≈ 0.34;
[0155] C: 0.785 / (0.895 + 0.875 + 0.785) ≈ 0.31;
[0156] This adaptive weighting mechanism can dynamically adjust the influence of the model in the final result according to the real-time performance of the model, improving the flexibility and robustness of the system.
[0157] Adopt a sliding window mechanism based on temporal correlation to combine the weighted model calculation results to obtain the final processing result for the original data. The sliding window mechanism takes into account the time continuity of the data and is especially suitable for processing time series data.
[0158] A fixed-size time window can be set, such as 5 minutes. Within this window, collect the weighted results of all models. As time goes by, the window slides forward continuously, always keeping the latest 5-minute data. For the data within the window, various combination methods can be adopted, such as simple average, weighted average, median, etc. The specific choice depends on the requirements of the application scenario.
[0159] Suppose within a certain 5-minute window, the three models made multiple predictions on the same target, and the weighted results are as follows:
[0160] A: [0.7, 0.72, 0.68, 0.71, 0.69];
[0161] B: [0.65, 0.67, 0.66, 0.68, 0.64];
[0162] C: [0.62, 0.63, 0.61, 0.64, 0.62];
[0163] Using the simple average method, the average prediction value of each model within this window can be obtained:
[0164] A: 0.7, B: 0.66, C: 0.624;
[0165] Considering the previously calculated model weights (A: 0.35, B: 0.34, C: 0.31), the final weighted combination result can be obtained:
[0166] 0.7×0.35 + 0.66×0.34 + 0.624×0.31 ≈ 0.663;
[0167] This value is the final result of the system's processing of the original data within this time window.
[0168] It should be noted that the size of the sliding window can be adjusted according to the actual application requirements. A larger window can provide more stable results, but may reduce the system's response speed to sudden events; a smaller window can reflect data changes faster, but may introduce more noise. In practical applications, the most suitable window size can be determined through experiments.
[0169] It is also possible to set an update frequency, such as updating the result once per second or per minute. Each time an update occurs, the credibility of the model is recalculated, the weights are adjusted, and a new processing result is generated based on the latest sliding window data. This can ensure that the system always uses the latest and most relevant information to generate results.
[0170] The edge computing node can effectively aggregate and process the calculation results of multiple target artificial intelligence models to obtain high-quality processing results for the original data. This method not only considers the historical performance and current state of each model, but also captures the temporal characteristics of the data through the sliding window mechanism, thus providing a flexible and reliable data processing solution.
[0171] Figure 4 The flowchart of the multi-model result fusion based on credibility evaluation for the embodiments of the present invention is as follows:
[0172] This figure details a complete fusion process from the model calculation results to the final processing results. The process starts with the input of the calculation results of each target artificial intelligence model, and then a credibility evaluation function based on three key dimensions is constructed: historical accuracy (reflecting the long-term performance of the model), model execution quality (reflecting the processing ability under the current resource state), and result output time (characterizing the processing efficiency). These three dimensions are combined through weighted combination to form the credibility evaluation function. On this basis, the system performs adaptive weighted processing on the model calculation results, where the weight coefficients will be dynamically adjusted according to the changes in the credibility evaluation results, and through normalization, it is ensured that the sum of all weights is one, guaranteeing the scientificity and rationality of the weighting process. Finally, the system adopts a sliding window mechanism based on temporal correlation to combine the weighted results within the sliding window, and finally generates the processing results for the original data. This fusion scheme not only considers the historical performance of the model but also takes into account the current execution state, and at the same time ensures the timeliness and continuity of the results through the consideration of temporal correlation, forming a complete multi-model result fusion technical solution.
[0173] In the second aspect of the embodiments of the present invention,
[0174] There is provided an artificial intelligence model service system based on edge computing, including:
[0175] A first unit for receiving an artificial intelligence model service request sent by a mobile terminal, where the artificial intelligence model service request includes raw data to be processed;
[0176] A second unit for performing data feature analysis on the raw data based on an edge computing node, dividing the raw data into multiple data subsets according to the data features of the raw data, and selecting corresponding target artificial intelligence models from a preset artificial intelligence model library based on each data subset, so that each data subset is processed by a target artificial intelligence model that matches its data features;
[0177] A third unit for the edge computing node to determine the calculation priority of each target artificial intelligence model according to the calculation complexity of each target artificial intelligence model and the calculation resource state of the edge computing node, and sequentially call the calculation resources of the edge computing node to execute each target artificial intelligence model according to the calculation priority, generating corresponding model calculation results;
[0178] A fourth unit for the edge computing node to summarize and process the model calculation results of each target artificial intelligence model, obtain the processing results for the raw data, and return the processing results to the mobile terminal.
[0179] In the third aspect of the embodiments of the present invention,
[0180] Provided is an electronic device, comprising:
[0181] a processor;
[0182] a memory for storing instructions executable by the processor;
[0183] wherein the processor is configured to call the instructions stored in the memory to execute the method described above.
[0184] In a fourth aspect of the embodiments of the present invention,
[0185] a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0186] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present invention.
[0187] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An artificial intelligence model service method based on edge computing, characterized in that, Including: Receiving an artificial intelligence model service request sent by a mobile terminal, where the artificial intelligence model service request includes raw data to be processed; Based on an edge computing node, performing data feature analysis on the raw data, dividing the raw data into multiple data subsets according to the data features of the raw data, and selecting a corresponding target artificial intelligence model from a preset artificial intelligence model library based on each data subset, so that each data subset is processed by a target artificial intelligence model matching its data features; The edge computing node determines the calculation priority of each target artificial intelligence model according to the calculation complexity of each target artificial intelligence model and the calculation resource status of the edge computing node, and sequentially calls the calculation resources of the edge computing node to execute each target artificial intelligence model according to the calculation priority, generating corresponding model calculation results; The edge computing node performs summary processing on the model calculation results of each target artificial intelligence model, obtains a processing result for the raw data, and returns the processing result to the mobile terminal; Obtaining a local density value according to the relationship between the distance and a preset truncation distance by calculating the distance between the data points of the raw data and other data points through a feature vector includes: Calculating the distance between data points in the raw data, where the distance is obtained by weighted summation of a dedicated distance function for each feature dimension; Calculating the structural similarity degree between the data point and its neighboring data points, multiplying the structural similarity degree by a regulation coefficient and then adding one to obtain a structure-aware density value; Calculating the arithmetic mean of the structure-aware density values by arithmetic mean based on the regulation coefficient, dividing the arithmetic mean by the structure-aware density value to obtain a local anomaly factor; Multiplying the difference between the local anomaly factor and a preset anomaly threshold by an attenuation coefficient, taking the negative of the multiplication result and performing an exponential operation to obtain an attenuation factor, and multiplying the attenuation factor by the structure-aware density value to obtain a corrected density value; Performing weighted combination of the corrected density value and a preset density coefficient to obtain a multi-scale density representation, calculating the product of a smoothing coefficient and the multi-scale density representation, and the product of the smoothing coefficient and the weighted multi-scale density value of neighboring data points, and adding the two products to obtain a local density value.
2. The method according to claim 1, characterized in that Dividing the raw data into multiple data subsets according to the data features of the raw data, and selecting a corresponding target artificial intelligence model from a preset artificial intelligence model library based on each data subset, so that each data subset is processed by a target artificial intelligence model matching its data features includes: Extracting the time series features, statistical features, and information entropy features in the raw data to construct a feature vector; calculating the distance between the data points of the raw data and other data points through the feature vector, and obtaining a local density value according to the relationship between the distance and a preset truncation distance; Performing clustering division on the raw data based on the local density value to obtain multiple data subsets; The matching degree score is composed of the weighted sum of the feature vectors of multiple said data subsets and a preset model feature descriptor; a priority queue is constructed with the said data subsets, the corresponding model feature descriptors, and the matching degree score. The optimal matching relationship between each said data subset and the model is determined by minimizing the allocation cost, and each said data subset is matched with the model that matches its data features based on the optimal matching relationship.
3. The method according to claim 1, wherein The edge computing node determines the computing priority of each said target artificial intelligence model according to the computing complexity of each said target artificial intelligence model and the computing resource status of the edge computing node, and sequentially calls the computing resources of the edge computing node to execute each said target artificial intelligence model according to the computing priority, generating corresponding model calculation results, including: The computing complexity is obtained by performing a weighted sum of the products of the operands and layer weight coefficients of each computing layer in the target artificial intelligence model based on the edge computing node. A resource status vector including the processor occupancy rate, memory occupancy rate, and bandwidth occupancy rate is constructed, and the resource availability is determined based on the resource status vector according to the minimum value of the ratio of the remaining amount to the maximum capacity of various resources. The normalized result of the computing complexity and the resource availability are respectively multiplied by the corresponding weight coefficients to obtain a priority score; a multi-objective optimization function is constructed based on the priority score, and the multi-objective optimization function includes objective constraint conditions, resource constraint conditions, and latency constraint conditions. The execution order of the target artificial intelligence models is determined according to the solution result of the multi-objective optimization function, and the computing resources of the edge computing node are called sequentially according to the execution order to execute each said target artificial intelligence model, generating corresponding model calculation results.
4. The method according to claim 3, characterized in that, The execution order of the target artificial intelligence models is determined according to the solution result of the multi-objective optimization function, and the computing resources of the edge computing node are called sequentially according to the execution order to execute each said target artificial intelligence model, generating corresponding model calculation results, including: The execution order is obtained by solving the multi-objective optimization function, and the solution process of the multi-objective optimization function is restricted by the resource constraint condition that the resource usage does not exceed the available resource capacity, and the latency constraint condition that the model execution time does not exceed the maximum allowable latency. The computing resources of the edge computing node are called sequentially according to the execution order to execute each target artificial intelligence model, and corresponding model calculation results are generated.
5. The method according to claim 1, characterized in that, The edge computing node aggregates the model calculation results of each said target artificial intelligence model to obtain the processing result for the original data, including: A credibility evaluation function for the model results is constructed, and the credibility evaluation function is obtained based on the weighted combination of the historical accuracy of each target artificial intelligence model, the model execution quality under the current resource status, and the result output time, and the credibility evaluation function is applied to the model calculation results of each said target artificial intelligence model. Adaptive weighting is performed on the calculation results of each of the models based on the credibility evaluation function, and the weight coefficient of the adaptive weighting is dynamically adjusted according to the change of the credibility evaluation function. The weight coefficient ensures that the sum of the weights of the calculation results of each of the models is one through normalization processing; The calculation results of each of the weighted models are combined by using a sliding window mechanism based on time series correlation to obtain a processing result for the original data.
6. An artificial intelligence model service system based on edge computing, which is used to implement the method described in any one of the foregoing claims 1-5, and is characterized in that, It includes: A first unit, configured to receive an artificial intelligence model service request sent by a mobile terminal, where the artificial intelligence model service request includes original data to be processed; A second unit, configured to perform data feature analysis on the original data based on an edge computing node, divide the original data into multiple data subsets according to the data features of the original data, and select a corresponding target artificial intelligence model from a preset artificial intelligence model library for each data subset, so that each data subset is processed by a target artificial intelligence model that matches its data features; A third unit, configured to determine the calculation priority of each target artificial intelligence model according to the calculation complexity of each target artificial intelligence model and the calculation resource status of the edge computing node, and sequentially call the calculation resources of the edge computing node to execute each target artificial intelligence model according to the calculation priority to generate corresponding model calculation results; A fourth unit, configured to summarize the model calculation results of each of the target artificial intelligence models by the edge computing node to obtain a processing result for the original data, and return the processing result to the mobile terminal.
7. An electronic device, characterized in that: It includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Cloud data processing system based on artificial intelligence algorithm
CN119473645A