Artificial intelligence-oriented model warehouse platform and model service platform

By integrating data management, inference task scheduling, and computing power management modules, the unified architecture problem of model repository and service platform is solved, realizing efficient model management and resource optimization, and improving the efficiency of model storage and deployment as well as platform stability.

CN121560466APending Publication Date: 2026-02-24COVID (NANJING) INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511404681.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

The lack of a unified architecture in existing model repositories and service platforms leads to low efficiency in the connection between model storage and inference services, increases engineering costs and deployment cycles, and makes it difficult to achieve seamless management and optimization of the entire model lifecycle, especially when scheduling cross-data source and heterogeneous computing resources.

Method used

The data management module processes heterogeneous data from different data sources, the model repository and service module provides storage and version control, the inference task scheduling and optimization module performs task scheduling and resource allocation, the computing power management module allocates resources, and the monitoring module performs multi-dimensional monitoring and predicts resource requirements through predictive models, thereby achieving optimal utilization of computing resources and improving the quality of platform services.

Benefits of technology

It enables centralized management and one-stop service of models, simplifies the process from storage to deployment, improves management efficiency and reusability, and ensures optimal utilization of computing resources and platform stability and service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121560466A_ABST
    Figure CN121560466A_ABST
Patent Text Reader

Abstract

The invention provides an artificial intelligence-oriented model warehouse platform and model service platform, and relates to the technical field of artificial intelligence and information, and the platform comprises a data management module which is used for accessing and processing cross-data-source heterogeneous data from an external data source, and preprocessing the cross-data-source heterogeneous data to obtain preprocessed heterogeneous data. According to the artificial intelligence-oriented model warehouse platform and model service platform, task scheduling and resource allocation are performed according to performance indexes through the cooperation of the reasoning task scheduling and optimization module and the computing power management module, so that optimal utilization of computing resources is ensured. The monitoring module combines multi-dimensional monitoring and an artificial intelligence algorithm, can monitor the operation state of the platform in real time, pre-judges resource demands and service quality changes through a prediction model, automatically triggers alarms and optimizes measures, and improves the stability and service quality of the platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and information technology, specifically to an artificial intelligence model repository platform and model service platform. Background Technology

[0002] As the application scope of artificial intelligence (AI) technology continues to expand, the number and complexity of models are also increasing. To facilitate centralized management and efficient access, research institutions and enterprises have gradually established model repositories and service platforms. Existing model repositories are mainly used for model storage, version management, and retrieval, ensuring model reusability and traceability to a certain extent. Service platforms provide online deployment and inference capabilities for models, leveraging cloud computing and containerization technologies to support multi-tenant management, elastic scaling, and distributed services, enabling AI models to be quickly applied to production systems and business scenarios. Some systems have also introduced modules for task scheduling, computing power management, and resource monitoring to improve service quality and resource utilization efficiency.

[0003] However, existing technologies still have shortcomings. Model repositories and service platforms are mostly designed to operate independently, lacking a unified architecture, resulting in inefficient integration from model storage to inference services. Model deployment typically requires additional migration and adaptation operations, increasing engineering costs and deployment cycles. More importantly, the model versions, performance metrics recorded in the repository, and the actual operational performance on the server side do not form a complete closed loop, making it difficult for users to achieve comprehensive management and optimization throughout the model's entire lifecycle. This problem is even more pronounced in complex scenarios involving cross-data sources and heterogeneous computing resource scheduling, hindering the large-scale deployment and efficient application of artificial intelligence models. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides an artificial intelligence model repository platform and model service platform. The technical problem this invention aims to solve is: how to use an inference task scheduling and optimization module to schedule tasks and allocate resources based on performance indicators, thereby achieving optimal utilization of computing resources and improving the quality of platform services.

[0005] To achieve the above objectives, the present invention provides the following technical solution: an artificial intelligence model repository platform and model service platform, comprising: a data management module, wherein the data management module is used to access and process cross-data source heterogeneous data from external data sources, and to preprocess the cross-data source heterogeneous data to obtain preprocessed heterogeneous data to ensure the consistency of input data.

[0006] The model repository and service module receives the preprocessed heterogeneous data, extracts metadata information of the artificial intelligence model, provides storage, version control and one-stop management of the artificial intelligence model, and supports access to multiple model formats and frameworks.

[0007] The inference task scheduling and optimization module extracts performance metrics from the metadata information, generates task scheduling results based on the performance metrics, and executes inference tasks based on the task scheduling results.

[0008] The computing power management module allocates resources to the heterogeneous computing hardware resources based on the task scheduling results.

[0009] The monitoring module performs multi-dimensional monitoring of the platform and generates operational status monitoring indicators. The monitoring module establishes a prediction model based on artificial intelligence algorithms, which include automatic alarms and adaptive optimization strategies. When the operational status monitoring indicators exceed the threshold, the automatic alarm is triggered. The prediction model is used to predict the future trends of the computing resource utilization and service quality indicators, thereby improving platform stability and service quality.

[0010] Preferably, the external data source includes external business databases, data warehouses, data lakes, and real-time data streaming systems; the cross-data source heterogeneous data includes relational database data, non-relational database data, data warehouse data, data lake data, and real-time data streams; and the preprocessing includes cleaning, optimization, and preprocessing recommendations.

[0011] Preferably, the preprocessing recommends using a performance scoring function to comprehensively evaluate the heterogeneous data from different data sources. The calculation formula for the performance scoring function is as follows: .

[0012] in, This represents the preprocessing workflow score, which is dimensionless. This indicates the preprocessing time, in seconds. This indicates resource utilization rate, expressed as a percentage. and By mapping the normalized coefficients to a space of the same dimensions These are weighting coefficients, dimensionless. The weighting coefficients are dimensionless and satisfy the following conditions: .

[0013] Preferably, the metadata information includes model version information, training parameters and performance metrics. The model repository and service module provide a unified calling interface, which supports loading and calling the artificial intelligence model via RPC to achieve rapid deployment and service-oriented architecture of the model.

[0014] Preferably, the performance metrics include inference latency, throughput, and accuracy.

[0015] Preferably, the heterogeneous computing hardware resources include a central processing unit, a graphics processing unit, and a tensor processing unit.

[0016] Preferably, the operational status monitoring indicators include computing resource utilization and service quality indicators. The monitoring module collects real-time monitoring data and stores it as historical monitoring data during operation. The multi-dimensional monitoring includes inference response latency, system throughput, and task completion rate. The threshold is determined based on the statistical distribution of the historical monitoring data. The statistical distribution includes the mean and variance of the computing resource utilization and the fluctuation range of the service quality indicators. When the value of the real-time monitoring data exceeds the threshold, the monitoring module triggers the automatic alarm.

[0017] Preferably, the prediction model predicts future computing resource requirements based on the historical monitoring data and the real-time monitoring data, and the computing resource requirements include processor computing power requirements, graphics processor computing power requirements, and storage bandwidth requirements.

[0018] This invention provides an artificial intelligence model repository platform and model service platform. It has the following beneficial effects: The artificial intelligence model repository platform and model service platform provided by this invention enable centralized management, version control, and one-stop service for models, supporting the access of multiple model formats and frameworks. Through the integration of the model repository and service modules, the platform can efficiently store and manage artificial intelligence models, simplifying the process from model storage to deployment and greatly improving model management efficiency and reusability.

[0019] This AI model repository and model service platform, through the collaboration of an inference task scheduling and optimization module and a computing power management module, schedules tasks and allocates resources based on performance indicators, thereby ensuring optimal utilization of computing resources. The monitoring module, combining multi-dimensional monitoring and AI algorithms, can monitor the platform's operational status in real time and predict resource demands and service quality changes through predictive models, automatically triggering alarms and optimization measures to improve platform stability and service quality. Attached Figure Description

[0020] Figure 1 This is an architecture diagram of an artificial intelligence model repository and service platform.

[0021] Figure 2 This is a diagram illustrating the data flow and processing procedure.

[0022] Figure 3 This is a schematic diagram of the model repository and version management module. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Example 1

[0024] like Figure 1-3 As shown in the illustration, this invention provides an artificial intelligence model warehouse platform and model service platform, including a data management module. The data management module is used to access and process heterogeneous data from external data sources across different data sources. It preprocesses the heterogeneous data to obtain preprocessed heterogeneous data to ensure input data consistency. External data sources include external business databases, data warehouses, data lakes, and real-time data streaming systems. The heterogeneous data from different data sources includes relational database data, non-relational database data, data warehouse data, data lake data, and real-time data streams. Preprocessing includes cleaning, optimization, and preprocessing recommendation.

[0025] External business databases: such as traditional relational database systems.

[0026] Data warehouse: historical data and complex query results.

[0027] Data lake: Stores structured and unstructured data, usually in its raw format.

[0028] Real-time data streaming systems, such as message queues or stream processing systems, are used to receive real-time event data.

[0029] Data cleaning: Remove redundant data and invalid or missing data items to ensure data quality.

[0030] Optimization: Adjust the data format according to requirements, convert unstructured data into structured data, and standardize the data format.

[0031] Preprocessing recommendation: Based on the characteristics of different data sources and application scenarios, the data management module uses a performance scoring function to evaluate the data and generate an optimal preprocessing scheme to ensure that the data can provide consistency and efficiency for subsequent machine learning models.

[0032] Preprocessing is recommended to use a performance scoring function to comprehensively evaluate heterogeneous data across data sources. The formula for calculating the performance scoring function is as follows: .

[0033] in, This represents the preprocessing workflow score, which is dimensionless. This indicates the preprocessing time, in seconds. This indicates resource utilization rate, expressed as a percentage. and By mapping the normalized coefficients to a space of the same dimensions These are weighting coefficients, dimensionless. The weighting coefficients are dimensionless and satisfy the following conditions: .

[0034] The model repository and service module receives preprocessed heterogeneous data, extracts metadata information from AI models, and provides storage, version control, and one-stop management for AI models. It supports access to various model formats and frameworks. Metadata information includes model version information, training parameters, and performance metrics. The module provides a unified API that supports loading and calling AI models via RPC, enabling rapid model deployment and service-oriented architecture.

[0035] Model versions: Different versions of each model at different points in time, making it easier to track and manage.

[0036] Training parameters: These include configurations for training such as learning rate, number of iterations, and batch size.

[0037] Performance metrics include evaluation metrics such as model accuracy, precision, and recall.

[0038] Framework information: The framework used by the model.

[0039] Input / output format: The input and output data format of the model, ensuring compatibility with data in real-world application scenarios.

[0040] The inference task scheduling and optimization module extracts performance metrics from metadata, generates task scheduling results based on these metrics, and then executes the inference tasks based on these results. Performance metrics include inference latency, throughput, and accuracy.

[0041] The computing power management module allocates resources to heterogeneous computing hardware resources based on task scheduling results. These heterogeneous computing hardware resources include a central processing unit (CPU), a graphics processing unit (GPU), and a tensor processing unit (TPU).

[0042] Central Processing Unit (CPU): Suitable for handling general computing tasks and computing tasks with low parallelism, such as data preprocessing and control logic.

[0043] Graphics Processing Unit (GPU): Suitable for large-scale parallel computing tasks, such as matrix operations and convolution operations in deep learning.

[0044] Tensor Processing Unit (TPU): Hardware designed specifically for deep learning tasks, offering higher efficiency than graphics processing units (GPUs).

[0045] The monitoring module performs multi-dimensional monitoring of the platform, generating operational status monitoring indicators. Based on artificial intelligence algorithms, the module establishes a predictive model, including automatic alarms and adaptive optimization strategies. Automatic alarms are triggered when operational status monitoring indicators exceed thresholds. The predictive model forecasts future trends in computing resource utilization and service quality indicators, thereby improving platform stability and service quality. The predictive model forecasts future computing resource requirements based on historical and real-time monitoring data, including processor computing power, graphics processing unit (GPU) computing power, and storage bandwidth requirements. Operational status monitoring indicators include computing resource utilization and service quality indicators. During operation, the monitoring module collects real-time monitoring data and stores it as historical monitoring data. Multi-dimensional monitoring includes inference response latency, system throughput, and task completion rate. Thresholds are determined based on the statistical distribution of historical monitoring data, including the mean and variance of computing resource utilization and the fluctuation range of service quality indicators. When the value of real-time monitoring data exceeds the threshold, the monitoring module triggers an automatic alarm.

[0046] Example 2 This embodiment provides powerful decision support for the data management module based on an automated optimization mechanism using performance scoring.

[0047] 1. Data optimization Data compression: Compressing data to reduce storage space.

[0048] Index creation: Create indexes for data tables or datasets to improve query and access efficiency.

[0049] Data partitioning and sharding: Partitioning large-scale data to optimize query and access speed.

[0050] 2. Preprocessing Recommendations and Performance Scoring To ensure efficient data processing, the system uses a performance scoring function for data preprocessing recommendations. The formula for calculating the performance scoring function is as follows: .

[0051] in, This represents the preprocessing workflow score, which is dimensionless. This indicates the preprocessing time, in seconds. This indicates resource utilization rate, expressed as a percentage. and By mapping the normalized coefficients to a space of the same dimensions These are weighting coefficients, dimensionless. The weighting coefficients are dimensionless and satisfy the following conditions: .

[0052] 2.1 Assumptions about the data source Suppose that data is accessed from the following three data sources: Data source graphics processor 1: External business database.

[0053] Preprocessing time =12 seconds.

[0054] resource utilization =40%.

[0055] Data source 2: Data lake.

[0056] Preprocessing time =20 seconds.

[0057] resource utilization =60%.

[0058] Real-time data stream.

[0059] Preprocessing time =5 seconds.

[0060] resource utilization =30%.

[0061] 2.2 Calculate the preprocessing score According to the formula Set the weighting coefficients: =0.6.

[0062] =0.4.

[0063] Calculate the inverse function of preprocessing time and resource utilization, and substitute it into the scoring formula.

[0064] 2.2.1 Data Source 1 Calculate the inverse function: .

[0065] .

[0066] Substitute into the formula: .

[0067] 2.2.2 Data Source 2 Calculate the inverse function: .

[0068] .

[0069] Substitute into the formula: .

[0070] 2.2.3 Data Source 3 Calculate the inverse function: .

[0071] .

[0072] Substitute into the formula: .

[0073] 3. Scoring Results The preprocessing scores for the three data sources are as follows: Data source 1: =0.06.

[0074] Data source 2: =0.03667.

[0075] Data source 3: =0.13332.

[0076] 4. Optimization and Decision Making Data source 3 has the highest preprocessing score. =0.13332, processing real-time data streams with limited resources.

[0077] Data source 2 received the lowest preprocessing score. =0.03667, indicating that the data processing flow needs further optimization to reduce resource consumption and processing time.

[0078] Example 3 This embodiment, based on the monitoring module, provides strong support for the stable operation of the platform through real-time monitoring and historical data analysis, combined with adaptive optimization and predictive models.

[0079] 1. Operational status monitoring indicators The monitoring module monitors the platform's operational status through the following dimensions: Computing resource utilization: This includes the usage of processor, graphics processor, and storage bandwidth.

[0080] Service quality metrics include inference response latency, system throughput, and task completion rate.

[0081] Predicting computing resource requirements: Based on historical and real-time data, predict future resource requirements to help the platform provide early warnings and adjust resource allocation.

[0082] 2. Threshold setting and automatic alarm mechanism Thresholds are a core component of the monitoring module, used to trigger alarms and control adaptive optimization. When a monitored metric exceeds a set threshold, the system will automatically trigger an alarm and take optimization measures. Thresholds are typically calculated based on the statistical distribution of historical monitoring data.

[0083] 2.1 Threshold Calculation Method Threshold settings are typically based on statistical analysis of historical monitoring data and involve the following steps: Calculate the statistical distribution of resource utilization rate: Mean: The average utilization rate of all computing resources in historical data.

[0084] Variance: Calculates the variance of resource utilization rate, measuring the magnitude of fluctuations in resource use.

[0085] Statistical distribution of service quality indicators: Mean: The average of inference response latency, throughput, and task completion rate in historical data.

[0086] Fluctuation range: The fluctuation range of service quality indicators, usually expressed by variance or standard deviation.

[0087] Based on statistical distribution, the system can set thresholds to trigger alarms. For example, if the utilization rate of computing resources exceeds the mean plus one standard deviation, an alarm will be triggered.

[0088] 2.2 Substituting specific values Assume the monitoring module uses the following historical data to calculate the threshold: Calculate resource utilization rate: Central Processing Unit (CPU) Utilization: The historical mean is 70% for the GPU and the standard deviation is 5% for the GPU.

[0089] Graphics processor utilization: The historical mean is 60% for graphics processors, and the standard deviation is 8% for graphics processors.

[0090] Set threshold: For the central processing unit: If the current central processing unit utilization exceeds 70% + 5% = 75% of the graphics processing unit, an alarm will be triggered.

[0091] For the graphics processor: If the current graphics processor utilization exceeds 60% + 8% = 68%, an alarm will be triggered.

[0092] Inference response delay: The mean inference response latency in historical data is 200ms for the graphics processor and the standard deviation is 50ms for the graphics processor.

[0093] Set threshold: If the current inference response delay exceeds 250ms, trigger an alarm.

[0094] System throughput: The historical data has a mean of 1000 requests / second for the graphics processor and a standard deviation of 200 requests / second for the graphics processor.

[0095] Set threshold: If the current throughput is below 800 requests / second, trigger an alarm.

[0096] Task completion rate: The mean of the historical data is 98% for graphics processors, and the standard deviation is 2% for graphics processors.

[0097] Set threshold: If the current task completion rate is below 96%, trigger an alarm.

[0098] 2.3 Implementation of Automatic Alarm When real-time monitoring data exceeds the threshold, the monitoring module will trigger an alarm. The following are the specific alarm handling methods: Real-time alerts: Notify relevant personnel via SMS, email, or platform push notifications.

[0099] Alarm Classification: The system sets different alarm levels based on the severity of the alarm.

[0100] Alarm Response: When an alarm is triggered, the platform administrator will manually or automatically adjust the resource configuration based on the alarm information.

[0101] For example, when the CPU utilization exceeds 75%, the system sends an emergency alert to the administrator, suggesting resource expansion or load balancing.

[0102] 3. Predictive Models and Resource Demand Forecasting Predictive models, based on historical and real-time data, forecast future computing resource requirements, facilitating proactive resource planning and dynamic optimization. The predicted computing resource requirements include: Processor computing power requirements: Predicting the usage of processors such as central processing units and graphics processing units in the near future.

[0103] Storage bandwidth requirements: Predict the access needs of storage resources to ensure that data storage does not become a bottleneck.

[0104] For example, by analyzing historical data, the system can predict that the demand for the central processing unit may increase by 20% in the next hour, and can prepare more computing resources in advance to cope with the increased load.

[0105] Specific prediction examples: Assuming we use historical data and real-time monitoring data to predict resource requirements for the next hour: CPU Demand: Historical data shows that CPU utilization reached 90% during peak hours over the past week. The system predicts that CPU demand may increase by 20% in the next hour.

[0106] Graphics Processing Unit (GPU) Demand: Historical data shows that GPU usage in deep learning tasks reaches 70% during training. Predictive models calculate that GPU demand will increase by 10% in the next hour.

[0107] Storage bandwidth requirements: By analyzing historical I / O load, the predictive model forecasts that storage bandwidth requirements will increase by 15% in the next hour.

[0108] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A model repository platform and model service platform for artificial intelligence, characterized in that, include: The data management module is used to access and process cross-data source heterogeneous data from external data sources, and to preprocess the cross-data source heterogeneous data to obtain preprocessed heterogeneous data. The model repository and service module receives the preprocessed heterogeneous data, extracts metadata information of the artificial intelligence model, and provides storage, version control, and one-stop management of the artificial intelligence model. The inference task scheduling and optimization module extracts performance metrics from the metadata information, generates task scheduling results based on the performance metrics, and executes inference tasks based on the task scheduling results. A computing power management module, which allocates resources to the heterogeneous computing hardware resources based on the task scheduling results; The monitoring module performs multi-dimensional monitoring of the platform and generates operational status monitoring indicators. The monitoring module establishes a prediction model based on artificial intelligence algorithms, which include automatic alarm and adaptive optimization strategies. The automatic alarm is triggered when the operational status monitoring indicators exceed the threshold.

2. The artificial intelligence model repository platform and model service platform according to claim 1, characterized in that: The external data sources include external business databases, data warehouses, data lakes, and real-time data streaming systems. The heterogeneous data across data sources includes relational database data, non-relational database data, data warehouse data, data lake data, and real-time data streams. The preprocessing includes cleaning, optimization, and preprocessing recommendations.

3. The artificial intelligence model repository platform and model service platform according to claim 2, characterized in that: The preprocessing recommendation employs a performance scoring function to comprehensively evaluate the heterogeneous data from different data sources. The formula for calculating the performance scoring function is as follows: , in, Indicates the preprocessing process score. Indicates the preprocessing time. Indicates resource utilization rate. These are the weighting coefficients. For the weighting coefficients, satisfying .

4. The artificial intelligence model repository platform and model service platform according to claim 1, characterized in that: The metadata information includes model version information, training parameters and performance metrics. The model repository and service module provides a unified calling interface, which supports loading and calling the artificial intelligence model via RPC.

5. The artificial intelligence model repository platform and model service platform according to claim 1, characterized in that: The performance metrics include inference latency, throughput, and accuracy.

6. The artificial intelligence model repository platform and model service platform according to claim 1, characterized in that: The heterogeneous computing hardware resources include a central processing unit, a graphics processing unit, and a tensor processing unit.

7. The artificial intelligence model repository platform and model service platform according to claim 1, characterized in that: The operational status monitoring indicators include computing resource utilization and service quality indicators. The monitoring module collects real-time monitoring data during operation and stores it as historical monitoring data. The multi-dimensional monitoring includes inference response latency, system throughput, and task completion rate. The threshold is determined based on the statistical distribution of the historical monitoring data. The statistical distribution includes the mean and variance of the computing resource utilization and the fluctuation range of the service quality indicators.

8. The artificial intelligence model repository platform and model service platform according to claim 7, characterized in that: The prediction model predicts future computing resource requirements based on the historical monitoring data and the real-time monitoring data. The computing resource requirements include processor computing power requirements, graphics processor computing power requirements, and storage bandwidth requirements.