Enterprise-oriented AI service unified management method and system combined with MCP

By unifying the registration and management of enterprise AI services, and combining AI gateways and MCP status trackers, the management and security issues of enterprises in the deployment of AI models are solved, enabling efficient and secure use and unified management of AI services.

CN121585571APending Publication Date: 2026-02-27HUAFU SECURITIES CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511544585.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-28
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Enterprises face challenges such as high technical barriers, fragmented management, security risks, uncontrollable costs, operational difficulties, and difficulty in reusing accumulated knowledge when introducing and deploying AI models. Existing tools have failed to effectively address the usage needs of non-technical personnel.

Method used

The system standardizes AI services by uniformly registering private deployment models, third-party API models, and SaaS models on servers. It combines AI gateways for identity authentication and access control, utilizes load balancing and security gateways to ensure communication security, and monitors abnormal events and records execution logs in real time through MCP status trackers.

Benefits of technology

It enables centralized management, enhanced security, and improved ease of use of AI services, lowers the barrier to entry, improves management efficiency and data security, and supports enterprise-level unified auditing and compliance requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121585571A_ABST
    Figure CN121585571A_ABST
Patent Text Reader

Abstract

The invention provides an enterprise-oriented MCP-combined AI service unified management method and system in the technical field of enterprise information technology management and artificial intelligence application. The method comprises the steps that S1, a server registers all models as AI services in a unified mode; s2, the AI gateway obtains an AI task request sent by an employee terminal to perform identity authentication, and matches authority information; step S3, interacting with the server through the load balancing and security gateway, querying AI services of which authority information can be called, and pushing the AI services to a visual interface of an employee terminal; s4, arranging each AI service by the visual interface based on the triggered drag signal and click signal to generate an AI task, and sending the AI task to a load balancing and security gateway for execution; s5, in the AI task execution process, an abnormal event is captured to give an alarm; and S6, recording an execution log in real time. The method has the advantages that the centralization, the safety and the usability of AI service management are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of enterprise information technology management and artificial intelligence application, and particularly discloses an AI service unified management method and system for enterprises combined with MCP. BACKGROUND

[0002] With the rapid development of artificial intelligence technology, especially the continuous enhancement of the performance of large language models (LLM), more and more enterprises hope to apply it to daily business scenarios, such as intelligent customer service, report generation, data analysis and content creation, etc. However, when enterprises introduce and deploy AI models, especially various third-party large language model services, they face a series of technical and management challenges, and are generally in a double dilemma of "employee use difficulty and enterprise management lack".

[0003] From the perspective of employee use, the existing technical solutions have high technical threshold and operational complexity. Currently, if an employee needs to call different large model services (such as Tongyi Qianwen, KIMI, Wenxin Yiyang, DeepSeek, etc.), he or she usually needs to complete account registration, API key application and management on each third-party platform, and bear the corresponding fees or perform a dispersed reimbursement process. In addition, the employee also needs to deploy and configure complex calling tools or development frameworks (such as LangChain, MCP client, etc.) by himself or herself. Such operations are not only tedious, but also require the employee to have certain programming ability and technical understanding. For non-technical business personnel, it is extremely difficult to continuously track the rapid evolution of model capabilities, select the optimal model and master its calling method, resulting in low AI technology popularization rate in enterprises and difficult to guarantee application effect.

[0004] From the perspective of enterprise management, the existing technology lacks unified control ability for enterprise-level AI service use, and there are significant hidden dangers in safety, cost, operation and maintenance and compliance, which are specifically manifested as: 1. Dispersed account and permission management: employees use personal accounts and API keys to operate independently on each platform, and it is difficult for enterprises to achieve unified identity authentication, access control and operation audit. When an employee leaves or changes position, the permission recovery is not timely, which can easily lead to data leakage risk.

[0005] 2. Uncontrollable calling cost: API calling fees are usually paid by employees or reimbursed in scattered manner, and it is difficult for enterprises to conduct overall budgeting, cost collection and effective control of AI resource use, and to accurately assess the input-output benefit.

[0006] 3. Security and compliance risks are prominent: sensitive business data may be transmitted to third-party services not supervised by the enterprise through the employee's personal account, and the enterprise cannot effectively audit and filter data export, which poses a risk of data leakage and violation.

[0007] 4. Lack of Operation and Maintenance Monitoring and Fault Response Mechanisms: When the model service experiences call failures, response timeouts, or abnormal results, the existing model lacks system-level real-time monitoring and alarm capabilities. Employees find it difficult to proactively detect anomalies, and the IT operations department cannot intervene in a timely manner, affecting business continuity and stability.

[0008] 5. Difficulty in knowledge accumulation and process reuse: Even if individual employees build effective AI application processes (such as "automatic generation of weekly reports" and "customer intent recognition"), due to the lack of a unified sharing and management platform, such best practices are difficult to promote and reuse within the team or enterprise, resulting in resource waste and redundant development.

[0009] While some AI workflow orchestration tools or API gateway products exist in the market for developers, these tools typically require users to be familiar with the MCP protocol, possess directed acyclic graph process design capabilities, and have specialized skills in containerized deployment. Essentially, they remain geared towards technical users and fail to fundamentally address the core enterprise needs of non-technical personnel for secure, convenient, and efficient use of large-scale models. Enterprises urgently require a system-level solution that can transform AI capabilities into manageable, out-of-the-box internal services, supporting unified scheduling, management, and auditing.

[0010] Therefore, how to provide a unified management method and system for AI services that combines MCP for enterprises, so as to improve the centralization, security and ease of use of AI service management, has become an urgent technical problem to be solved. Summary of the Invention

[0011] The technical problem to be solved by this invention is to provide a unified management method and system for AI services combined with MCP for enterprises, so as to improve the centralization, security and ease of use of AI service management.

[0012] In a first aspect, the present invention provides a unified management method for AI services combined with MCP for enterprises, comprising the following steps: Step S1: The server registers all private deployment models, third-party API models, and SaaS big models as standardized AI services. Step S2: The AI ​​gateway obtains the AI ​​task request sent by the employee terminal, performs identity authentication based on the AI ​​task request, and matches the corresponding permission information. Step S3: The AI ​​gateway interacts with the server through the load balancer and security gateway to query the AI ​​services that can be called based on the permission information, and pushes them to the visual interface of the employee terminal for display. Step S4, the visualization interface arranges each AI service based on the triggered drag signal and click signal to generate an AI task, and sends the AI task to a load balancing and security gateway through an AI gateway to call the corresponding AI service to execute the AI task; Step S5, during the execution of the AI task, an exception event is automatically captured through an MCP state tracker, and an alarm is given based on the captured exception event; Step S6, the server records an execution log of the execution of the AI task in real time, and stores the execution log.

[0013] Further, the step S1 further comprises: the server configures model metadata of each AI service, including at least a model name, a capability label, a permission label, a department, whether it involves a secret, an API Key, and a calling address; The step S2 is specifically: The AI gateway obtains an AI task request sent by an employee terminal through a visualization interface, carries a department and a position role, performs identity authentication including at least identity legality, permission compliance, and quota availability based on the AI task request, and matches corresponding permission information after authentication. The step S3 is specifically: The AI gateway interacts with the server based on the TLS protocol through the load balancing and security gateway, queries the AI service that can be called by the permission information, and pushes it to the visualization interface of the employee terminal in the form of a service card for display.

[0014] Further, the step S4 is specifically: The visualization interface arranges each AI service through a DAG editor based on the triggered drag signal and click signal, including at least combination, sequence adjustment, and connection relationship adjustment to generate an AI process, converts the AI process into an AI task, and configures task parameters of the AI task including at least a timeout time, a maximum number of concurrent tasks, a number of retries, and a log path; The AI task is sent to the load balancing and security gateway through the AI gateway to call the corresponding AI service to execute the AI task.

[0015] Further, the step S5 is specifically: During the execution of the AI task, the MCP state tracker tracks exception events including at least timeout exceptions, resource exceptions, communication exceptions, service exceptions, and logic exceptions, captures the triggered exception events, pushes the captured exception events to a management terminal in real time through WeChat Enterprise, email, SMS, or telephone to give an alarm, and performs a recovery operation based on a corresponding recovery measure matched by a preset recovery mechanism; When the MCP state tracker detects that the number of execution events of the computing node exceeds a preset threshold through a timer, it triggers the timeout exception event. When the MCP status tracker detects through probes that a container on a compute node reports a memory overflow, GPU memory usage >90%, or CPU usage >90%, it triggers the abnormal event of the aforementioned resource anomaly. When the MCP status tracker detects a TCP connection interruption or connection timeout, it triggers the abnormal event of the communication anomaly. When the MCP status tracker detects that the response status code of the AI ​​service is an error code, it triggers an abnormal event for the service. When the MCP state tracker determines that the output result of the AI ​​service is abnormal according to the preset output verification rules, it triggers the abnormal event of the logical abnormality.

[0016] Furthermore, step S6 specifically includes: The server records in real time the execution logs of the AI ​​task, including at least employee identity, abnormal events, invocation behavior, invocation process, input and output content, and execution time. The execution logs are stored in the TiSQL database and pushed synchronously to the enterprise monitoring platform. The server reads the stored execution logs at a preset period to generate an AI service usage report.

[0017] Secondly, this invention provides a unified AI service management system for enterprises that integrates MCP, including the following modules: The AI ​​service registration module is used by the server to uniformly register various private deployment models, third-party API models, and SaaS large models as standardized AI services. The identity authentication module is used by the AI ​​gateway to obtain AI task requests sent by employee terminals, perform identity authentication based on the AI ​​task requests, and match the corresponding permission information. The AI ​​service display module is used by the AI ​​gateway to interact with the server through load balancing and security gateway, query the AI ​​services that can be called with the permission information, and push them to the visual interface of the employee terminal for display. The AI ​​task execution module is used by the visualization interface to orchestrate the AI ​​services based on the triggered drag and click signals to generate AI tasks, and send the AI ​​tasks to the load balancer and security gateway through the AI ​​gateway to call the corresponding AI service to execute the AI ​​tasks. An abnormal event capture module is used to automatically capture abnormal events through the MCP status tracker during the execution of the AI ​​task and to issue alarms based on the captured abnormal events. An execution log management module is configured to record an execution log of the AI task in real time and store the execution log.

[0018] Further, the AI service registration module is further configured to configure the AI services with model metadata including at least a model name, a capability label, a permission label, a department, a secret-related indication, an API Key, and a calling address. The identity authentication module is specifically configured to: The AI gateway acquires an AI task request sent by the employee terminal through a visual interface and carrying a department and a role, performs identity authentication including at least identity legality, permission compliance, and quota availability based on the AI task request, and matches corresponding permission information after authentication. The AI service display module is specifically configured to: The AI gateway interacts with the server based on a TLS protocol through a load balancing and security gateway, queries AI services that can be called by the permission information, and pushes the AI services to the visual interface of the employee terminal in the form of service cards for display.

[0019] Further, the AI task execution module is specifically configured to: The visual interface generates an AI flow by arranging the AI services including at least combination, sequence adjustment, and connection relationship adjustment through a DAG editor based on triggered drag signals and click signals, converts the AI flow into an AI task, and configures task parameters of the AI task including at least a timeout time, a maximum concurrency number, a retry number, and a log path. The AI task is sent to the load balancing and security gateway through the AI gateway to call corresponding AI services to execute the AI task.

[0020] Further, the abnormal event capture module is specifically configured to: During the execution of the AI task, the MCP state tracker tracks abnormal events including at least timeout exceptions, resource exceptions, communication exceptions, service exceptions, and logic exceptions, captures the triggered abnormal events, pushes the captured abnormal events to a management terminal in real time through WeChat for Enterprise, email, short message, or telephone to perform alarm, and performs recovery operations based on a preset recovery mechanism and matching corresponding recovery measures. When the MCP state tracker detects that a computing node execution event exceeds a preset threshold through a timer, the timeout exception is triggered. When the MCP state tracker detects that a computing node's container reports memory overflow, GPU memory usage > 90%, or CPU usage > 90% through a probe, the resource exception is triggered. When the MCP state tracker monitors that the TCP connection is interrupted or the connection is timed out, triggering the abnormal event of the communication exception; When the MCP state tracker monitors that the response status code of the AI service is an error code, triggering the abnormal event of the service exception; When the MCP state tracker judges that the output result of the AI service is abnormal through a preset output check rule, triggering the abnormal event of the logic exception.

[0021] Further, the execution log management module is specifically used for: The server records at least the employee identity, the abnormal event, the calling behavior, the calling process, the input and output content, and the execution time consumption of the AI task execution in real time, stores the execution log into a TiSQL database, synchronously pushes the execution log to an enterprise supervision platform, and generates an AI service use report based on a preset period of reading the stored execution log.

[0022] The application has the following advantages: 1. The server registers each private deployment model, third-party API model, and SaaS large model as a standardized AI service; the AI gateway obtains an AI task request sent by an employee terminal, performs identity authentication based on the AI task request, matches corresponding permission information, then interacts with the server through a load balancing and security gateway, queries the AI service that can be called by the permission information, and pushes the AI service to a visual interface of the employee terminal for display; the visual interface arranges each AI service based on a triggered drag signal and a click signal to generate an AI task, sends the AI task to the load balancing and security gateway through the AI gateway to call the corresponding AI service to execute the AI task; the MCP state tracker automatically captures an abnormal event during the AI task execution, and alarms based on the captured abnormal event; the server records an execution log of the AI task execution in real time, and stores the execution log; that is, by uniformly registering various AI models as standard services and controlling and managing through the AI gateway as a core hub, centralized management of service access, identity authentication, and permission allocation is realized, fundamentally solving the problems of account dispersion and uncontrollable cost; on this basis, the load balancing and security gateway guarantee communication security, the full-link execution log record, and the real-time monitoring and alarm mechanism of the MCP state tracker are combined to build a deep security and reliability guarantee system covering data flow, service state, and business logic; finally, the above centralized control and security base provide the front-end employee terminal with visual service display and low-code drag arrangement capabilities, greatly reducing the use threshold of the AI service, and ultimately greatly improving the centralization, security, and ease of use of AI service management.

[0023] 2、By registering private deployment models, third-party API models, and SaaS large models as standardized AI services and configuring model metadata (such as model name, capability label, etc.), this standardized registration avoids fragmentation of AI services within the enterprise, reduces compatibility issues between different models, and improves management efficiency.

[0024] 3、The AI gateway performs multi-dimensional identity authentication (including identity legitimacy, permission compliance, and quota availability) based on employee terminal requests and securely interacts with the server through the TLS protocol, ensuring that only authorized users can access specific AI services and preventing unauthorized access and data leakage. The advantage is that it binds permissions with departments and roles, enabling fine-grained access control and improving enterprise data security.

[0025] 4、By introducing load balancing and security gateways combined with the TLS protocol, the high availability and stability of AI service calls are ensured. Load balancing can distribute request pressure and avoid single-point failures, while security gateways provide encrypted transmission and reduce network risks. This advantage not only improves AI task execution efficiency but also reduces system downtime probability.

[0026] 5、Through a visual interface that supports drag-and-drop and click signals, the DAG editor is used to orchestrate AI services (such as combination and sequence adjustment) and configure task parameters (such as timeout time and retry count). This allows non-technical employees to easily create complex AI processes, reduces the use threshold, and improves work efficiency.

[0027] 6、The MCP state tracker automatically captures various abnormal events (such as timeout, resource, and communication exceptions) and sends real-time alerts (through WeChat, email, etc.), while triggering recovery mechanisms to achieve proactive monitoring and rapid response, reducing AI task interruption time, and improving system robustness.

[0028] 7、By recording execution logs (including employee identity, abnormal events, input and output, etc.) in real time and storing them in TiSQL databases, usage reports are generated synchronously, providing complete audit trails for the enterprise, facilitating compliance checks and performance optimization. The advantage is that it enhances transparency and traceability, meeting the needs of enterprise regulation.

[0029] 8、By unifying private deployment, third-party API and SaaS large model as standardized AI services, the efficient integration and management of intra-enterprise AI resources are realized, and the management uniformity and compatibility are significantly improved; at the same time, multi-dimensional identity authentication and permission control are carried out with the help of AI gateway, and load balancing and security gateway are combined to ensure system high availability and data transmission security, thereby enhancing the overall security and reliability; in addition, the visual interface supports drag-and-drop task arrangement, which reduces the user operation threshold and improves the flexibility of task configuration, and the real-time exception monitoring and automatic alarm mechanism of the MCP state tracker can quickly respond to faults and reduce interruption risks; finally, by recording execution logs in detail and generating usage reports, the scheme provides complete audit tracking capabilities, supports enterprise compliance supervision and performance optimization, thereby showing comprehensive advantages in efficiency, security, user experience and operation and maintenance. BRIEF DESCRIPTION OF DRAWINGS

[0030] The application will be further described below with reference to the accompanying drawings and embodiments.

[0031] Fig. 1 is a flowchart of an AI service unified management method for enterprises combined with MCP according to the application.

[0032] Fig. 2 is a structural schematic diagram of an AI service unified management system for enterprises combined with MCP according to the application. DETAILED DESCRIPTION

[0033] The technical scheme in the embodiments of the present application has the following general idea: by uniformly registering various AI models as standard services and controlling them through the AI gateway as the core hub, centralized management of service access, identity authentication and permission allocation is realized, fundamentally solving the problems of account dispersion and uncontrollable cost; on this basis, the communication security is ensured through load balancing and security gateway, and the real-time monitoring and alarm mechanism of the MCP state tracker is combined with the whole-link execution log recording to build a deep security and reliability guarantee system covering data flow, service state and business logic; finally, the above centralized control and security base provide the front-end employee terminal with visual service display and low-code drag-and-drop arrangement capabilities, greatly reducing the use threshold of AI services, and thereby improving the centralization, security and ease of use of AI service management.

[0034] Please refer to Figs. 1-2 The preferred embodiment of an AI service unified management method for enterprises combined with MCP according to the application is shown in the figure, which includes the following steps: Step S1, the server registers each private deployment model (such as QWQ, Gemma3), third-party API model (such as Hengsheng Juyuan), and SaaS large model (such as KIMI, Wenxin Yanyan, DeepSeek) as standardized AI services; employees do not need to apply for individual accounts or configure keys, and all calls are completed through the enterprise unified portal (AI gateway); Step S2, the AI gateway obtains the AI task request sent by the employee terminal, performs identity authentication based on the AI task request, and matches the corresponding permission information; By building a unified AI gateway, private deployment models, third-party APIs, and SaaS large models are encapsulated as standardized AI services within the enterprise. Employees do not need to register accounts, apply for API keys, or manage billing accounts on their own. They only need to directly call AI capabilities through a visual interface, completely eliminating the technical background barrier in the traditional mode, significantly reducing the use threshold, and enabling non-technical personnel to easily apply large models, thereby improving the overall AI usage efficiency and popularity of the enterprise.

[0035] Step S3, the AI gateway interacts with the server through load balancing and security gateway, queries the AI services that can be called by the permission information, and pushes them to the visual interface of the employee terminal for display; Step S4, the visual interface arranges each AI service based on the triggered drag signal and click signal to generate an AI task, sends the AI task to the load balancing and security gateway through the AI gateway, and calls the corresponding AI service to execute the AI task; The drag-and-drop arrangement capability is provided, allowing business personnel to quickly build reusable AI processes by combining nodes (such as "intent recognition → large model → email sending"). The system automatically completes MCP protocol conversion and task scheduling without the need for code writing or deployment of complex tools. This not only reduces technical dependence but also enables business departments to innovate independently, such as quickly implementing "weekly report generation" or "customer complaint classification" scenarios, shortening the AI application development cycle, and promoting the agility of enterprise digital transformation.

[0036] Step S5, during the execution of the AI task, the MCP state tracker automatically captures abnormal events based on the captured abnormal events to perform alarm; Through the built-in MCP state tracker, model call timeouts, service unavailability, or logic exceptions can be detected in real time, and alarms can be pushed to employees and IT support personnel within seconds. This supports manual intervention or automatic switching to backup computing nodes. This "use-monitor-alarm-recovery" closed-loop mechanism significantly improves the availability and stability of AI services, avoids the problem of no awareness of faults in the traditional mode, and reduces the risk of business interruption.

[0037] Step S6, the server records the execution log of the AI task execution in real time, and stores the execution log.

[0038] The application adopts a three-layer network architecture, and is not only for security isolation, but also constructs an AI service unified portal which is transparent to enterprise employees and controllable to enterprise IT, specifically including an external network area, a DMZ area (buffer area), and an internal network area; the external network area is the only access portal for employee terminals, integrates an AI gateway, employees do not need to register any large model account, and all requests enter the system after unified identity authentication through the AI gateway; the DMZ area is deployed with a load balancing and security gateway, is responsible for traffic scheduling and safe proxy calling of external APIs (such as KIMI and Tongyi Qianwen), and ensures that the third-party model calling behavior is controlled and auditable; the internal network area is the server, and is deployed with AI services, a visual arrangement platform, an MCP scheduling component, and a TiSQL database; all large models (including private models and third-party APIs) are uniformly registered as standardized services, and employees do not need to care about the underlying deployment details when calling. The three-layer network architecture makes employees only need to click templates such as DeepSeek and Tongyi Qianwen on the visual interface, and can automatically complete the whole process of identity authentication, model calling, and result returning, and completely eliminates complex operations such as account application, API Key management, and MCP tool deployment.

[0039] The step S1 further includes that the server configures model metadata of each AI service, at least including a model name, an ability label, a permission label, a department, whether it involves a secret, an API Key, and a calling address; The step S2 is specifically: The AI gateway (integrating an enterprise SSO / OAuth 2.0 authentication system) obtains an AI task request sent by an employee terminal through a visual interface (such as an enterprise WeChat workbench), carries a department and a position role, performs identity authentication based on the AI task request, at least including identity legality (whether it is an enterprise employee), permission compliance (whether it is authorized to use the model, such as “only calling Tongyi Qianwen, and not calling overseas models”), and quota availability (whether the number of calling times in the day is over the limit), and matches the corresponding permission information after the authentication is passed; and records complete audit logs, ensures that sensitive data is not leaked, and the calling behavior is traceable; The AI gateway is not only a traffic portal, but also an enterprise AI unified outlet, which automatically injects employee identity information, completes authentication, uniformly charges all model calling fees (without employees paying in advance), and records and filters the logs of third-party API calling.

[0040] By integrating enterprise SSO / OAuth2.0 authentication system, fine permission control and quota management according to departments and roles are realized, all AI calls are subject to unified authentication, and complete audit logs are recorded, effectively preventing the risk of sensitive data leakage through uncontrolled API, ensuring that the calling behavior is traceable, meeting the requirements of enterprise data security and compliance, and solving the pain point of traditional scattered use that cannot be audited uniformly.

[0041] In specific implementation, all process control operations (such as pause, retry, and terminate) need to pass through enterprise SSO identity authentication; ordinary employees can only operate the tasks they initiate; IT administrators can intervene in abnormal processes across departments.

[0042] The step S3 is specifically: The AI gateway interacts with the server based on the TLS protocol through load balancing and security gateway, queries the AI services that can be called according to the permission information, and pushes the AI services to the employee terminal in real time in the form of service cards for display on the visual interface.

[0043] The step S4 is specifically: The visual interface arranges the AI services through a DAG editor based on the triggered drag signal and click signal to generate AI processes (such as "meeting reservation", "security financial report generation", and "customer complaint classification") at least including combination, sequence adjustment, and connection relationship adjustment, converts the AI processes into AI tasks (users do not need to write code or understand MCP protocols), and configures task parameters of the AI tasks at least including timeout time, maximum concurrency number, retry number, and log path (for example, setting "20 times of calling large models per day" for the "finance department" to prevent resource abuse, task parameters are in JSON / YAML format, support version control and gray release, and facilitate IT centralized management); The AI tasks are sent to the load balancing and security gateway through the AI gateway to call the corresponding AI services to execute the AI tasks.

[0044] DAG (Directed Acyclic Graph) is a directed acyclic graph used to visually represent the dependencies and execution order of nodes in AI workflow.

[0045] The step S5 is specifically: During the AI task execution process, the MCP state tracker tracks abnormal events including at least timeout exceptions, resource exceptions, communication exceptions, service exceptions, and logic exceptions, captures the triggered abnormal events, pushes the captured abnormal events to a management terminal in real time through WeChat, email, SMS, or telephone for alarm, and performs recovery operations based on a preset recovery mechanism matching corresponding recovery measures; manual intervention or automatic switching of backup computing nodes is supported to form a "use-monitor-alarm-recovery" closed loop; once the MCP state tracker detects an exception (such as a model returning an empty result, a timeout, or a service being unavailable), it immediately generates an alarm event and pushes it to the applicant and IT support personnel in real time to achieve "second-level fault perception and minute-level response".

[0046] When the MCP state tracker detects that the computing node execution event exceeds the preset threshold through the timer, the timeout exception event is triggered; When the MCP state tracker detects that the computing node's container reports memory overflow (OOM), GPU memory usage > 90%, or CPU usage > 90% through the probe, the resource exception event is triggered; When the MCP state tracker detects that the TCP connection is interrupted or timed out, the communication exception event is triggered; When the MCP state tracker detects that the AI service response status code is an error code (such as HTTP 5xx), the service exception event is triggered; When the MCP state tracker determines that the AI service output result is abnormal through a preset output verification rule (such as the node returning an empty result, missing a key field, or containing a large model that cannot be returned), the logic exception event is triggered.

[0047] MCP (Model Context Protocol) is a model context protocol used to standardize the protocol for scheduling AI service processes and is transparent to employees.

[0048] The step S6 specifically includes: The server records the execution logs of the AI task execution in real time, including at least employee identity, abnormal events, calling behavior, calling process, input and output content, and execution time consumption, and stores the execution logs in a TiSQL database and synchronously pushes them to an enterprise supervision platform, and generates an AI service usage report based on a preset period (such as a monthly "department AI usage report", "application usage report", and "expense allocation list").

[0049] The TiSQL database records all calling behavior, process logs, and abnormal events, supports enterprise auditing of AI usage by department / personnel, and provides a data basis for alarm pushing.

[0050] With the unified billing function of the AI gateway and the full-link log recording of the TiSQL database, AI usage reports and cost allocation lists can be generated by department or project, enabling enterprises to clearly monitor AI investment and output; this solves the problem of dispersed costs and difficulty in statistics, realizes controllable costs and measurable effects, provides data support for enterprise budget management and resource optimization, and improves the return on investment.

[0051] A preferred embodiment of an AI service unified management system combined with MCP for enterprises, comprising the following modules: An AI service registration module is used for the server to uniformly register private deployment models (such as QWQ and Gemma3), third-party API models (such as the Everbright Juyuan Investment Consultant Assistant), and SaaS large models (such as KIMI, Wenxin Yanyan, and DeepSeek) as standardized AI services; employees do not need to apply for separate accounts or configure keys, and all calls are completed through the enterprise unified portal (AI gateway); An identity authentication module is used for the AI gateway to obtain an AI task request sent by an employee terminal, perform identity authentication based on the AI task request, and match corresponding permission information; By constructing a unified AI gateway, private deployment models, third-party APIs, and SaaS large models are encapsulated as standardized AI services within the enterprise; employees do not need to register accounts, apply for API keys, or manage billing accounts on multiple platforms; they only need to directly call AI capabilities through a visual interface, completely eliminating the technical background barrier in the traditional mode, greatly reducing the use threshold, and enabling non-technical personnel to easily apply large models, thereby improving the overall AI usage efficiency and popularity of the enterprise.

[0052] An AI service display module is used for the AI gateway to interact with the server through load balancing and a security gateway, query AI services that can be called by the permission information, and push to the visual interface of the employee terminal for display; An AI task execution module is used for the visual interface to arrange AI services based on triggered drag signals and click signals to generate AI tasks, and send the AI tasks to the load balancing and security gateway through the AI gateway to call corresponding AI services to execute AI tasks; Provide drag-and-drop orchestration capabilities, allowing business personnel to quickly build reusable AI processes by combining nodes (such as "intent recognition -> large model -> email sending"), and the system automatically completes MCP protocol conversion and task scheduling, without the need to write code or deploy complex tools, which not only reduces technical dependence, but also enables business departments to innovate independently, such as quickly implementing "weekly report generation" or "customer complaint classification" scenarios, shortening the AI application development cycle, and promoting the agility of enterprise digital transformation.

[0053] An abnormal event capture module is configured to automatically capture abnormal events during the AI task execution process through the MCP state tracker, and to perform an alarm based on the captured abnormal events; By using the built-in MCP state tracker, model call timeouts, service unavailability, or logic abnormalities can be detected in real time, and alarms can be pushed to employees and IT support personnel in seconds, supporting manual intervention or automatic switching to backup computing nodes. This "use-monitor-alarm-recovery" closed-loop mechanism significantly improves the availability and stability of AI services, avoids the problem of no awareness of faults in the traditional mode, and reduces the risk of business interruption.

[0054] An execution log management module is configured to record and store execution logs of the AI task execution in real time.

[0055] The application adopts a three-layer network architecture, and not only for security isolation, but also to build an AI service unified portal that is transparent to enterprise employees and controllable to enterprise IT, specifically including an external network area, a DMZ area (buffer area), and an internal network area. The external network area serves as the only access portal for employee terminals, integrates an AI gateway, and employees do not need to register any large model account, all requests enter the system after unified identity authentication through the AI gateway; the DMZ area is deployed with a load balancing and security gateway, responsible for traffic scheduling and secure proxy calling of external APIs (such as KIMI and Tongyi Qianwen), ensuring that third-party model calling behavior is controlled and auditable; the internal network area is the server, which deploys AI services, a visual orchestration platform, MCP scheduling components, and a TiSQL database, etc.; all large models (including private models and third-party APIs) are registered as standardized services, and employees do not need to care about the underlying deployment details when calling. The three-layer network architecture enables employees to simply click on templates such as "DeepSeek" and "Tongyi Qianwen" on the visual interface, to automatically complete the entire process of identity authentication, model calling, and result returning, completely eliminating complex operations such as account application, API Key management, and MCP tool deployment.

[0056] The AI service registration module is further configured to configure the server with model metadata of each AI service, including at least model name, capability label, permission label, department, whether it involves secrets, API Key, and calling address; The identity authentication module is specifically used for: The AI gateway (integrated enterprise SSO / OAuth 2.0 authentication system) obtains an AI task request sent by an employee terminal through a visual interface (such as an enterprise WeChat workbench), the AI task request carrying a department and a position role, performs identity authentication based on the AI task request, and the identity authentication at least includes identity legality (whether it is an enterprise employee), permission compliance (whether it is authorized to use the model, such as “only calling general questions, not calling overseas models”), and quota availability (whether the number of calls per day is over the limit), and matches the corresponding permission information after the authentication is passed; and records complete audit logs to ensure that sensitive data is not leaked and the calling behavior is traceable. The AI gateway is not only a traffic entrance, but also a unified outlet of enterprise AI, which automatically injects employee identity information and completes authentication; unifies the charging of all model calling fees (without employees paying in advance); and records logs and performs security filtering on third-party API calls.

[0057] Through the integrated enterprise SSO / OAuth2.0 authentication system, fine permission control and calling quota management are realized according to departments and roles, all AI calls are subjected to unified authentication, and complete audit logs are recorded, which effectively prevents the risk of sensitive data leakage through uncontrolled API, ensures that the calling behavior is traceable, meets the requirements of enterprise data security and compliance, and solves the pain points of traditional scattered use that cannot be audited uniformly.

[0058] In specific implementation, all process control operations (such as suspension, retry, and termination) need to pass through enterprise SSO identity authentication; ordinary employees can only operate the tasks they initiate; and IT administrators can intervene in abnormal processes across departments.

[0059] The AI service display module is specifically used for: The AI gateway interacts with the server based on the TLS protocol through load balancing and a security gateway, queries AI services that can be called according to the permission information, and pushes them to the visual interface of the employee terminal in the form of service cards for display.

[0060] The AI task execution module is specifically used for: The visual interface generates an AI process (such as "meeting reservation", "stock financial report generation", "customer complaint classification") by arranging each AI service through the DAG editor based on the triggered drag signal and click signal, and converts the AI process into an AI task (without the need for the user to write code or understand the MCP protocol), and configures the task parameters of the AI task, including at least the timeout time, the maximum number of concurrent calls, the number of retries, and the log path (for example, setting "the maximum number of calls to the large model is 20 per day" for the "finance department" to prevent resource abuse, and the task parameters are in JSON / YAML format, support version control and gray release, and are convenient for IT centralized management); The AI task is sent to the load balancing and security gateway through the AI gateway to call the corresponding AI service to execute the AI task.

[0061] DAG (Directed Acyclic Graph) is a directed acyclic graph used to visually represent the dependencies and execution order of each node in the AI workflow.

[0062] The abnormal event capture module is specifically configured to: During the execution of the AI task, the MCP state tracker tracks abnormal events including at least timeout exceptions, resource exceptions, communication exceptions, service exceptions, and logic exceptions, captures the triggered abnormal events, and pushes the captured abnormal events to the management terminal in real time through WeChat, email, SMS, or telephone for alarm, and performs recovery operations based on the pre-set recovery mechanism and corresponding recovery measures; support manual intervention or automatic switching of backup computing nodes to form a "use-monitoring-alarm-recovery" closed loop; once the MCP state tracker detects an exception (such as an empty result returned by the model, a timeout, or a service unavailable), an alarm event is generated immediately and pushed to the applicant and IT support personnel in real time, achieving "second-level fault perception and minute-level response".

[0063] When the MCP state tracker detects that the execution event of the computing node exceeds the pre-set threshold value through the timer, the timeout exception event is triggered; When the MCP state tracker detects that the computing node's container reports memory overflow (OOM), GPU memory usage > 90%, or CPU usage > 90% through the probe, the resource exception event is triggered; When the MCP state tracker detects that the TCP connection is interrupted or the connection times out, the communication exception event is triggered; When the MCP state tracker detects that the response status code of the AI service is an error code (such as HTTP 5xx), the service exception event is triggered; When the MCP state tracker determines that the output result of the AI service is abnormal through the preset output check rule (such as the node returning an empty result, missing a key field, or containing a large model that cannot be returned), an abnormal event of the logical exception is triggered.

[0064] MCP (Model Context Protocol) is a model context protocol, which is a standardized protocol for scheduling AI service processes and is transparent to employees.

[0065] The execution log management module is specifically used for: The server records at least the employee identity, abnormal event, calling behavior, calling process, input and output content, and execution time consumption of the AI task execution in real time, stores the execution log into a TiSQL database, synchronously pushes the execution log to an enterprise supervision platform, reads the stored execution log based on a preset period to generate an AI service use report (such as generating a “department AI use report”, an “application use report”, and a “cost allocation list” monthly).

[0066] The TiSQL database records all calling behaviors, process logs, and abnormal events, supports auditing AI use conditions by departments / personnel of the enterprise, and provides a data basis for alarm pushing.

[0067] With the unified billing function of the AI gateway and the full-link log recording of the TiSQL database, an AI use report and a cost allocation list can be generated according to departments or projects, so that the enterprise can clearly monitor AI investment and output; this solves the previous problem of scattered costs and difficulty in statistics, realizes controllable costs and measurable effects, provides data support for budget management and resource optimization of the enterprise, and improves the return on investment.

[0068] In summary, the advantages of the present application are: 1. Register various private deployment models, third-party API models, and SaaS large models as standardized AI services through a server; an AI gateway obtains an AI task request sent by an employee terminal, performs identity authentication based on the AI task request, matches corresponding permission information, then interacts with the server through a load balancing and security gateway, queries AI services that can be called by the permission information, and pushes to a visual interface of the employee terminal for display; the visual interface arranges AI services based on triggered drag signals and click signals to generate AI tasks, sends the AI tasks to the load balancing and security gateway through the AI gateway to call corresponding AI services to execute the AI tasks; during AI task execution, an MCP state tracker automatically captures exception events based on the captured exception events to perform alarm; the server records execution logs of AI task execution in real time and stores the execution logs; that is, by uniformly registering various AI models as standard services and controlling them through the AI gateway as a core hub, centralized management of service access, identity authentication, and permission allocation is achieved, fundamentally solving the problems of account dispersion and uncontrollable costs; on this basis, the load balancing and security gateway ensure communication security, combined with real-time monitoring and alarm mechanisms of full-link execution log recording and MCP state tracker, a deep security and reliability protection system covering data flow, service status, and business logic is constructed; ultimately, the above centralized control and security base provides visual service display and low-code drag arrangement capabilities for front-end employee terminals, greatly reducing the use threshold of AI services, and ultimately greatly improving the centralization, security, and ease of use of AI service management.

[0069] 2. By uniformly registering private deployment models, third-party API models, and SaaS large models as standardized AI services and configuring model metadata (such as model name, capability label, etc.), this standardized registration avoids fragmentation of AI services within an enterprise, reduces compatibility issues between different models, and thus improves management efficiency.

[0070] 3. The AI gateway performs multi-dimensional identity authentication (including identity legality, permission compliance, and quota availability) based on requests from employee terminals, and securely interacts with the server through the TLS protocol, ensuring that only authorized users can access specific AI services, preventing unauthorized access and data leakage, and the advantage is that the permissions are bound to departments and roles, achieving fine-grained access control and improving enterprise data security.

[0071] 4. By introducing a load balancing and security gateway combined with the TLS protocol, the high availability and stability of AI service calls are ensured; load balancing can disperse request pressure and avoid single point of failure, while the security gateway provides encrypted transmission and reduces network risks, which not only improves the execution efficiency of AI tasks but also reduces the probability of system downtime.

[0072] 5. The visual interface supports drag-and-drop and click signals, and the DAG editor allows for the arrangement of AI services (such as combination and order adjustment) and the configuration of task parameters (such as timeout and number of retries). This enables non-technical employees to easily create complex AI processes, lowers the barrier to entry, and improves work efficiency.

[0073] 6. The MCP status tracker automatically captures various abnormal events (such as timeouts, resource and communication anomalies) and issues real-time alerts (via WeChat, email, etc.) while triggering a recovery mechanism. This enables proactive monitoring and rapid response, reduces AI task interruption time, and improves system robustness.

[0074] 7. By recording execution logs in real time (including employee identity, abnormal events, input and output, etc.) and storing them in the TiSQL database, and generating usage reports synchronously, it provides enterprises with a complete audit trail, facilitating compliance checks and performance optimization; the advantage is enhanced transparency and traceability, meeting enterprise regulatory requirements.

[0075] 8. By unifying the registration of private deployments, third-party APIs, and SaaS models as standardized AI services, the solution achieves efficient integration and management of enterprise AI resources, significantly improving management uniformity and compatibility. Simultaneously, leveraging an AI gateway for multi-dimensional identity authentication and access control, combined with load balancing and a security gateway, ensures high system availability and secure data transmission, enhancing overall security and reliability. Furthermore, the visual interface supports drag-and-drop task orchestration, lowering the user's operational threshold and increasing the flexibility of task configuration. The MCP status tracker's real-time anomaly monitoring and automatic alarm mechanism can quickly respond to faults and reduce the risk of interruption. Finally, by recording detailed execution logs and generating usage reports, the solution provides complete audit trail capabilities, supporting enterprise compliance and performance optimization, thus demonstrating comprehensive advantages in efficiency, security, user experience, and operations.

[0076] While specific embodiments of the present invention have been described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and not intended to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A unified management method for AI services combined with MCP for enterprises, characterized in that: Includes the following steps: Step S1: The server registers all private deployment models, third-party API models, and SaaS big models as standardized AI services. Step S2: The AI ​​gateway obtains the AI ​​task request sent by the employee terminal, performs identity authentication based on the AI ​​task request, and matches the corresponding permission information. Step S3: The AI ​​gateway interacts with the server through the load balancer and security gateway to query the AI ​​services that can be called based on the permission information, and pushes them to the visual interface of the employee terminal for display. Step S4: The visualization interface orchestrates each AI service to generate AI tasks based on the triggered drag and click signals, and sends the AI ​​tasks to the load balancer and security gateway through the AI ​​gateway to call the corresponding AI service to execute the AI ​​tasks. Step S5: During the execution of the AI ​​task, abnormal events are automatically captured by the MCP status tracker, and alarms are issued based on the captured abnormal events; Step S6: The server records the execution log of the AI ​​task in real time and stores the execution log.

2. The unified management method for enterprise-oriented AI services combined with MCP as described in claim 1, characterized in that: Step S1 further includes: configuring the server to configure the model metadata of each AI service, including at least the model name, capability tag, permission tag, department, whether it involves confidentiality, API key, and call address; Step S2 specifically involves: The AI ​​gateway obtains AI task requests sent by employees through a visual interface, which carry their department and job role. Based on the AI ​​task requests, it performs identity authentication, including at least identity legitimacy, permission compliance, and quota availability. After successful authentication, it matches the corresponding permission information. Step S3 specifically involves: The AI ​​gateway interacts with the server via load balancing and a security gateway, based on the TLS protocol, to query the AI ​​services that can be invoked with the permission information, and pushes the information to the employee's terminal visualization interface in real time for display in the form of service cards.

3. The unified management method for enterprise-oriented AI services combined with MCP as described in claim 1, characterized in that: Step S4 specifically involves: The visualization interface, based on the triggered drag and click signals, uses a DAG editor to arrange each AI service, including at least combination, order adjustment, and connection relationship adjustment, to generate an AI process. The AI ​​process is then converted into an AI task, and the AI ​​task is configured with task parameters including at least timeout, maximum concurrency, number of retries, and log path. The AI ​​task is sent to the load balancer and security gateway through the AI ​​gateway to invoke the corresponding AI service to execute the AI ​​task.

4. The unified management method for enterprise-oriented AI services combined with MCP as described in claim 1, characterized in that: Step S5 specifically involves: During the execution of the AI ​​task, the MCP status tracker tracks at least timeout exceptions, resource exceptions, communication exceptions, service exceptions, and logic exceptions, captures the triggered exception events, and pushes the captured exception events to the management terminal in real time via WeChat, email, SMS, or telephone to issue an alarm. Based on the preset recovery mechanism, the corresponding recovery measures are matched and the recovery operation is performed. When the MCP state tracker detects that the number of execution events of the computing node exceeds a preset threshold through a timer, it triggers the timeout exception event. When the MCP status tracker detects through probes that a container on a compute node reports a memory overflow, GPU memory usage >90%, or CPU usage >90%, it triggers the abnormal event of the aforementioned resource anomaly. When the MCP status tracker detects a TCP connection interruption or connection timeout, it triggers the abnormal event of the communication anomaly. When the MCP status tracker detects that the response status code of the AI ​​service is an error code, it triggers an abnormal event for the service. When the MCP state tracker determines that the output result of the AI ​​service is abnormal according to the preset output verification rules, it triggers the abnormal event of the logical abnormality.

5. The unified management method for enterprise-oriented AI services combined with MCP as described in claim 1, characterized in that: Step S6 specifically involves: The server records in real time the execution logs of the AI ​​task, including at least employee identity, abnormal events, invocation behavior, invocation process, input and output content, and execution time. The execution logs are stored in the TiSQL database and pushed synchronously to the enterprise monitoring platform. The server reads the stored execution logs at a preset period to generate an AI service usage report.

6. A unified AI service management system for enterprises, incorporating MCP, characterized in that: Includes the following modules: The AI ​​service registration module is used by the server to uniformly register various private deployment models, third-party API models, and SaaS large models as standardized AI services. The identity authentication module is used by the AI ​​gateway to obtain AI task requests sent by employee terminals, perform identity authentication based on the AI ​​task requests, and match the corresponding permission information. The AI ​​service display module is used by the AI ​​gateway to interact with the server through load balancing and security gateway, query the AI ​​services that can be called with the permission information, and push them to the visual interface of the employee terminal for display. The AI ​​task execution module is used by the visualization interface to orchestrate the AI ​​services based on the triggered drag and click signals to generate AI tasks, and send the AI ​​tasks to the load balancer and security gateway through the AI ​​gateway to call the corresponding AI service to execute the AI ​​tasks. An abnormal event capture module is used to automatically capture abnormal events through the MCP status tracker during the execution of the AI ​​task and to issue alarms based on the captured abnormal events. The execution log management module is used by the server to record and store the execution logs of the AI ​​task in real time.

7. The unified AI service management system for enterprises combined with MCP as described in claim 6, characterized in that: The AI ​​service registration module is also used for: configuring the server to include model metadata for each AI service, including at least the model name, capability tag, permission tag, department, whether it involves confidentiality, API key, and call address; The identity authentication module is specifically used for: The AI ​​gateway obtains AI task requests sent by employees through a visual interface, which carry their department and job role. Based on the AI ​​task requests, it performs identity authentication, including at least identity legitimacy, permission compliance, and quota availability. After successful authentication, it matches the corresponding permission information. The AI ​​service display module is specifically used for: The AI ​​gateway interacts with the server via load balancing and a security gateway, based on the TLS protocol, to query the AI ​​services that can be invoked with the permission information, and pushes the information to the employee's terminal visualization interface in real time for display in the form of service cards.

8. The unified AI service management system for enterprises combined with MCP as described in claim 6, characterized in that: The AI ​​task execution module is specifically used for: The visualization interface, based on the triggered drag and click signals, uses a DAG editor to arrange each AI service, including at least combination, order adjustment, and connection relationship adjustment, to generate an AI process. The AI ​​process is then converted into an AI task, and the AI ​​task is configured with task parameters including at least timeout, maximum concurrency, number of retries, and log path. The AI ​​task is sent to the load balancer and security gateway through the AI ​​gateway to invoke the corresponding AI service to execute the AI ​​task.

9. The unified AI service management system for enterprises combined with MCP as described in claim 6, characterized in that: The exception event capture module is specifically used for: During the execution of the AI ​​task, the MCP status tracker tracks at least timeout exceptions, resource exceptions, communication exceptions, service exceptions, and logic exceptions, captures the triggered exception events, and pushes the captured exception events to the management terminal in real time via WeChat, email, SMS, or telephone to issue an alarm. Based on the preset recovery mechanism, the corresponding recovery measures are matched and the recovery operation is performed. When the MCP state tracker detects that the number of execution events of the computing node exceeds a preset threshold through a timer, it triggers the timeout exception event. When the MCP status tracker detects through probes that a container on a compute node reports a memory overflow, GPU memory usage >90%, or CPU usage >90%, it triggers the abnormal event of the aforementioned resource anomaly. When the MCP status tracker detects a TCP connection interruption or connection timeout, it triggers the abnormal event of the communication anomaly. When the MCP status tracker detects that the response status code of the AI ​​service is an error code, it triggers an abnormal event for the service. When the MCP state tracker determines that the output result of the AI ​​service is abnormal according to the preset output verification rules, it triggers the abnormal event of the logical abnormality.

10. The unified AI service management system for enterprises combined with MCP as described in claim 6, characterized in that: The execution log management module is specifically used for: The server records in real time the execution logs of the AI ​​task, including at least employee identity, abnormal events, invocation behavior, invocation process, input and output content, and execution time. The execution logs are stored in the TiSQL database and pushed synchronously to the enterprise monitoring platform. The server reads the stored execution logs at a preset period to generate an AI service usage report.

Citation Information

Cited By

  • MCP multi-service collaborative cross-platform AI test resource arrangement method and system

    CN121807727A