AI component calling method and device, medium and computer program product

By adopting a microservice architecture with a unified protocol and containerized encapsulation in the AI ​​system, the problem of low calling efficiency caused by the monolithic architecture of AI components is solved, and efficient AI component calling and resource management are achieved.

CN120909681AActive Publication Date: 2025-11-07ZHONGDIAN DATA IND CO LTD

Patent Information

Application Number
CN202511407170.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2025-11-07
Estimated Expiration
2045-09-29

AI Technical Summary

Technical Problem

In existing AI systems, the monolithic architecture of AI components leads to low invocation efficiency, mainly due to high system coupling, rigid resource allocation, heterogeneous protocols, and lack of elasticity, resulting in low efficiency when large model applications call AI components.

Method used

We adopt a microservice architecture based on a unified protocol. Each AI component is encapsulated in containers and built into a microservice. The MCP protocol is used to achieve unified scheduling at the protocol layer. Combined with a two-level scheduling mechanism and a closed-loop health monitoring system, we can achieve independent operation and maintenance and on-demand scaling of AI components.

Benefits of technology

It improves the efficiency of large-scale model applications calling AI components, avoids the inefficiency of interface development caused by protocol heterogeneity, and ensures service continuity and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909681A_ABST
    Figure CN120909681A_ABST
Patent Text Reader

Abstract

The invention discloses an AI component calling method and device, a medium and a computer program product, and relates to the technical field of container mirroring, the method comprises the following steps: when a first calling request sent by a large model application through a gateway is detected, determining a target AI component container instance corresponding to the first calling request, the AI component container instance is a container instance obtained after containerization packaging processing is carried out on an AI component capable of being deployed independently, and comprises a model file and a protocol adapter; performing format conversion on the first calling request through a protocol adapter in the target AI component container instance according to a preset protocol format to obtain a second calling request; calling a model corresponding to the model file according to the second calling request for task processing to obtain a processing result; and packaging the processing result into a response message through the protocol adapter according to a preset protocol format, and sending the response message to the large model application through the gateway. According to the method, the calling efficiency of calling each AI component by the large model application is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of container image, and particularly to an AI component calling method, device, medium and computer program product. BACKGROUND

[0002] In current AI (Artificial Intelligence) systems, each AI component (such as a Natural Language Processing (NLP) service, a Computer Vision (CV) service, a speech recognition service, etc.) is usually deployed using a monolithic architecture. When an external system (such as a large model application) needs to call each AI component, various defects exist, which in turn leads to low calling efficiency. For example, due to the protocol heterogeneity of AI components, the large model application needs to develop different protocol interfaces to call different AI components, which in turn leads to low calling efficiency. Therefore, how to improve the calling efficiency of the large model application calling each AI component has become a problem that needs to be solved urgently.

[0003] The above content is only used to assist in understanding the technical solutions of the present application and does not represent the acknowledgement of the above content as prior art. SUMMARY

[0004] The main purpose of the present application is to provide an AI component calling method, device, medium and computer program product, which aims to solve the technical problem of how to improve the calling efficiency of the large model application calling each AI component.

[0005] To achieve the above purpose, the present application provides an AI component calling method, which comprises: When a first calling request sent by a large model application through a gateway is detected, it is determined that an AI component container instance corresponding to the first calling request is a target AI component container instance, wherein the AI component container instance is a container instance obtained by containerizing and encapsulating an AI component that can be independently deployed, and the AI component container instance includes a model file and a protocol adapter; The protocol adapter in the target AI component container instance is used to format convert the first calling request according to a preset protocol format, to obtain a second calling request; The model corresponding to the model file in the target AI component container instance is called according to the second calling request to perform task processing, to obtain a processing result; The protocol adapter in the target AI component container instance is used to encapsulate the processing result into a response message according to the preset protocol format, and the response message is sent to the large model application through the gateway.

[0006] Optionally, the AI component calling method further comprises: Decouple a plurality of monolithic architectures containing AI-capable components to obtain a plurality of AI components that can be deployed individually; Determine a dependency file corresponding to each AI component, the dependency file containing at least one dependency item, wherein the dependency item includes at least one of a model file of the AI component, an inference engine library, an operating system package, a protocol adapter, and a health monitoring script; For each AI component, containerize and package the AI component and the dependency file of the AI component to obtain an AI component container instance.

[0007] Optionally, the step of containerizing and packaging the AI component and the dependency file of the AI component to obtain an AI component container instance includes: Encapsulating the AI component and at least one dependency item in the dependency file into a preset container image to obtain an AI component container image; Determine a declaration file corresponding to the AI component container image, and deploy the AI component container image in a preset container orchestration platform using the declaration file; Determine an AI component container instance to be registered for service and published for metadata in the container orchestration platform based on the deployed declaration file, wherein the AI component container instance is registered for service and published for metadata in the container orchestration platform, so that a large model application can call the AI component container instance for task processing through a gateway.

[0008] Optionally, the preset container image includes two different first and second base images, and the step of encapsulating the AI component and the dependency file of the AI component into the preset container image to obtain an AI component container image includes: In the compilation environment stage, copy the AI component and at least one dependency item in the dependency file to the first base image to obtain a temporary build image, wherein a preset Python library and an operating system package are installed in the first base image, an inference engine in the inference engine library is downloaded and compiled to obtain a compiled inference engine binary file, and a model file is preheated based on the inference engine corresponding to the compiled inference engine binary file to obtain a preheating mark file; In the running environment stage, if the preheating state corresponding to the preheating mark file is preheating complete, copy the model file, the protocol adapter, the Python library, and the compiled inference engine binary file in the temporary build image to the second base image to obtain the AI component container image.

[0009] Optionally, the step of determining a declaration file corresponding to the AI component container image includes: determine the computing resource required by the AI component container instance according to a preset resource configuration rule, wherein the resource configuration rule comprises a scaling rule determined according to a real load of the computing resource and a preset load threshold, a computing resource support rule of the at least two AI component container instances, and a node affinity and stain tolerance rule of the computing resource; determine a declaration file according to the computing resource, wherein the declaration file comprises a specified AI component container image, a replica number and the computing resource, and the replica number is an instance number of the AI component container instance.

[0010] Optionally, the AI component calling method further comprises: When the large model application calls the target AI component container instance to process a task, a preset container probe is used to monitor a service state of the target AI component container instance in real time; When the service state is monitored to be abnormal, the task processing operation of the target AI component container instance is stopped, and a new target AI component container instance is reconstructed to continue the task processing according to the new target AI component container instance.

[0011] Optionally, the AI component calling method further comprises: receiving, by the AI component container instance corresponding to the first calling request, the first calling request sent by the gateway according to a preset intelligent routing decision logic, wherein the preset intelligent routing decision logic comprises that the gateway filters a plurality of available first AI component container instances from a preset registration center according to the first calling request sent by the large model application, and filters a first AI component container instance with the best performance from the plurality of first AI component container instances as the AI component container instance corresponding to the first calling request according to real-time monitoring data of the at least one first AI component container instance and a preset service quality requirement, and sends the first calling request to the AI component container instance corresponding to the first calling request.

[0012] In addition, to achieve the above-mentioned purpose, the present application further provides an AI component calling device, comprising: A determination module is configured to determine, when detecting that the large model application sends a first calling request through a gateway, an AI component container instance corresponding to the first calling request as a target AI component container instance, wherein the AI component container instance is a container instance obtained by containerizing and encapsulating an AI component that can be independently deployed, and the AI component container instance comprises a model file and a protocol adapter; A conversion module is configured to convert, by the protocol adapter in the target AI component container instance, a first calling request into a second calling request according to a preset protocol format; A processing module is configured to call a model corresponding to a model file in the target AI component container instance to process a task according to the second calling request, and obtain a processing result. The sending module is configured to encapsulate the processing result into a response message according to a preset protocol format through a protocol adapter in the target AI component container instance, and send the response message to the large model application through the gateway.

[0013] In addition, to achieve the above-mentioned purpose, the present application also provides an AI component calling device, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the computer program is configured to implement the steps of the AI component calling method as described above.

[0014] In addition, to achieve the above-mentioned purpose, the present application also provides a medium, which is a computer readable storage medium, and the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the AI component calling method as described above.

[0015] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps of the AI component calling method as described above.

[0016] In the embodiments of the present application, since the AI component container instance is a container instance obtained by containerizing and encapsulating the AI component that can be deployed independently, and at least contains a model file and a protocol adapter, the high coupling degree of the monolithic architecture system can be avoided, and the phenomenon of poor service continuity caused by the need for overall maintenance when upgrading or replacing the AI component can be avoided. When the first calling request sent by the large model application through the gateway is detected, the target AI component container instance corresponding to the first calling request is determined, and the protocol adapter in the target AI component container instance is used to convert the format of the first calling request according to a preset protocol format to obtain a second calling request. Then, the model corresponding to the model file in the target AI component container instance is called to perform task processing, and the processing result is obtained. Then, the protocol adapter is used to encapsulate the processing result into a response message according to the preset protocol format, and the response message is sent to the large model application through the gateway. Thus, the phenomenon of low calling efficiency caused by the protocol heterogeneity of the AI component and the need for developing different protocol interfaces when the large model application calls different AI components can be avoided. The calling of heterogeneous components by the large model application through a single protocol can be implemented, and the calling efficiency of the large model application for calling each AI component is improved. BRIEF DESCRIPTION OF DRAWINGS

[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and serve to explain the principles of the present application together with the specification.

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, for those skilled in the art, other drawings can also be obtained based on these drawings without any creative effort.

[0019] Figure 1 The flowchart provided by the first embodiment of the AI component calling method of the present application; Figure 2 The flowchart provided by the second embodiment of the AI component calling method of the present application; Figure 3 The flowchart of the AI component calling method of the present application; Figure 4 The module diagram of the AI component calling device of the embodiment of the present application; Figure 5 The device structure diagram of the hardware running environment involved in the AI component calling method of the embodiment of the present application.

[0020] The purpose implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0021] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application, and are not used to limit the present application.

[0022] In order to better understand the technical solutions of the present application, the following will be described in detail in combination with the drawings in the specification and specific embodiments.

[0023] In the current AI system, each AI component (such as natural language processing NLP service, computer vision CV service, speech recognition service, etc.) is usually deployed by using monolithic architecture, and thus has the following defects: High system coupling degree: AI component upgrade or replacement needs overall downtime maintenance, poor service continuity; Resource allocation rigidity: in high concurrency scenarios, the AI component with high frequency cannot be independently expanded, causing inefficient AI component to occupy computing resources (such as GPU); Protocol heterogeneity: different AI components need to be customized to develop interfaces, and there is a problem of non-uniform cross-component communication protocol, low development efficiency; Lack of elastic capability: the traditional deployment method cannot dynamically scale instances according to the load, and the fault recovery depends on manual intervention.

[0024] The above defects seriously affect the availability and resource utilization of the AI system, resulting in low calling efficiency of the AI component by the large model application.

[0025] In this embodiment, AI component decoupling deployment is achieved by constructing a microservice architecture based on a unified protocol. The microservice architecture based on a unified protocol can be a microservice architecture based on MCP (Model Calling Protocol) protocol. MCP is a standard protocol for communication between models and tools. It allows models to dynamically call external tools (such as APIs (Application Programming Interface), databases, etc.) according to task requirements, thereby enhancing model capabilities.

[0026] In this embodiment, when a large model application calls AI components through a microservice architecture based on a unified protocol, the calling efficiency can be improved, and the following advantages can be achieved: Unified protocol layer scheduling: define a standardized MCP communication protocol (including fields such as service identification, input and output format, Qos (Quality of Service) quality requirements, and timeout mechanism), so that a large model can call heterogeneous components (i.e., AI components) through a single protocol; Container encapsulation: encapsulate each AI component (such as OCR (Optical Character Recognition), speech recognition) as an independent container instance (i.e., AI component container instance), which includes an inference engine for model inference and a protocol adapter (such as MCP protocol adapter); Double-level scheduling mechanism: dynamically route requests by a gateway (such as MCP gateway) (select the optimal AI component container instance according to Qos requirements and real-time load), and link a container orchestration platform to achieve resource elastic supply; Health monitoring closed loop: real-time detection of AI component service state through container probes, and instance reconstruction or traffic switching when abnormal.

[0027] Therefore, the microservice architecture based on a unified protocol in this embodiment can achieve independent operation and on-demand scaling of AI components through protocol decoupling and service isolation.

[0028] It should be noted that the execution subject of the present embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, an AI component calling device, etc. that can realize the above functions. The following will take the AI component calling device as an example to illustrate the present embodiment and the following embodiments.

[0029] Based on this, the present application provides an AI component calling method, which is described with reference to Figure 1 , Figure 1 The flowchart of the first embodiment of the AI component calling method of the present application is shown in the figure.

[0030] In this embodiment, the AI component calling method includes steps S10-S40.

[0031] Step S10, when detecting that the first calling request sent by the large model application through the gateway, determine the AI component container instance corresponding to the first calling request as the target AI component container instance; It should be noted that the AI component container instance is a container instance obtained by containerizing and encapsulating the AI component that can be independently deployed, and the AI component container instance includes a model file and a protocol adapter; Optionally, the model file can be a file embodying the function of the model, such as a file containing the code of the model, and the model corresponding to the model file can be a model embodying the function of the AI component, such as a speech recognition model, a text recognition model, etc.

[0032] Optionally, the protocol adapter can be an MCP protocol adapter, and each AI component container instance has one protocol adapter, such as an MCP protocol adapter.

[0033] Optionally, the terminal where the large model application is located, the gateway, and the terminal where the microservice architecture based on the unified protocol are connected in turn. The terminal where the microservice architecture based on the unified protocol can include multiple AI component container instance pools, and can be provided with a container orchestration platform, AI components (such as AI components with natural language processing NLP service functions, AI components with computer vision CV service functions, and AI components with speech recognition services) and the like.

[0034] Optionally, the large model application can be an application in a terminal (such as a server) that can use a large model. The large model can be a natural language large model or other types of large models, which are not limited here.

[0035] Optionally, the external system such as the large model application can call the capabilities of the AI component through the gateway, and can perform request analysis of the large model application in the gateway when sending a request to the gateway, determine the required service capability type, and then select the optimal AI component container instance as the target AI component container instance according to the service capability type from the registration center. The instance pool containing multiple AI component container instances can be combined with real-time monitoring data (such as the current load, historical response delay, error rate, GPU (Graphics Processing Unit, graphics processor) utilization, and geographic location of each AI component container instance) and Qos requirements, etc., and the request of the large model application is forwarded to the target AI component container instance as the first calling request.

[0036] Optionally, the AI component calling method further comprises: receiving, by the first calling request corresponding AI component container instance, the first calling request sent by the gateway according to the preset intelligent routing decision logic.

[0037] It should be noted that the preset intelligent routing decision logic comprises: screening, by the gateway, a plurality of available first AI component container instances from the preset registration center according to the first calling request sent by the large model application, and screening, from the plurality of first AI component container instances, a first AI component container instance with optimal performance as the first calling request corresponding AI component container instance according to real-time monitoring data of at least one first AI component container instance and a preset service quality requirement, and sending the first calling request to the first calling request corresponding AI component container instance.

[0038] Optionally, the first AI component container instance can be an AI component container instance that has completed service registration and metadata publishing in the preset registration center through the container orchestration platform.

[0039] Optionally, the real-time monitoring data can include current load, historical response delay, error rate, GPU utilization, geographical location, etc. of the AI component container instance.

[0040] Optionally, the service quality requirement can be a Qos requirement, such as low delay, high precision, and specific version, etc.

[0041] Optionally, the first AI component container instance with optimal performance can be an AI component container instance with the lowest delay, an AI component container instance with the smallest number of connections, etc.

[0042] Optionally, the large model application sends the first calling request to the gateway, and the gateway can screen the first calling request corresponding AI component container instance from the registration center according to the preset intelligent routing decision logic, i.e., the gateway screens a plurality of available first AI component container instances from the preset registration center according to the first calling request sent by the large model application, and screens, from the plurality of first AI component container instances, a first AI component container instance with optimal performance as the first calling request corresponding AI component container instance according to real-time monitoring data of at least one first AI component container instance and a preset service quality requirement, and forwards the first calling request to the first calling request corresponding AI component container instance. The first calling request corresponding AI component container instance receives the first calling request sent by the gateway and performs subsequent task processing.

[0043] In this embodiment, the gateway first determines a plurality of available first AI component container instances in the registration center, then filters the first AI component container instance with the best performance according to real-time monitoring data and preset service quality requirements as the AI component container instance corresponding to the first invocation request, and forwards the first invocation request to the AI component container instance corresponding to the first invocation request. The AI component container instance corresponding to the first invocation request receives the first invocation request sent by the gateway and performs subsequent task processing, thereby ensuring that the large model application invokes an effective AI component container instance to efficiently complete the corresponding task.

[0044] In step S20, the protocol adapter in the target AI component container instance performs format conversion on the first invocation request according to the preset protocol format to obtain a second invocation request. Optionally, the preset protocol format can be a format consistent with the protocol adapter, such as the protocol format of the MCP protocol.

[0045] Optionally, after the target AI component container instance receives the first invocation request, the protocol adapter in the target AI component container instance can perform format conversion on the first invocation request, such as converting it into a standardized MCP communication protocol format invocation request, and taking it as the second invocation request.

[0046] In step S30, the model corresponding to the model file in the target AI component container instance is invoked to perform task processing according to the second invocation request, and a processing result is obtained. Optionally, in the target AI component container instance, the model corresponding to the encapsulated model file can be invoked according to the second invocation request, so as to perform task processing according to the invoked model, such as image recognition, voice translation, etc., thereby obtaining the processing result of the model. For example, if image recognition is required, the image to be recognized included in the second invocation request is input to the model, and the recognition result for the image to be recognized is output as the processing result.

[0047] In step S40, the protocol adapter in the target AI component container instance encapsulates the processing result into a response message according to the preset protocol format, and sends it to the large model application through the gateway.

[0048] Optionally, when the protocol adapter is an MCP protocol adapter, the MCP protocol adapter can be used to convert the processing result of the model into a standard MCP response format response message, send it to the gateway, and send it to the large model application through the gateway.

[0049] In addition, to assist in understanding the process of invoking the AI component by the large model application in this embodiment, the following examples are provided.

[0050] For example, an AI capability call initiated by an external system (such as a large model application) will be expressed as a MCP request in a standard format, which is first sent to the gateway. After receiving the request, the gateway will perform critical intelligent routing decision logic, including: request parsing, instance matching and selection, protocol conversion and invocation, result adaptation and return, and finally response.

[0051] Request parsing: The gateway parses the service_id (explicitly required service capability type) and Qos requirements (such as low latency, high precision, and specific version, etc.) in the request.

[0052] Instance matching and selection: According to the service_id, the gateway can filter out the available instance pool (including multiple available AI component container instances) that meets the conditions from the registration center. Then, combined with real-time monitoring data (such as the current load, historical response delay, error rate, GPU utilization, and geographical location of each AI component container instance, etc.) and Qos requirements, the gateway calculates and selects the optimal AI component container instance for request forwarding through optimized routing algorithms (such as lowest latency first, weighted round robin, minimum connection number, etc.), that is, the gateway sends the first invocation request to the target AI component container instance.

[0053] Protocol conversion and invocation: After the gateway sends the first invocation request to the target AI component container instance, the protocol adapter in the target AI component container instance, such as the MCP protocol adapter, processes the first invocation request and invokes the encapsulated model for task processing.

[0054] Result adaptation and return: After the model processing is completed, the processing result is generated, and the protocol adapter converts the processing result into a response message in the standard MCP response format and sends it to the gateway.

[0055] Final response: The gateway sends the response message in the standard MCP response format to the large model application.

[0056] This complete invocation process ensures that the invoker (i.e., the large model application) only needs to follow one protocol to transparently access all different backend services.

[0057] In the embodiment, since the AI component container instance is a container instance obtained after containerization encapsulation processing is performed on an AI component that can be separately deployed, and at least contains a model file and a protocol adapter, the high coupling degree of a monolithic architecture system can be avoided, the phenomenon that the AI component needs to be upgraded or replaced and the whole system needs to be shut down for maintenance, and the service continuity is poor, and when it is detected that the first invocation request sent by the large model application passes through the gateway, the target AI component container instance corresponding to the first invocation request is determined, and the protocol adapter in the target AI component container instance is used to convert the first invocation request into a second invocation request according to a preset protocol format, the model corresponding to the model file in the target AI component container instance is invoked according to the second invocation request to process a task, a processing result is obtained, the protocol adapter is used to encapsulate the processing result into a response message according to the preset protocol format, and the response message is sent to the large model application through the gateway, so that the phenomenon that the large model application needs to develop different protocol interfaces to call different AI components due to the protocol heterogeneity of the AI components, and the calling efficiency is low, the large model application can call heterogeneous components through a single protocol, and the calling efficiency of the large model application for calling each AI component is improved.

[0058] The second embodiment, Based on the first embodiment, the second embodiment is proposed, and the first embodiment is referred to Figure 2 In the embodiment, the AI component calling method further includes steps S100-S300.

[0059] Step S100, decoupling a plurality of monolithic architectures containing AI capability components to obtain a plurality of AI components that can be separately deployed; Optionally, component micro-service reconstruction and protocol adapter development can be performed to build a micro-service architecture based on a unified protocol.

[0060] Optionally, each AI capability component (such as a natural language processing (NLP) service, a computer vision (CV) service, and a speech recognition service) in the original monolithic architecture can be decoupled and split into independent, separately deployable and maintainable functional units, that is, independent, separately deployable and maintainable AI components.

[0061] Optionally, a protocol adapter dedicated to each AI component, such as an MCP protocol adapter, can be developed and deployed, and the protocol adapter can be integrated into the AI component.

[0062] Optionally, the protocol adapter can realize bidirectional conversion, at the input end of the AI component, receiving and parsing the standard MCP request from the gateway, extracting the payload (such as image data, text content), and converting it into an input format understandable by the internal API of the AI component; at the output end of the AI component, receiving the processing result (such as recognized objects, translated text) of the internal AI component, and encapsulating it into standardized response information according to the MCP protocol specification, so that the heterogeneous components expose a unified communication protocol to the outside.

[0063] Step S200, determining the dependency file corresponding to each AI component, which contains at least one dependency; It should be noted that the dependency includes at least one of the model file of the AI component, the inference engine library, the operating system package, the protocol adapter, and the health monitoring script. Optionally, the model file can include a model weight file, a model body code, a model identifier, etc. The model can be a model that implements the basic function of the AI component, such as a speech recognition model if the basic function of the AI component is speech recognition.

[0064] Optionally, the inference engine library includes at least one inference engine, such as a TensorRT engine.

[0065] Optionally, the operating system package can include various operating systems.

[0066] Optionally, the protocol adapter can be a unified protocol adapter, such as an MCP protocol adapter.

[0067] Optionally, the health monitoring script can be used to control the container probe to monitor the service state of the AI component container, and automatically trigger AI component container instance reconstruction or traffic switching when an exception occurs.

[0068] Optionally, the dependency file can include multiple dependencies, including but not limited to a model file, an inference engine or inference engine library, a necessary operating system package, an MCP protocol adapter (which can be embodied by MCP protocol adapter code), and a health monitoring script, etc.

[0069] Step S300, for each AI component, containerizing and encapsulating the AI component and the dependency file of the AI component to obtain an AI component container instance.

[0070] Optionally, after the component micro-service reconstruction and protocol adapter development, container image construction and environment encapsulation can be performed to build a micro-service architecture based on a unified protocol.

[0071] Optionally, each independent AI component and each dependency in the dependency file can be containerized and encapsulated to obtain an AI component container instance.

[0072] Optionally, the AI component and the dependency can be encapsulated into a Docker container image, and the build steps can be defined by writing a Dockerfile (Dockerfile is a description file for Docker image construction), including selecting a base image (such as an NVIDIA PyTorch image containing a specific CUDA version), copying code and model files, installing runtime dependencies, exposing service ports, configuring health check probes (such as HTTP GET / healthz endpoints), and setting commands to run at startup, such as starting the MCP protocol adapter service, finally building an AI component container image with high portability, which can be deployed in any compatible Docker / Kubernetes environment with the same runtime consistency, thereby solving the coupling and deployment rigidity problem of monolithic applications.

[0073] Optionally, the AI component container image can be deployed and resource configured by a container orchestration management platform (such as Kubernetes (K8s)), and service registration and metadata publishing can be performed so that the gateway can discover and call the AI component container instance.

[0074] In the embodiment, by decoupling a plurality of monolithic architectures, a plurality of AI components are obtained, dependency files corresponding to each AI component are determined, and for each AI component, the AI component and the dependency file are containerized and encapsulated to obtain an AI component container instance, thereby ensuring the effectiveness of the obtained AI component container instance.

[0075] A third embodiment, Based on the first or second embodiment, a third embodiment is proposed, in which step S300, the step of containerizing and encapsulating the AI component and the dependency file of the AI component to obtain an AI component container instance, includes steps a10-a30.

[0076] Step a10, encapsulating the AI component and at least one dependency in the dependency file into a preset container image to obtain an AI component container image; Optionally, each dependency in the dependency file of the AI component can be determined, and the AI component and at least one dependency can be encapsulated into a preset container image to obtain an AI component container image.

[0077] Optionally, each independent AI component can correspond to an AI component container image.

[0078] Optionally, after obtaining the AI component container image, the AI component container image can be stored in a preset image repository for subsequent invocation by a container orchestration platform.

[0079] Optionally, the preset container image includes two different first and second base images.

[0080] Optionally, the first base image can be a base image containing a complete compilation tool chain, used to install compilation-time dependencies, download and compile source code of a high-performance inference engine (such as TensorRT, ONNX Runtime), and install Python dependency packages, etc.

[0081] Optionally, the second base image can be a minimal base image, such as nvidia / cuda:11.8.0-runtime-ubuntu22.04.

[0082] Optionally, in step a10, the AI component and the dependency file of the AI component are encapsulated into the preset container image to obtain an AI component container image, including steps a11-a12.

[0083] Step a11, in the compilation environment phase, at least one dependency in the AI component and the dependency file is copied to the first base image to obtain a temporary build image. It should be noted that the preset Python library and operating system package are installed in the first base image, the inference engine in the inference engine library is downloaded and compiled to obtain a compiled inference engine binary file, and the model file is preheated according to the inference engine corresponding to the compiled inference engine binary file to obtain a preheating mark file. Optionally, the temporary build image can include the first base image with the copied AI component, dependency, preheating mark file, Python, complete compilation tool chain, etc.

[0084] Optionally, when building the AI component container image, a multi-stage building method can be used to optimize the image size and security.

[0085] Optionally, the multi-stage building method can include a compilation environment phase (builder phase) and a running environment phase (runtime phase).

[0086] Optionally, in the compilation environment phase, a base image containing a complete compilation tool chain can be selected as the first base image, and in the first base image, operating system packages and Python libraries are installed, and inference engines are downloaded and compiled in the inference engine library. The AI component and at least one dependency (such as a protocol adapter, a model file, etc.) of the AI component are copied to the first base image.

[0087] Optionally, in the first base image, the downloaded inference engine can be compiled according to the complete compilation tool chain and the Python library, and the model file can be loaded into the inference engine for model warm-up processing. That is, in the compilation environment stage, the model warm-up and caching mechanism can be implemented, and a special RUN command can be added in the build instruction of the Dockerfile, which can execute a model warm-up script. The model warm-up script can download the required model weight file from the model repository (such as Hugging Face, S3 storage bucket) in the build stage, and perform simulation inference on the model file.

[0088] Optionally, the model corresponding to the model file can be simulated to complete the model warm-up processing (such as GPU memory allocation, CUDA kernel compilation, etc.), and a warm-up marker file can be generated to determine whether the image construction operation in the running environment stage needs to continue according to the warm-up marker file.

[0089] Step a12, in the running environment stage, if the warm-up state corresponding to the warm-up marker file is warm-up completed, the model file, protocol adapter, Python library and compiled inference engine binary file in the temporary build image are copied to the second base image to obtain the AI component container image.

[0090] Optionally, in the running environment stage, the warm-up state corresponding to the warm-up marker file can be judged. If the warm-up state is warm-up completed, the model file, protocol adapter, Python library and compiled inference engine binary file in the temporary build image in the compilation environment stage can be copied to a very simple second base image to obtain a running stage image, which is used as the AI component container image.

[0091] In this embodiment, by encapsulating the AI component and at least one dependency into the preset container image, and using a multi-stage construction method when performing container encapsulation, that is, copying the AI component and the dependency to the first base image in the compilation environment stage to obtain a temporary build image, and copying the model file, protocol adapter, Python library and compiled inference engine binary file in the temporary build image to the second base image in the running environment stage to obtain the AI component container image, the size of the AI component container image is reduced, and the effectiveness of the obtained AI component container image is guaranteed.

[0092] Step a20, determine the declaration file corresponding to the AI component container image, and deploy the declaration file of the AI component container image in the preset container orchestration platform; Optionally, the declaration file of the AI component container corresponding to the AI component container image can be edited and deployed in the container orchestration platform.

[0093] Optionally, in step a20, a declaration file corresponding to the AI component container image is determined, including steps a21-a22.

[0094] In step a21, the computing resources required by the AI component container instance are determined according to a preset resource configuration rule, wherein the resource configuration rule includes an expansion and contraction rule determined according to a real load of the computing resource and a preset load threshold, a computing resource support rule of at least two AI component container instances, and a node affinity and taint tolerance rule of the computing resource. Optionally, the computing resource can be a GPU resource or a CPU resource. The following is an example of a GPU resource.

[0095] A GPU allocation strategy can be configured for the GPU resource to meet the computing resource requirements of the AI component container instance when performing task processing.

[0096] Optionally, GPU resource declaration and request can be performed: in the resources.limiits (resource limit) and requests (resource request) of the AI component container, nvidia.com / gpu:1 is explicitly declared, and the scheduler of the container orchestration platform (such as Kubernetes) is forced to allocate a complete and exclusive GPU device for the Pod (the smallest scheduling unit). Optionally, the Pod can include the AI component container instance.

[0097] Optionally, GPU resource allocation can be performed according to the resource configuration rule to determine the computing resources, such as GPU resources, required by at least one AI component container instance.

[0098] Optionally, the node affinity and taint tolerance rule of the computing resource: through nodeAffinity (node affinity), the Pod is forced to be scheduled only on the node with a specific model GPU (such as NVIDIA-A100) indicated in the label. At the same time, the taint (Taints) of the GPU node is set, such as kubectl taint nodes gpu-node-1 nvidia.com / gpu=true:NoSchedule. The Pod can tolerate the taint through tolerations (taint tolerance), ensuring that the expensive GPU node is only used by the Pod that needs GPU, and achieving logical isolation of physical resources.

[0099] Optionally, the computing resource support rules of the at least two AI component container instances, such as multi-instance GPU (MIG) support: for a GPU supporting MIG (such as NVIDIA-A100), a MIG partition instead of the entire GPU can be requested in the resource declaration (a declaration item included in the declaration file) (such as nvidia.com / mig-1g.5gb: 1), so as to split a single physical GPU into multiple independent micro-service instances (i.e., AI component container instances), realizing finer-grained resource sharing and higher utilization.

[0100] Optionally, the horizontal pod autoscaling (HPA) rules determined according to the actual load of the computing resource and the preset load threshold can be based on the set CPU and / or GPU utilization target threshold (TargetUtilization) to make the scaling decision. For example, the HPA can be configured to increase the number of replicas when the average GPU utilization of the AI component container is continuously higher than 70%, and to decrease the number of replicas when the average GPU utilization is continuously lower than 30%. K8s ensures that the AI component container instances are scheduled and run on the Node according to the strategy, realizing the elastic supply and isolation of resources.

[0101] Optionally, the horizontal pod autoscaling rules determined according to the actual load of the computing resource and the preset load threshold can include GPU actual load-based elastic scaling: the HorizontalPodAutoscaler (HPA) is configured to no longer depend on indirect indicators such as CPU or memory, but to be directly associated with the GPU actual utilization (nvidia_gpu_utilization) custom indicator collected by the Prometheus monitoring system. When the average GPU utilization of all instances (i.e., AI component container instances) is continuously higher than the preset threshold (such as 70%), the HPA will automatically increase the number of replicas, and vice versa, realizing the precise and dynamic response to the computing resource demand.

[0102] Step a22, determining a declaration file according to the computing resource, wherein the declaration file includes a specified AI component container image, a number of replicas, and a computing resource, and the number of replicas is the number of instances of the AI component container instances.

[0103] Optionally, each declaration item can be written in the container orchestration management platform according to the available computing resources (such as GPU resources, CPU resources, etc.), to form the declaration file.

[0104] Optionally, the declaration item can include a specified AI component container image, a number of instances of the AI component container instances, and a computing resource configured for the AI component container.

[0105] Optionally, the AI component container image stored in the image repository and its version number can be specified in the declaration file, and the number of replicas of the AI component container image is defined through the replicas field, and the required computing resources are consumed.

[0106] Optionally, the declaration file can be submitted to the K8s cluster, and the K8s schedules the AI component container instance according to the node resources and labels, and allocates the required computing resources (such as GPU) for the scheduled AI component container instance.

[0107] Optionally, the AI component container image specified in the declaration file is pulled from the image repository by the container orchestration platform (such as K8s), and is run on the nodes of the K8s cluster. The K8s automatically creates or destroys the AI component container instance corresponding to the specified AI component container image according to the replicas field, to ensure that there is always a specified number of AI component container instances running.

[0108] In this embodiment, the computing resources are determined according to the resource configuration rules, and the declaration file is determined according to the computing resources, thereby ensuring the validity of the determined declaration file.

[0109] Step a30, determining the AI component container instance to be registered for service and metadata publishing in the container orchestration platform according to the deployed declaration file. In the container orchestration platform, the AI component container instance is registered for service and metadata publishing, so that the large model application can call the AI component container instance for task processing through the gateway.

[0110] Optionally, a service object can be created according to a declaration file deployed in a container orchestration platform to perform a service registration and metadata publishing process. Optionally, a group of pods (such as AI component container instances) can be exposed as accessible service objects. When the AI component container instance is successfully started in K8s, the internal startup script automatically performs a service registration process, and the AI component container instance reports its key registration information by calling a range registration API endpoint provided by the MCP gateway. Optionally, the key registration information can include the running address (IP address or service discovery domain name) and port number of the AI component container instance, the core capability identifier (Service ID) it implements (such as 'nlp-translation-en2zh'), the service version number, the maximum processing capability supported (such as the maximum query rate per second Max QPS), the current health status (Healthy / Unhealthy), and other metadata that may be needed (such as supported input parameter options, model version). The MCP gateway receives and persists these registration information, and builds a global view of the entire microservice cluster capability, providing basic data support for subsequent intelligent routing decisions. That is, at this time, the large model application can call each AI component container instance in the K8s cluster through the gateway to perform task processing, such as voice translation, image recognition, etc.

[0111] Optionally, after the component microservice reconstruction and protocol adapter development, as well as the container image building and environment encapsulation, the cluster deployment and resource configuration, and the service registration and metadata publishing can be performed to build a microservice architecture based on a unified protocol.

[0112] Optionally, when performing cluster deployment and resource configuration, the Kubernetes (K8s) container orchestration platform can be used to manage the deployed and encapsulated AI component containers, and a container orchestration platform deployment declaration (Deployment) file can be written to specify the container image (such as a specified AI component container image), the number of replicas (Replicas), and the required resources (i.e., computing resources such as GPU resources, CPU resources).

[0113] Optionally, when performing service registration and metadata publishing, the declaration file can be used to determine the service object (such as a group of AI component container instances), and the service registration API endpoint provided by the gateway can be used to perform service registration and metadata publishing.

[0114] In the embodiment, the AI component and at least one dependency are encapsulated into a preset container image to obtain an AI component container image, a declaration file of the AI component container image is deployed, and service registration and metadata are published, so that the large model application can normally call the AI component container instance through the gateway to perform task processing, thereby ensuring the calling efficiency of the large model application in calling each AI component.

[0115] A fourth embodiment, Based on the first embodiment, the second embodiment or the third embodiment, the fourth embodiment is provided. In the embodiment, the AI component calling method further includes steps b10-b20.

[0116] In step b10, when the large model application calls a target AI component container instance for task processing, a preset container probe is used to monitor the service state of the target AI component container instance in real time. In step b20, when the service state is monitored to be abnormal, the task processing operation of the target AI component container instance is stopped, and a new target AI component container instance is reconstructed to continue the task processing according to the new target AI component container instance.

[0117] Optionally, at any link where the large model application starts to call the target AI component container instance for task processing until the task processing is completed and a feedback response message is received, the service state of the AI component container of the target AI component container instance can be detected by the container probe set in advance to determine the health state.

[0118] When the service state is monitored to be abnormal, the corresponding task processing operation can be stopped, a new target AI component container instance is reconstructed, and the task processing is continued according to the new target AI component container instance. Flow switching and other operations, such as reselecting the AI component container instance, can also be performed.

[0119] For example, as shown in Figure 3 The AI component can be sequentially subjected to component micro-service reconstruction, container image construction, cluster deployment and service registration. When the external large model application requests to dispatch and execute any AI component container instance of the AI component, operation and maintenance monitoring loop is performed to update the health state, such as real-time monitoring of the service state by the container probe to determine the health state of the AI component, and the expansion and contraction of the computing resources, such as GPU, are fed back.

[0120] In the embodiment, when the large model application calls the target AI component container instance for task processing, the container probe is used to monitor the service state corresponding to the target AI component container instance, and when it is determined that the service state is abnormal, the task processing operation of the target AI component container instance is stopped, and a new target AI component container instance is reconstructed to continue the task processing according to the new target AI component container instance, thereby ensuring the effective performance of the task processing.

[0121] In addition, in order to achieve the above-mentioned purpose, referring to Figure 4 The application further provides an AI component calling device, which comprises: The determining module A10 is configured to determine that the AI component container instance corresponding to the first calling request sent by the large model application through the gateway is a target AI component container instance when it is detected that the first calling request is sent by the large model application through the gateway, wherein the AI component container instance is a container instance obtained by containerizing and encapsulating an AI component that can be independently deployed, and the AI component container instance comprises a model file and a protocol adapter; The conversion module A20 is configured to convert the first calling request into a second calling request according to a preset protocol format through the protocol adapter in the target AI component container instance. The processing module A30 is configured to call a model corresponding to the model file in the target AI component container instance to perform task processing according to the second calling request, and obtain a processing result. The sending module A40 is configured to encapsulate the processing result into a response message according to the preset protocol format through the protocol adapter in the target AI component container instance, and send the response message to the large model application through the gateway.

[0122] The AI component calling device provided by the application adopts the AI component calling method in the above-mentioned embodiments, and can improve the calling efficiency of the large model application in calling each AI component. Compared with the prior art, the AI component calling device provided by the application has the same beneficial effects as the AI component calling method provided by the above-mentioned embodiments, and the other technical features in the AI component calling device are the same as the features disclosed in the above-mentioned embodiments, which will not be repeated here.

[0123] The application provides an AI component calling device, which comprises at least one processor and a memory in communication connection with the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the AI component calling method in the above-mentioned embodiment one.

[0124] The following refers to Figure 5FIG. 1 shows a structural diagram of an AI component calling device suitable for implementing embodiments of the present application. The AI component calling device in embodiments of the present application can include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), car terminals (e.g., car navigation terminals), and the like, as well as fixed terminals such as digital TVs, desktop computers, and the like. The device shown in the figure is merely an example and should not impose any limitation on the functions and use range of embodiments of the present application.

[0125] As shown, the AI component calling device can include a processing device 1001 (e.g., a central processor, a graphics processor, etc.) that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 1002 or loaded from a storage device 1003 into a random access memory (RAM) 1004. Various programs and data required for device operation are also stored in the RAM 1004. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. In general, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, and the like; output devices 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; the storage device 1003 including, for example, a magnetic tape, a hard disk, and the like; and a communication device 1009. The communication device 1009 can allow the AI component calling device to communicate wirelessly or by wire with other devices to exchange data. Although the AI component calling device having various systems is shown in the figure, it should be understood that all of the systems shown are not required to be implemented or provided. More or fewer systems can be alternatively implemented or provided.

[0126] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program codes for executing the method shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiments disclosed in the present application are executed.

[0127] The AI component calling device provided by the present application adopts the AI component calling method in the above-mentioned embodiments, which can improve the calling efficiency of the large model application in calling each AI component. Compared with the prior art, the AI component calling device provided by the present application has the same beneficial effects as the AI component calling method provided by the above-mentioned embodiments, and other technical features in the AI component calling device are the same as the features disclosed in the previous embodiment method, which will not be repeated here.

[0128] It should be understood that parts of the present application can be realized by hardware, software, firmware or a combination thereof. In the description of the above-mentioned embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0129] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0130] The present application provides a computer readable storage medium having stored thereon computer readable program instructions (i.e. computer program) for executing the AI component calling method in the above-mentioned embodiments.

[0131] The computer readable storage medium provided in the present application may, for example, be a U disk, but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, system, or device, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more conductive wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present embodiment, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer readable storage medium can be transmitted in any suitable medium, including but not limited to electrical wires, optical cables, RF (Radio Frequency), and the like, or any suitable combination of the above.

[0132] The above computer readable storage medium can be contained in the AI component calling device, or can exist separately without being assembled into the AI component calling device.

[0133] The above computer readable storage medium carries one or more programs, which, when executed by the AI component calling device, enable the AI component calling device to perform the step flow of the above AI component calling method.

[0134] Computer program code for carrying out operations of the present application can be written in one or more programming languages or combinations of languages including object oriented programming languages such as Java, Smalltalk, C++ or conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider through the Internet).

[0135] The flow and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present application. In this regard, each block in the flow and block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may be executed in the reverse order, depending on the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.

[0136] The modules involved in the embodiments of the present application can be implemented in software or in hardware. In some cases, the names of the modules do not limit the modules themselves.

[0137] The readable storage medium provided by the present application is a computer readable storage medium, which stores computer readable program instructions (i.e. computer programs) for executing the AI component calling method described above, and can improve the calling efficiency of the large model application in calling each AI component. Compared with the prior art, the computer readable storage medium provided by the present application has the same beneficial effects as the AI component calling method provided by the above embodiments, and will not be repeated here.

[0138] The present application also provides a computer program product comprising a computer program, which, when executed by a processor, implements the steps of the AI component calling method as described above.

[0139] The computer program product provided by the present application can improve the calling efficiency of the large model application in calling each AI component. Compared with the prior art, the computer program product provided by the present application has the same beneficial effects as the AI component calling method provided by the above embodiments, and will not be repeated here.

[0140] The above is only some embodiments of the present application, and does not limit the patent scope of the present application, and any equivalent structural transformation made by using the contents of the present application specification and drawings, or direct / indirect application in other related technical fields is included in the patent protection scope of the present application.

Claims

1. An AI component calling method, characterized in that, The method comprises: When a first call request sent by a large model application through a gateway is detected, it is determined that an AI component container instance corresponding to the first call request is a target AI component container instance, wherein the AI component container instance is obtained by containerizing encapsulation processing of an AI component that can be independently deployed, and the AI component container instance comprises a model file and a protocol adapter; The first call request is format-converted by the protocol adapter in the target AI component container instance according to a preset protocol format, to obtain a second call request; A model corresponding to the model file in the target AI component container instance is called according to the second call request to perform task processing, to obtain a processing result; The processing result is encapsulated into a response message by the protocol adapter in the target AI component container instance according to the preset protocol format, and is sent to the large model application through the gateway. 2.The AI component invocation method of claim 1, wherein, The AI component calling method further comprises: A plurality of monolithic architectures containing AI capability components are decoupled to obtain a plurality of AI components that can be independently deployed; A dependency file corresponding to each AI component is determined, wherein the dependency file contains at least one dependency of the AI component, and the dependency comprises at least one of a model file, an inference engine library, an operating system package, a protocol adapter and a health monitoring script of the AI component; For each AI component, the AI component and the dependency file of the AI component are containerized and encapsulated to obtain an AI component container instance. 3.The AI component invocation method of claim 2, wherein, The step of containerizing and encapsulating the AI component and the dependency file of the AI component to obtain an AI component container instance comprises: At least one dependency in the AI component and the dependency file is encapsulated into a preset container image to obtain an AI component container image; A declaration file corresponding to the AI component container image is determined, and the AI component container image is deployed by using the declaration file in a preset container orchestration platform; An AI component container instance to be registered for service and published for metadata in the container orchestration platform is determined according to the deployed declaration file, wherein the AI component container instance is registered for service and published for metadata in the container orchestration platform, so that a large model application can call the AI component container instance through a gateway to perform task processing. 4.The AI component invocation method of claim 3, wherein, The preset container image comprises two different first and second base images, and the step of encapsulating the AI component and the dependency file of the AI component into a preset container image to obtain an AI component container image comprises: In a compilation environment stage, at least one dependency in the AI component and the dependency file is copied to the first base image to obtain a temporary build image, wherein a preset Python library and an operating system package are installed in the first base image, an inference engine in the inference engine library is downloaded and compiled to obtain a compiled inference engine binary file, and a model preheating process is performed on the model file according to the inference engine corresponding to the compiled inference engine binary file to obtain a preheating mark file; In the running environment stage, if the preheating state corresponding to the preheating mark file is preheating completion, the model file, protocol adapter, Python library and compiled inference engine binary file in the temporary construction image are copied to the second base image to obtain an AI component container image. 5.The AI component invocation method of claim 3, wherein, The step of determining the declaration file corresponding to the AI component container image comprises: According to a preset resource configuration rule, determine the computing resource required by the AI component container instance, wherein the resource configuration rule comprises an expansion and contraction rule determined according to a real load of the computing resource and a preset load threshold, a computing resource support rule of at least two AI component container instances, and a node affinity and stain tolerance rule of the computing resource; According to the computing resource, determine the declaration file, wherein the declaration file comprises a specified AI component container image, a copy number and a computing resource, and the copy number is the instance number of the AI component container instance.

6. The AI component invocation method of any one of claims 1-5, wherein, The AI component calling method further comprises: When the large model application calls the target AI component container instance for task processing, a preset container probe is used to monitor the service state of the target AI component container instance in real time; When the service state is monitored to be abnormal, stop the task processing operation of the target AI component container instance, and reconstruct a new target AI component container instance to continue the task processing according to the new target AI component container instance.

7. The AI component invocation method of any one of claims 1-5, wherein, The AI component calling method further comprises: The AI component container instance corresponding to the first calling request receives the first calling request sent by the gateway according to a preset intelligent routing decision logic, wherein the preset intelligent routing decision logic comprises that the gateway filters a plurality of available first AI component container instances from a preset registration center according to the first calling request sent by the large model application, and filters the first AI component container instance with the best performance from the plurality of first AI component container instances as the AI component container instance corresponding to the first calling request according to the real-time monitoring data of at least one first AI component container instance and the preset service quality requirement, and sends the first calling request to the AI component container instance corresponding to the first calling request.

8. An AI component calling apparatus comprising: The device comprises a memory, a processor and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the AI component calling method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the AI component calling method according to any one of claims 1 to 7.

10. A computer program product, characterised in that, The computer program product comprises a computer program, and the computer program is executed by the processor to implement the steps of the AI component calling method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Request processing method of AI model, computer equipment, medium and product

    CN118972431A

  • Industrial software protocol proxy method and system based on swan mongolian operating system

    CN119052331A

  • AI large model and low-code platform interactive integration method and system

    CN119903506A

  • Associated plugin management method, device and system

    WO2015013936A1

Cited By

  • Multi-agent resource collaborative scheduling management system and method based on MCP protocol

    CN121277716A

  • Multi-agent-based data interaction method and device, server and program product

    CN121691485A