Universal support method and system for AI application in Linux system

By providing a unified AI application interface and task scheduling mechanism under the Linux system, the interface inconsistency and resource management difficulties of AI applications under the Linux system are solved, efficient development and cross-platform migration are achieved, and development efficiency and user experience are improved.

CN120429115APending Publication Date: 2025-08-05KYLIN CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510557237.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

AI applications under existing Linux systems have problems such as lack of unified interface standards, high interface learning costs, complex environment configuration, and difficult resource management, resulting in low development efficiency and high cross-platform migration costs.

Method used

It provides a general support method for AI applications under Linux systems. Through interface modules, communication layer modules and runtime modules, a unified AI application interface standard is realized, and the differences between various AI manufacturers are blocked, and the system resources are dynamically managed through the task scheduling module to optimize the environment configuration and resource use.

Benefits of technology

It reduces the adaptation and migration costs of developers, improves development efficiency, ensures rapid development environment construction and integration of AI functions, and improves the fluency and user experience of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429115A_ABST
    Figure CN120429115A_ABST
Patent Text Reader

Abstract

The invention discloses a universal support method and system for an AI application under a Linux system, and the method comprises the steps that an interface module receives a service request of the AI application, the interface module forwards the service request to a communication layer module when the received service request is valid, and the interface module comprises a public interface and a private interface; the communication layer module establishes connection for the AI application and the runtime module and transmits a service request to the runtime module; the runtime module responds to a service request of the AI application, generates a response result and returns the response result to the communication layer module; and the communication layer module returns a response result of the runtime module to the AI application initiating the service request. The invention aims to solve the problems of complex environment configuration, difficulty in resource management, low development efficiency and the like in the prior art, so that a developer of the AI application can quickly establish a development environment and quickly integrate AI functions, the development efficiency is improved, and the cross-platform migration cost and maintenance cost of the application are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of operating systems, and in particular to a general support method and system for AI applications under a Linux system. Background Art

[0002] With the continuous advancement of technologies such as deep learning and machine learning, the field of artificial intelligence (AI) has achieved significant breakthroughs. In particular, large-scale model technologies such as BERT and GPT have demonstrated powerful performance and broad application potential in areas such as natural language processing, computer vision, and speech processing. More and more companies and developers are looking to leverage these technologies to enhance the intelligence of their products and services, such as smart customer service, smart homes, autonomous driving, and smart offices. As an open-source and highly customizable operating system, Linux has become a preferred choice for many individuals and businesses due to its stability and flexibility. Furthermore, the recent rise of domestic operating systems based on Linux has made Linux a popular choice. However, the inherent complexity and specialized nature of AI technology makes it difficult for many non-specialized companies and developers to directly apply it. While most AI vendors offer development kits that lower the barrier to entry for AI applications by providing pre-packaged algorithms and interfaces, the interface standards vary across vendors. Applications that integrate with multiple AI vendors or switch between vendors require re-adaptation, which is a significant waste of manpower, material, and financial resources. Therefore, developing a universal AI application development solution for Linux is crucial. Currently, there are a wide variety of AI frameworks available for Linux systems, with the following key technical challenges: 1. Lack of unified interface standards: The interfaces and tools used by different AI frameworks vary significantly, requiring developers to learn multiple technology stacks. This high learning curve hinders the maintenance and expansion of AI-related functionality in later stages of an application. 2. High interface learning costs: The interfaces of some AI frameworks often require specialized knowledge to understand and access, resulting in a high learning curve. 3. Complex environment configuration: AI application development requires multiple dependent libraries and frameworks. Manually configuring these environments is time-consuming and error-prone, requiring developers to spend significant time setting up and debugging the environment. 4. Inadequate resource management: Effectively managing and scheduling computing resources in a multi-tasking and multi-user environment is a challenge. Therefore, achieving universal support for AI applications on Linux systems has become a key technical issue that needs to be addressed urgently. Summary of the Invention

[0003] Technical problem to be solved by the present invention: In response to the above-mentioned problems in the prior art, a general support method and system for AI applications under the Linux system are provided. The present invention aims to solve the problems existing in the prior art, such as complex environment configuration, difficult resource management, and low development efficiency, so that developers of AI applications can quickly build a development environment, quickly integrate AI functions, improve development efficiency, and reduce the cross-platform migration cost and maintenance cost of applications.

[0004] In order to solve the above technical problems, the technical solution adopted by the present invention is: A general support method for AI applications in Linux systems, including: S101: The interface module receives a service request from an AI application. When the service request is valid, the interface module forwards the service request to the communication layer module. The interface module includes a public interface and a private interface. The interface types provided by the public interface include natural language processing, visual processing, speech processing, text generation, and image generation. The private interface is embedded with an interface authentication module, a model configuration module, and an AI assistant. The interface authentication module is used to authenticate the service request. The model configuration module is used to convert the service request into a configuration service request for the AI model. The AI assistant is used to provide natural language processing, visual processing, speech processing, text generation, and image generation services, as well as interaction with the operating system. The model configuration module and the AI assistant execute the service request only after authentication is passed. S102: The communication layer module establishes a connection between the AI application and the runtime module, and transmits the service request to the runtime module; S103: The runtime module responds to the service request of the AI application. If the service request calls a public interface, the runtime module calls the AI model on the client or cloud side to generate a response result of the service request and returns it to the communication layer module. If the service request calls a private interface, if the service request is a configuration service request, the runtime module performs the configuration of the AI model on the client or cloud side and generates a configuration response result and returns it to the communication layer module. If the service request is a service request for an AI assistant, the runtime module provides natural language processing, visual processing, speech processing, text generation, and image generation services or interacts with the operating system, and generates an interaction response result and returns it to the communication layer module. S104: The communication layer module returns the response result of the runtime module to the AI application that initiated the service request.

[0005] Optionally, the runtime module includes a task scheduling module and a task queue. Before the runtime module responds to the service request of the AI application in step S103, it schedules the service request: S201, obtaining the current load status of the system. If the system is currently in a high load state, generating a response result rejecting the service request of the AI application and jumping to step S104; otherwise, jumping to step S202; S202: Put the service request of the AI application into the task queue, and take out the service request from the task queue and execute it according to the task priority and system resource load through the task scheduling module to respond to the service request of the AI application.

[0006] Optionally, obtaining the current load status of the system in step S201 includes: S301, obtaining a load status of a specified load type through a system interface provided by a system API center of a runtime module, where the specified load type includes part or all of CPU load, memory load, disk load, and GPU load; S302, the load state of each load type is determined to have a corresponding load level according to its preset grading threshold, wherein the load level includes a high load level; if the load level of more than a preset number of load types is a high load level, the system is determined to be in a high load state, otherwise the system is determined to be in a non-high load state.

[0007] Optionally, in step S202, when the task scheduling module takes out and executes the service request in the task queue according to the task priority and system resource load to respond to the service request of the AI application, it includes periodically obtaining the service requests of the AI model on the calling side and the cloud in the task queue according to the preset scheduling time window, scheduling the service requests of the AI model on the calling side and the cloud in a first-in-first-out manner, and only taking out and executing one service request of the AI model on the calling side and the cloud in each scheduling time window; including obtaining the service request for calling the local AI model in the task queue according to the preset scheduling period, calculating the scheduling priority of the service request for calling the local AI model according to the task priority and system resource load carried by the service request for calling the local AI model, and taking out and executing a service request of the local AI model with the highest scheduling priority according to the scheduling priority in each scheduling period.

[0008] Optionally, the runtime module also includes a privacy data management module, which is used to encrypt and store user data involved in the service request when the runtime module responds to the service request of the AI application, and decrypt and read the user data when the runtime module needs to read it; the runtime module also includes a context management module for managing historical messages of conversations with the AI model.

[0009] Optionally, the runtime module also includes a prompt word configuration management module, which is used to generate a preset prompt word for the service request when the runtime module responds to the service request of the AI application. When the AI model on the terminal side or the cloud is called to generate a response result of the service request, the prompt word and the service request are combined as inputs of the AI model on the terminal side or the cloud to generate a response result of the service request; the runtime module also includes an AI engine selection module, which is used to first screen out candidate AI models from the AI models on the terminal side and the cloud according to the computing power requirements of the service request when the AI model on the terminal side or the cloud is called to generate a response result of the service request and return it to the communication layer module, and then select the AI model with the highest priority according to the priority settings of each AI model in the candidate AI models, and call the AI model with the highest priority to generate a response result of the service request and return it to the communication layer module.

[0010] Optionally, the runtime module also includes a toolset, which includes a picture viewing plug-in, an instruction plug-in, a retrieval enhancement generation service and a system audio service. The picture viewing plug-in is used to display and output the generated image when generating the image; the instruction plug-in is used to provide a set of instructions supported by the runtime module for user selection; the retrieval enhancement generation service is used to generate a knowledge graph from the document input by the user for AI model call; the system audio service is used to provide system audio input and output functions for speech processing Speech.

[0011] In addition, the present invention also provides a general support system for AI applications under a Linux system, comprising a microprocessor and a memory connected to each other, wherein the microprocessor is programmed or configured to execute the general support method for AI applications under a Linux system.

[0012] In addition, the present invention also provides a computer-readable storage medium, which stores a computer program or instruction, and the computer program or instruction is programmed or configured to execute the general support method for AI applications under the Linux system through a processor.

[0013] In addition, the present invention also provides a computer program product, including a computer program or instructions, which are programmed or configured to execute the general support method for AI applications under the Linux system through a processor.

[0014] Compared with the prior art, the present invention can mainly achieve the following beneficial effects: the method of the present invention includes an interface module receiving a service request of an AI application, and the interface module forwards the service request to a communication layer module when the received service request is valid, and the interface module includes a public interface and a private interface; the communication layer module establishes a connection between the AI application and the runtime module, and passes the service request to the runtime module; the runtime module responds to the service request of the AI application, generates a response result and returns it to the communication layer module; the communication layer module returns the response result of the runtime module to the AI application that initiated the service request. The present invention implements a general development kit for AI applications under the Linux system through the interface module, the communication layer module and the runtime module, which can provide a unified standard and easy-to-understand interface for various AI capabilities, shield the differences implemented by various AI manufacturers, reduce the adaptation and migration costs of application developers, and solve the problems of complex environment configuration, difficult resource management, low development efficiency, etc. in the prior art, so that AI application developers can quickly build a development environment, quickly integrate AI functions, improve development efficiency, and reduce the cross-platform migration cost and maintenance cost of applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 Schematic diagram of the basic process of the method of the embodiment of the present invention.

[0016] Figure 2 Schematic diagram of the system architecture in an embodiment of the present invention.

[0017] Figure 3 Schematic diagram of the interaction process in an embodiment of the present invention. DETAILED DESCRIPTION

[0018] In order to enable those skilled in the art to better understand the technical solution of the present invention, the technical solution of the present invention will be further described in detail below with reference to the accompanying drawings in the embodiments of the present invention.

[0019] like Figure 1 As shown, the general support method for AI applications in the Linux system of this embodiment includes: S101: The interface module receives a service request from an AI application. When the service request is valid, the interface module forwards the service request to the communication layer module. The interface module includes a public interface and a private interface. The interface types provided by the public interface include natural language processing (NLP), vision processing (Vision), speech processing (Speech), text generation, and image generation. The private interface is embedded with an interface authentication module, a model configuration module, and an AI assistant. The interface authentication module is used to authenticate the service request. The model configuration module is used to convert the service request into a configuration service request for the AI model. The AI assistant is used to provide natural language processing, vision processing, speech processing, text generation, and image generation services, as well as interaction with the operating system. The model configuration module and the AI assistant execute the service request only after authentication is passed. S102: The communication layer module establishes a connection between the AI application and the runtime module, and transmits the service request to the runtime module; S103: The runtime module responds to the service request of the AI application. If the service request calls a public interface, the runtime module calls the AI model on the client or cloud side to generate a response result of the service request and returns it to the communication layer module. If the service request calls a private interface, if the service request is a configuration service request, the runtime module performs the configuration of the AI model on the client or cloud side and generates a configuration response result and returns it to the communication layer module. If the service request is a service request for an AI assistant, the runtime module provides natural language processing, visual processing, speech processing, text generation, and image generation services or interacts with the operating system, and generates an interaction response result and returns it to the communication layer module. S104: The communication layer module returns the response result of the runtime module to the AI application that initiated the service request.

[0020] like Figure 2 As shown, the interface module (named AI SDK in this embodiment) includes public interfaces and private interfaces. In this embodiment, the division into public interfaces and private interfaces is based on the following considerations: (1) The AI assistant (named OsAssistant in this embodiment) is an important entry point for users to interact with the operating system AI and is also the core competitiveness of the operating system. The interface used cannot be made public; (2) Some interactive capabilities of the operating system cannot be made public, such as intent recognition, model scheduling strategies, etc., which will be optimized for the operating system. (3) Some functions may be connected to private services; 4) The AI experience will continue to be optimized, and the interface changes may also be relatively frequent, making it unsuitable for external provision.

[0021] like Figure 2As shown, the interface types provided by the public interface in this embodiment include natural language processing NLP, visual processing Vision, speech processing Speech, text generation and image generation. Natural language processing NLP, visual processing Vision, and speech processing Speech are basic AI capabilities, and text generation and image generation are generative AI capabilities. There are certain differences in the usage scenarios of generative AI and basic AI, and from a functional perspective, the two also overlap. Applications may use one of them alone or in combination. It will be clearer if the interfaces are distinguished. Generally speaking, most generative AI uses related technologies of large models, and basic AI capabilities are more targeted at traditional application scenarios.

[0022] like Figure 2 As shown, in this embodiment, the private interface embeds an interface authentication module, a model configuration module, and an AI assistant (named OsAssistant in this embodiment). The interface authentication module is used to authenticate service requests. The model configuration module is used to convert service requests into configuration service requests for the AI model. The AI assistant is used to convert service requests into interactive management service requests for the operating system and the AI model. The model configuration module and the AI assistant perform service request conversion only after authentication. The runtime module is the specific implementation of the various capability interfaces described above, responsible for managing various AI services, model configuration services, and task scheduling services. It can connect to multiple cloud-side and device-side model services to achieve final model inference. The AI assistant converts service requests into interactive management service requests for the operating system and the AI model, providing higher-level interfaces such as context management and multimodal interfaces, as well as interfaces for AI assistant business scenarios, such as interfaces that provide instruction information. The configuration of the model configuration service is global; any application that uses the interface will use the configured information, such as the cloud-side model key and the device-cloud model scheduling policy.

[0023] like Figure 2As shown, the runtime module in this embodiment has basic service management functions and is responsible for registering and deregistering AI basic capability services, the AI assistant, and the model configuration module. It routes requests from clients (AI applications) to the AI models. It also provides a configuration service, offering model configuration and authentication services. For example, when configuring the key for a Kedaxun cloud model, this service performs authentication and verifies the key's validity. This service can be used to obtain specific configuration information for the model adapted by the current subsystem. The runtime module includes an AI basic service module, which implements traditional and generative AI capability services by invoking AI models, including basic natural language processing, vision processing, and speech processing services. The runtime module also includes an AI assistant service (named the OsAssistant service in this embodiment) corresponding to the AI assistant. Building on the AI basic capability service, the OsAssistant service provides a multimodal interface. It enables the completion of more complex tasks by utilizing the subsystem's built-in toolset, search enhancement generation services, and other services. It also manages conversation history.

[0024] like Figure 2 As shown, the runtime module in this embodiment includes a task scheduling module and a task queue, which can dynamically schedule and limit the resource usage of AI tasks according to the current system load status, thereby ensuring the smoothness of system operation. (1) The evaluation criteria for system fluency are mainly evaluated according to the response time, which refers to the time interval from the user's operation (such as clicking, sliding, etc.) to the system UI making corresponding feedback (such as interface switching, icon display, etc.). An excellent system UI response time should be within 100 milliseconds, so that the user can hardly feel the delay and the operation experience is smooth. If the response time exceeds 300 milliseconds, the user may notice the delay and the fluency experience will decrease. In step S103 of this embodiment, before the runtime module responds to the service request of the AI application, it includes scheduling for the service request: S201, obtaining the current load status of the system. If the system is currently in a high load state, generating a response result rejecting the service request of the AI application and jumping to step S104; otherwise, jumping to step S202; S202: Put the service request of the AI application into the task queue, and take out the service request from the task queue and execute it according to the task priority and system resource load through the task scheduling module to respond to the service request of the AI application.

[0025] As an optional implementation, when the task queue space is limited, it is possible to add a judgment on whether the task queue is empty after jumping to step S202 in step S201 and before executing step S202. If the task queue is empty, jump to step S202; otherwise, wait for a specified time and then continue to judge whether the task queue is empty. When waiting for a specified time, the waiting time can be set according to the demand time of the service request, so that the longer the demand time of the service request, the longer the waiting time, so that the service request with a shorter demand time will enter the task queue first.

[0026] In step S201 of this embodiment, obtaining the current load status of the system includes: S301. Obtain the load status of a specified load type through a system interface provided by the system API center of the runtime module. The specified load type includes some or all of CPU load, memory load, disk load, and GPU load. The following applies: ① CPU load: AI tasks typically involve a large number of computational operations, such as model inference and data processing. When the CPU load is too high, it means that the CPU needs to handle too many tasks simultaneously. This reduces the computing resources (such as CPU time slices) allocated to each task, affecting system responsiveness, causing operational lag, and reducing the smoothness of the user experience. ② Memory load: When running, AI applications need to load model data, intermediate calculation results, etc. into memory. Excessive memory load may lead to insufficient memory, causing the system to frequently perform memory swap operations. This operation significantly reduces data read and write speeds because the data exchange speed between memory and disk is much lower than the internal read and write speed of memory. This affects system responsiveness, causing operational lag, and reducing the smoothness of the user experience. ③ Disk I / O load: AI tasks may frequently read model files, training data, and write intermediate results and logs. When the I / O load is too high, a large number of processes may be blocked waiting for I / O operations to complete. For example, if a program needs to read data from the hard drive, if the hard drive I / O load is high, data read speed will slow down, and the program will be stuck in a waiting state, unable to execute subsequent instructions. This will reduce the overall system responsiveness, causing users to experience system lag and unsmooth operation. To process I / O requests, the CPU needs to expend time and energy on operations such as interrupt handling and data transfer. When the I / O load is high, the CPU needs to frequently respond to I / O interrupts, which increases CPU overhead. If the CPU spends most of its time processing I / O-related tasks, it has less time for other critical tasks (such as application logic operations and graphics rendering), resulting in a decrease in overall system performance and impacting smoothness. ④ GPU Load: In AI applications, the GPU handles a large number of parallel computing tasks, especially the training and inference of deep learning models. The GPU is primarily responsible for graphics rendering. When the load is too high, it slows down its processing of graphics data. This can cause screen freezes and frame drops, as the GPU cannot send generated image frames to the display in a timely manner. Resource monitoring functionality has been added to the task management submodule within the runtime module. Obtain real-time data such as CPU usage, memory usage, disk I / O read / write rates, and GPU load through the system interface. For example, collect data at regular intervals (such as 1 second) to form time series data on resource usage; S302: Determine the corresponding load level for each load type based on its preset grading threshold, where the load level includes a high load level. If more than a preset number of load types have a high load level, the system is determined to be in a high load state; otherwise, the system is determined to be in a non-high load state. For example, load thresholds for different resources may be set, such as CPU usage exceeding 80% being considered high load, memory usage exceeding 90% being considered high load, and so on. If multiple resources reach a high load state simultaneously, the entire system is determined to be in a high load state.

[0027] The task scheduling strategy in this embodiment includes: 1) Priority scheduling: assigning priorities to tasks based on their importance and urgency. For example, the AI task created by the application that the current user is operating has the highest priority, while the AI task created by the application that the user has not operated for a long time has a relatively low priority. 2) Flexible resource allocation: dynamically adjust the resources allocated to tasks according to the load status. When the system load is low, the AI tasks will be subject to lower resource restrictions and higher allocations according to their priority. When the system load is high, the AI tasks will be subject to higher resource restrictions and lower allocations according to their priority. 3) Task queue management: The task management submodule maintains a task queue. When the system load is too high, the submission of new tasks is restricted and they are temporarily stored in the queue. At the same time, tasks are taken out of the queue for execution in sequence according to the task priority and the system resource load, to avoid the system crashing due to too many tasks and to ensure the overall smoothness of the system.

[0028] In step S202 of this embodiment, when the task scheduling module takes out the service request in the task queue and executes it in response to the service request of the AI application according to the task priority and the system resource load, it includes periodically obtaining the service requests of the AI model on the calling side and the cloud in the task queue according to the preset scheduling time window, scheduling the service requests of the AI model on the calling side and the cloud in a first-in-first-out manner, and only taking out and executing one service request of the AI model on the calling side and the cloud in each scheduling time window; it includes obtaining the service request for calling the local AI model in the task queue according to the preset scheduling period, calculating the scheduling priority of the service request for calling the local AI model according to the task priority and the system resource load carried by the service request for calling the local AI model, and taking out and executing a service request of the local AI model with the highest scheduling priority according to the scheduling priority in each scheduling period.

[0029] When calculating the scheduling priority of a service request that calls a local AI model based on the task priority it carries and the system resource load, a feasible calculation method can be adopted as needed. For example, as an optional implementation method, it is based on the following goals: (1) Static evaluation before scheduling: does not rely on runtime indicators (such as actual memory usage, execution time), but is only based on the requirements of the task declaration and the current resource status of the system. (2) Multi-resource collaboration: considers the supply and demand relationship of resources such as CPU, memory, and graphics card at the same time. (3) Priority retention: even if high-priority tasks have large resource requirements, they should not be completely suppressed. In this embodiment, the function expression for calculating the scheduling priority of a service request that calls a local AI model is: P1= P0×R avail / R task , Among them, P1 is the scheduling priority, P0 is the task priority carried by the service request, R avail is the current resource availability of the system, R task The task resource requirements carried by the service request, and have: R avail = w cpu ×C avail / C total + w mem ×M avail / M total + w gpu ×G avail / G total , R task =( w cpu ×C demand / C total + w mem ×M demand / M total + w gpu ×G demand / G total ) 1 / 2 , in, w cpu 、 w mem and w gpu are the weights of CPU, memory and graphics card respectively, C avail 、C demand and C total They are the available, required and total amount of CPU respectively, M avail 、Mdemand and M total They are the available amount, required amount and total amount of memory respectively, G avail , G demand and G total are the available, required and total GPU capacity of the graphics card respectively. Among them, the task resource demand R carried by the service request task In the calculation of , taking the square root can prevent high-demand tasks from being overly suppressed. The above scheduling priority function expression statically compares the task resource requirements with the system's currently available resources, determining the priority before scheduling. This has the following advantages: a. No runtime metrics are required: It relies only on the task's declared requirements and the system snapshot status. b. Multi-resource collaboration: It balances the supply and demand of CPU, memory, and graphics cards. Priority protection: It ensures the fundamental advantages of high-priority tasks through multiplication design. c. Flexibility: It can adapt to different scenarios through weight and nonlinear adjustment.

[0030] like Figure 2 As shown, the runtime module in this embodiment also includes a privacy data management module, which is used to encrypt and store the user data involved in the service request when the runtime module responds to the service request of the AI application, and decrypt and read the user data when the runtime module needs to read it.

[0031] like Figure 2 As shown, the runtime module in this embodiment also includes a context management module, which is used to manage historical messages (context) of conversations with the AI model when the runtime module responds to service requests from the AI application. As an optional implementation, the context management module supports automatic and manual management modes. If the context management mode is automatic, the context with the AI model is managed by the context management module; if the context management mode is manual, the context with the AI model is managed by the service request itself. The context management module provides the function of managing the conversation context, simplifying the complexity of the application user interface. However, the caller still retains the ability to manage the context.

[0032] like Figure 2 As shown, the runtime module in this embodiment also includes a prompt word configuration management module. The prompt word configuration management module is used to generate a preset prompt word for the service request when the runtime module responds to the service request of the AI application. When the AI model on the end or cloud is called to generate a response result for the service request, the prompt word and the service request are combined as input to the AI model on the end or cloud to generate a response result for the service request. The prompt word configuration management module has built-in prompt words for multiple scenarios, and new prompt words can also be created.

[0033] like Figure 2As shown, the runtime module in this embodiment also includes an AI engine selection module. The AI engine selection module is used to first screen out candidate AI models from the AI models on the end and cloud according to the computing power requirements of the service request when calling the AI model on the end or cloud to generate a response result of the service request and return it to the communication layer module. Then, according to the priority settings of each AI model in the candidate AI models, the AI model with the highest priority is selected, and the AI model with the highest priority is called to generate a response result of the service request and return it to the communication layer module. The runtime module in this embodiment also includes a system API center for managing the API interfaces of the AI models on the end and cloud.

[0034] like Figure 2 As shown, the runtime module in this embodiment also includes a toolset, which includes a picture viewing plug-in, an instruction plug-in, a retrieval-augmented generation (RAG) service and a system audio service. The picture viewing plug-in is used to display and output the generated image when generating the image; the instruction plug-in is used to provide a set of instructions supported by the runtime module for user selection; the retrieval-augmented generation service is used to generate a knowledge graph from the document input by the user for AI model call; the system audio service is used to provide system audio input and output functions for speech processing Speech.

[0035] like Figure 3 As shown in the figure, as an optional implementation, in this embodiment, after the application calls the interface module, the interface module sends the request information to the runtime module through the communication layer module for processing. Each interface has a server-side proxy module (Proxy) that establishes a link with the runtime. The interface can communicate with the runtime through the server-side proxy module. For example, a public interface accesses the runtime module through the communication layer module via service proxy 1, while a private interface accesses the runtime module through the communication layer module via service proxy 2.

[0036] In summary, the general support method for AI applications under the Linux system of this embodiment proposes the idea of a unified AI standard interface, shielding the specific implementation of each manufacturer. Developers do not need to pay attention to the specific underlying implementation, lowering the threshold for application integration of AI capabilities. Application developers only need to adapt a set of interfaces. The method of this embodiment proposes the idea of dividing the AI interface into internal interfaces and external interfaces. The internal interface can be quickly iterated and changed according to market demand to enhance the competitiveness of the operating system. The external interface provides basic capabilities, and third-party developers can quickly integrate AI capabilities. The method of this embodiment proposes the idea of separating the interface and implementation. The interface layer only provides the interface to the outside world, and the specific implementation is implemented by the runtime. The method of this embodiment proposes the idea of dynamically scheduling AI-related tasks according to the current load status of the system, efficiently and reasonably allocating system resources, and ensuring the AI intelligent user experience while ensuring the smoothness of the overall system.

[0037] In addition, this embodiment also provides a general support system for AI applications under a Linux system, including a microprocessor and a memory connected to each other, wherein the microprocessor is programmed or configured to execute the general support method for AI applications under a Linux system.

[0038] In addition, this embodiment also provides a computer-readable storage medium, which stores a computer program or instruction, and the computer program or instruction is programmed or configured to execute the general support method for AI applications under the Linux system through a processor.

[0039] In addition, this embodiment also provides a computer program product, including a computer program or instructions, which are programmed or configured to execute the general support method for AI applications under the Linux system through a processor.

[0040] Those skilled in the art should understand that the technical solution provided by the present invention may be in the form of a method, a system, or a computer program product. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The present invention is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, may be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the functions described in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including the instruction device, which implements the function specified in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0041] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A general support method for AI applications in Linux system, characterized by: include: S101: The interface module receives a service request from an AI application. When the service request is valid, the interface module forwards the service request to the communication layer module. The interface module includes a public interface and a private interface. The interface types provided by the public interface include natural language processing, visual processing, speech processing, text generation, and image generation. The private interface is embedded with an interface authentication module, a model configuration module, and an AI assistant. The interface authentication module is used to authenticate the service request. The model configuration module is used to convert the service request into a configuration service request for the AI model. The AI assistant is used to provide natural language processing, visual processing, speech processing, text generation, and image generation services, as well as interaction with the operating system. The model configuration module and the AI assistant execute the service request only after authentication is passed. S102: The communication layer module establishes a connection between the AI application and the runtime module, and transmits the service request to the runtime module; S103: The runtime module responds to the service request of the AI application. If the service request calls a public interface, the runtime module calls the AI model on the client or cloud side to generate a response to the service request and returns it to the communication layer module. If the service request is a private interface call, then when the service request is a configuration service request, the configuration of the AI model on the other side or in the cloud is executed, and a configuration response result is generated and returned to the communication layer module; when the service request is a service request for an AI assistant, natural language processing, visual processing, speech processing, text generation, and image generation services are provided, or interaction with the operating system is performed, and an interaction response result is generated and returned to the communication layer module; S104: The communication layer module returns the response result of the runtime module to the AI application that initiated the service request.

2. The general support method for AI applications under Linux system according to claim 1, characterized in that: The runtime module includes a task scheduling module and a task queue. Before the runtime module responds to the service request of the AI application in step S103, it schedules the service request: S201, obtaining the current load status of the system. If the system is currently in a high load state, generating a response result rejecting the service request of the AI application, and jumping to step S104; Otherwise, jump to step S202; S202: Put the service request of the AI application into the task queue, and take out the service request from the task queue and execute it according to the task priority and system resource load through the task scheduling module to respond to the service request of the AI application.

3. The general support method for AI applications in Linux system according to claim 2, characterized in that: Acquiring the current load status of the system in step S201 includes: S301, obtaining a load status of a specified load type through a system interface provided by a runtime module, where the specified load type includes part or all of CPU load, memory load, disk load, and GPU load; S302, the load state of each load type is determined to have a corresponding load level according to its preset grading threshold, wherein the load level includes a high load level; if the load level of more than a preset number of load types is a high load level, the system is determined to be in a high load state, otherwise the system is determined to be in a non-high load state.

4. The general support method for AI applications in Linux system according to claim 2, characterized in that: In step S202, when the task scheduling module takes out the service request in the task queue and executes it in response to the service request of the AI application according to the task priority and the system resource load, it includes periodically obtaining the service requests of the AI model on the calling side and the cloud in the task queue according to the preset scheduling time window, scheduling the service requests of the AI model on the calling side and the cloud in a first-in-first-out manner, and only one service request of the AI model on the calling side and the cloud can be taken out and executed in each scheduling time window; it includes obtaining the service request for calling the local AI model in the task queue according to the preset scheduling period, calculating the scheduling priority of the service request for calling the local AI model according to the task priority and the system resource load carried by the service request for calling the local AI model, and taking out and executing a service request of the local AI model with the highest scheduling priority according to the scheduling priority in each scheduling period.

5. The general support method for AI applications in Linux system according to claim 1, characterized in that: The runtime module also includes a privacy data management module, which is used to encrypt and store user data involved in a service request when the runtime module responds to a service request from an AI application, and decrypt and read the data when the runtime module needs to read it; the runtime module also includes a context management module for managing historical messages of conversations with the AI model.

6. The general support method for AI applications in Linux system according to claim 1, characterized in that: The runtime module also includes a prompt word configuration management module, which is used to generate a preset prompt word for the service request when the runtime module responds to the service request of the AI application. When the AI model on the terminal side or the cloud is called to generate a response result of the service request, the prompt word and the service request are combined as inputs of the AI model on the terminal side or the cloud to generate a response result of the service request; the runtime module also includes an AI engine selection module, which is used to first screen out candidate AI models from the AI models on the terminal side and the cloud according to the computing power requirements of the service request when the AI model on the terminal side or the cloud is called to generate a response result of the service request and return it to the communication layer module, and then select the AI model with the highest priority according to the priority settings of each AI model in the candidate AI models, and call the AI model with the highest priority to generate a response result of the service request and return it to the communication layer module.

7. The general support method for AI applications in Linux system according to claim 1, characterized in that: The runtime module also includes a tool set, which includes a picture viewing plug-in, an instruction plug-in, a search enhancement generation service, and a system audio service. The picture viewing plug-in is used to display and output the generated image when generating the image; the instruction plug-in is used to provide a set of instructions supported by the runtime module for user selection; The retrieval enhancement generation service is used to generate a knowledge graph from the documents input by the user for AI model call; the system audio service is used to provide system audio input and output functions for speech processing.

8. A general support system for AI applications under Linux system, comprising a microprocessor and a memory connected to each other, characterized in that: The microprocessor is programmed or configured to execute the general support method for AI applications under the Linux system as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program or instruction stored therein, characterized in that: The computer program or instruction is programmed or configured to execute, through a processor, the general support method for AI applications under a Linux system as recited in any one of claims 1 to 7.

10. A computer program product comprising a computer program or instructions, characterized in that The computer program or instruction is programmed or configured to execute, through a processor, the general support method for AI applications under a Linux system as recited in any one of claims 1 to 7.