A method and system for implementing operating system AI processing capabilities

By integrating AI model capabilities into the operating system, the limitations of traditional operating systems in terms of application scope and data privacy are solved. This enables the integration and natural interaction of AI capabilities across multiple systems, supports flexible adaptation and local processing of multimodal AI models, and enhances the intelligent task processing capabilities of devices.

CN120821469BActive Publication Date: 2025-12-05ZHIDA CHENGYUAN TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511333377.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2025-12-05
Estimated Expiration
2045-09-18

AI Technical Summary

Technical Problem

Traditional operating systems struggle to meet the demands of AI-era systems for intelligence, personalized services, and natural interaction. They have limited application scope, lack the flexibility to adapt to AI models, and cloud-based large-scale model solutions present data privacy issues.

Method used

The system integrates AI model capabilities into the operating system, supports plug-in design of various AI models through a logic processing engine and AI adapter, enables local offline processing, and combines sensor data for task analysis and execution.

Benefits of technology

It achieves cross-platform and multi-system AI capability integration, supports multimodal AI model invocation, ensures data privacy and provides flexible model management, and enhances the device's intelligent task processing and natural interaction capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120821469B_ABST
    Figure CN120821469B_ABST
Patent Text Reader

Abstract

The application discloses a kind of method and system for realizing operating system AI processing capacity, method includes: the user instruction captured is sent to the intent engine in logic processing engine by variable definition to AI adapter, and variable;The intent engine in logic processing engine is sent to AI model processing by corresponding processing generation inference flow to prompt word engine in logic processing engine;When AI model processing executes the inference flow, call corresponding AI model to reason;According to the inference result, generate task flow, and each execution flow in task flow is executed by corresponding task execution module and tool to call AI adapter.This application adds AI capacity to the OS of traditional terminal, supports multiple device platforms and multiple operating systems;Support language class, visual class, audio class, database index and multiple AI models;Support AI model plug-in flexible adaptation;With end side model as main as possible to avoid data privacy problem.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of operating systems, in particular to a method and system for realizing AI processing capability of an operating system, a computing device and a storage medium. BACKGROUND

[0002] With the rapid development of artificial intelligence (AI), especially the development of AI large models, the traditional operating systems (OS) of mobile phones, vehicle machines and other embedded devices, which take hardware resource management and application scheduling as core functions, have been difficult to meet the new demands for system intelligence, service personalization and natural interaction in the AI era.

[0003] Currently, a few technologies for adding AI capability to operating systems (OS) are mainly applied to computers or only enhance the local AI function of a single platform, have a single application range and are mainly applied to a relatively single platform or device, or only have a single LLM (Large Language Model) capability. At the same time, there is a lack of AI model adaptation flexibility, because the current technology is usually strongly coupled with AI large models, and thus lacks flexibility in updating AI models and capabilities. In addition, the existing technology is mainly based on a cloud model scheme, and the data privacy problem of the cloud model leads to a lack of data security. SUMMARY

[0004] One of the purposes of the embodiments of the present application is to address the deficiencies of the existing technology, and to provide a method and system for realizing AI processing capability of an operating system, a computing device and a storage medium, which provide a method for adding AI capability to an operating system (OS) for multiple systems such as Linux, Android, QNX, etc. of mobile phones, vehicle machines and other embedded devices, and give the OS more intelligent, more efficient and more natural interaction capabilities by integrating AI model capabilities.

[0005] To solve the above technical problems, in a first aspect, the embodiments of the present application provide a method for realizing AI processing capability of an operating system, which comprises:

[0006] defining variables of the captured user instructions through an AI adapter, and sending the variables to an intent engine in a logic processing engine;

[0007] The intent engine of the logic processing engine generates an inference flow through a prompt word engine in the logic processing engine, and sends the inference flow to AI model processing;

[0008] When the AI model processing executes the inference flow, it calls a corresponding AI model for inference;

[0009] The task flow is generated according to the inference result, and each execution flow in the task flow calls a corresponding task execution module through an AI adapter to perform execution.

[0010] Preferably, the AI model processing for executing the inference flow specifically includes:

[0011] The encapsulation plug-in of the AI model processing includes a session management interface and a model inference interface.

[0012] A model session is created through the session management, and the inference task is performed on the input inference flow through the model inference interface.

[0013] The inference result is transferred to the prompt word engine of the logic processing engine.

[0014] Preferably, the inference task performed on the input inference flow through the model inference interface specifically includes:

[0015] The inference flow is taken as the input of the AI model processing through the prompt word engine interface of the corresponding logic processing engine.

[0016] The AI model processing performs model inference, completes inference through the model inference interface, and generates a command.

[0017] The generated command is returned to the intent engine of the logic processing engine through the prompt word engine.

[0018] Preferably, the task flow is generated according to the inference result, and each execution flow in the task flow calls a corresponding task execution module through an AI adapter to perform execution specifically includes:

[0019] After the intent engine of the logic processing engine receives the command, it is determined whether the command is executable by comparing the current command with a pre-registered support function whitelist, if yes, the command is taken as an action flow string and is transmitted to the agent engine of the logic processing engine for multi-task arrangement, a serial graph of an execution node is generated by the multi-task arrangement, each node is an execution flow, and each execution flow is executed through an AI adapter calling a corresponding module.

[0020] Preferably, each execution flow is executed through an AI adapter calling a corresponding module and tool specifically includes:

[0021] For a task requiring AI capability, the AI adapter calls the agent task processing.

[0022] The agent task processing calls a corresponding model of the AI model processing through a task prompt word transmitted to the prompt word engine, and a sensor engine required by the model inference interface is called, and the sensor engine calls an OS module to perform data collection of a sensor.

[0023] The model inference interface combines the data collected by the sensor and the task prompt word of the prompt word engine to perform inference to obtain an inference result;

[0024] The inference result is returned to the agent engine in the logic processing engine, and a task execution workflow is updated;

[0025] The updated task execution workflow is passed to the OS module and then returned to the agent task processing of the AI adapter, a registered fixed task interface is called, and the task execution is completed,

[0026] A flag indicating that the execution is completed is returned through the agent task processing of the AI adapter.

[0027] In a second aspect, to solve the technical problem of the present application, the embodiments of the present application also provide a system for realizing AI processing capability of an operating system, the system comprising:

[0028] An OS module, an AI model processing module, a logic processing module, and an AI adapter, wherein the OS module is connected to the AI model processing module and an agent engine submodule in the logic processing module, the AI model processing module is connected to a sensor engine submodule in the logic processing module and a prompt word engine submodule in the logic processing module, the prompt word engine submodule in the logic processing module is connected to an intent engine submodule in the logic processing module, and the AI adapter is connected to the sensor engine submodule in the logic processing module, the prompt word engine submodule in the logic processing module, and the intent engine submodule in the logic processing module, respectively;

[0029] The AI model processing module supports an AI large model, includes various model inference engines, and is encapsulated into a model inference interface.

[0030] The sensor engine submodule of the logic processing module is configured to schedule and manage the sensor.

[0031] The prompt word engine submodule of the logic processing module is configured to perform prompt word caching, prompt word splicing, prompt word triggered inference, and provide a result returned by the inference interface of the AI model processing module.

[0032] The intent engine submodule of the logic processing module is configured to convert input into a formatted text command, i.e., to feed the input into a corresponding model plug-in and return a formatted text command generated by the model.

[0033] The agent engine submodule of the logic processing module is configured to generate a serial graph of an execution node from the formatted command, wherein each node in the graph is an execution flow, and each execution flow is executed by calling a corresponding module through the AI adapter.

[0034] The AI adapter is used to perform the task of the agent function.

[0035] Further, the intent engine submodule of the logic processing module is also used to compare the pre-registered support function whitelist according to the input formatted text command, judge whether the task is executable, if executable, take the above formatted text command as an action flow string, and pass it to the agent engine of the logic processing module for multi-task arrangement.

[0036] In a third aspect, the embodiments further provide a computing device, comprising a processor and a memory, the memory being configured to store a computer program, the computer program comprising program instructions, and the processor being configured to invoke the program instructions to execute the method as described above.

[0037] In a fourth aspect, the embodiments further provide a computer readable storage medium, which stores a computer program, the computer program comprising program instructions, and the program instructions, when executed by a processor or a computer, causing the processor to execute the method as described above.

[0038] Compared with the prior art, the system and method for realizing the AI processing capability of an operating system provided by the embodiments have at least the following beneficial effects: solving the problem of single application range, supporting application on multiple operating systems of multiple device platforms, such as Linux, Android, and QNX systems of devices of Qualcomm, Mediatek, and Nvidia; solving the problem of single AI capability, supporting multiple types of AI model capabilities, such as LLM (Large Language Model), VLM (Vision-Language Model), ASR (Automatic Speech Recognition), TTS (Text-to-Speech), and RAG (Retrieval-Augmented Generation); solving the problem of AI model adaptation flexibility, the technical solution of the embodiments can use plug-in design to support flexible use or replacement of models when applying AI large models; solving the problem of data privacy of cloud models, the technical solution of the embodiments mainly uses local offline lightweight end-side models, which greatly avoids the data privacy problem of cloud models. BRIEF DESCRIPTION OF DRAWINGS

[0039] The above-mentioned features, technical characteristics, advantages, and implementation modes of the present application will be further described in the following preferred embodiments in a clear and easy-to-understand manner, combined with the accompanying drawings.

[0040] Figure 1 A system block diagram for implementing an operating system AI processing capability according to an embodiment of the present application;

[0041] Figure 2 A method flowchart for implementing an operating system AI processing capability according to an embodiment of the present application;

[0042] Figure 3 A computing device structure schematic diagram according to an embodiment of the present application. DETAILED DESCRIPTION

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, specific embodiments of the present application will be described below with reference to the drawings. Obviously, the drawings in the following description only represent some embodiments of the present application, and for those skilled in the art, other drawings can be obtained from these drawings without creative labor, and other embodiments can also be obtained.

[0044] In order to make the drawing simple, only the parts related to the present application are shown in each drawing, which does not represent the actual structure of the product. In addition, in order to make the drawing simple and easy to understand, in some drawings, only one of the components with the same structure or function is shown, or only one of them is marked. In the embodiments of the present application, "one" not only means "only one", but also means "more than one". The embodiments of the technical solutions of the present application will be described in detail below mainly by taking some specific embodiments as examples.

[0045] As shown in Figure 1 To solve the technical problems of the present application, the embodiments of the present application also provide a system for implementing an operating system AI processing capability, the system comprising:

[0046] An OS module, an AI model processing module, a logic processing module and an AI adapter, wherein the OS module is connected with an agent engine submodule in the AI model processing module and the logic processing module, the AI model processing module is connected with a sensor engine submodule in the logic processing module and a prompt word engine submodule in the logic processing module, the prompt word engine submodule in the logic processing module is connected with an intent engine submodule in the logic processing module, and the AI adapter is connected with the sensor engine submodule in the logic processing module, the prompt word engine submodule in the logic processing module and the intent engine submodule in the logic processing module, respectively;

[0047] The AI model processing module supports an AI large model, includes various model inference engines and encapsulates into a model inference interface.

[0048] The sensor engine submodule of the logic processing module is used for scheduling and managing sensors.

[0049] The prompt word engine submodule of the logic processing module is used for prompt word caching, prompt word splicing, prompt word trigger reasoning, and returning the result returned by the inference interface of the AI model processing module.

[0050] The intent engine submodule of the logic processing module is used for converting input into a formatted text command, i.e., feeding the input into a corresponding model plug-in, and returning a formatted text command generated by the model.

[0051] The agent engine submodule of the logic processing module is used for generating a serial graph of execution nodes from the formatted command, where each node in the graph is an execution flow, and each execution flow is executed by calling a corresponding module through an AI adapter.

[0052] The AI adapter is used to perform tasks of the agent function.

[0053] Further, the intent engine submodule of the logic processing module is also used for comparing a pre-registered support function whitelist according to the input formatted text command to determine whether the task is executable, and if so, the formatted text command is transmitted to the agent engine of the logic processing module as an action flow string for multi-task arrangement.

[0054] The following embodiments are used to introduce the main technical solutions of the present application, based on the description of the accompanying drawings Figure 1 The following embodiments are used to introduce the main technical solutions of the present application, based on the description of the accompanying drawings

[0055] For the OS module:

[0056] The corresponding hardware devices are multiplexed, such as Qualcomm, Mediatek, other MCUs / SOCs, and OS systems such as Linux, Android, and QNX. Then, a C++ interface for sensor calling is provided, mainly through C++ open source libraries such as OpenCV vision library and SDL2 multimedia library, to realize the calling of sensor drivers, send hardware sensor calling requests to the OS, and realize the calling of sensor hardware.

[0057] For the AI model processing module:

[0058] (1) The AI model processing module supports AI large models (such as LLM, VLM, etc.) and model inference engines (such as llama.cpp, whisper.cpp, etc.), and rewrites the inference engine and encapsulates it into a new inference interface.

[0059] For example, based on the llama.cpp inference engine, encapsulate new C++ inference interfaces, such as the following main interfaces:

[0060] platform_model_plugin::Create(EngineConfig config) is used to read the configuration file to prepare to create an inference environment;

[0061] platform_model_plugin::Inference(InferenceData& data, Error** err, std::string &result) is used to call the interface implemented in llama.cpp to perform model inference and return the inference result by passing inference parameters;

[0062] platform_model_plugin::InferenceAsync(InferenceData& data) is also used to call the interface implemented in llama.cpp to perform model inference, and the result is returned in real time using a callback function.

[0063] The underlying calls to sensor C++ open source libraries (such as OpenCV, SDL2), calls to related processor backends (such as CPU / NPU / GPU), and calls to related system libraries (such as libc++.so, libm.so) and other work are all implemented in the inference engine llama.cpp, and are directly called when the model inference is called.

[0064] (2) Then implement an abstraction layer in C++, which can call different models using the same function interface, and implement functions such as model inference, plugin management, and session management for model inference.

[0065] For example, the main interfaces for plugin management are:

[0066] createPlugin() creates a plugin object;

[0067] releasePlugin() deletes the plugin object.

[0068] For model inference main interfaces:

[0069] Create, Inference, InferenceAsync, and other interfaces are abstracted to directly call the corresponding plugin interfaces in platform_model_plugin.

[0070] For session management main interfaces:

[0071] createSession(AiModelInfo& modelInfo, int sessionId, std::stringsessionName) creates a thread invoking model inference interface by reading in configuration and session id and name;

[0072] releaseSession(int sessionId) releases inference thread by session id.

[0073] Encapsulate the above as a model plugin so, such as lib_whisper_plugin.so.

[0074] Sensor engine submodule for logical processing module

[0075] Based on sensor C++ open source library (such as OpenCV, SDL2, ASLA), write C++ function interface to dispatch and manage sensors, implement sensor device identification and selection, sensor opening and closing pause, sensor data input and output, etc. When the sensor engine is called, according to the task type passed in (such as language class, vision class, audio class), send a hardware sensor call request to the OS and apply for the required sensor resources.

[0076] Prompt word engine submodule for logical processing module

[0077] Through C++ code, mainly implement prompt word cache, prompt word splicing, prompt word trigger inference, get inference return function interface.

[0078] For example, CachePrompt(const std::string& prompt) appends the user or system prompt to the internal cache std::vector;

[0079] BuildPrompt() splices several prompt words into a complete prompt word in order using C++ string method;

[0080] TriggerInference((InferenceData& data, Error** err,) gets prompt and input data as data parameter, then calls the inference interface in the model plugin for inference;

[0081] GetInferenceResult() returns the inference result through the callback function, or takes out the result from the internal buffer and returns it.

[0082] Intention engine submodule for logical processing module

[0083] Through C++ code, the input-to-formatted text command, inference invocation, and rationality evaluation function interfaces are mainly implemented. For example:

[0084] Input-to-formatted text command interface: InputToCommand(const InputData& raw_input) determines whether the input data type is a simple input stream (such as a text stream, audio stream) or a specified format text command based on the input data type. If it is a simple input stream, it will call the fixed inference process to convert the input into a specified format text command (such as " [{id:1,type:ai,tool:vlm_model,task: scene_mood}, {…}, …] "). For example, a text stream calls LLM, an audio stream calls ASR+LLM, and a video stream calls VLM+LLM to convert the input stream into a specified format text command.

[0085] Inference invocation interface: InvokeModel(const std::string& prompt, ModelType model) is called by InputToCommand, which feeds the prompt / audio features / video frames to the corresponding model plugin based on the input type judgment result, and returns the formatted text command generated by the model.

[0086] Rationality evaluation interface: ValidateCommand(const std::string& json_cmd, std::string& registry_cmd) compares the input formatted text command with the pre-registered support function whitelist registry_cmd to evaluate whether the task is executable. If it is executable, the above formatted command is used as an action flow string, which is passed to the agent engine for multi-task scheduling.

[0087] Agent engine sub-module for logical processing module

[0088] Draw a serial graph of execution nodes using Taskflow for formatted commands; each node is a call to an external tool execution flow, and each execution flow is executed by calling the corresponding module and tool through an AI adapter.

[0089] Through C++ code, the LLM inference task graph, generated execution task flow, and execution graph function interfaces are mainly implemented.

[0090] (1) LLM inference task graph: InferTaskGraph(const std::string& json_cmd, const std::vector <toolmeta>& tools) send user raw formatted command and system prompt (including available tool list, constraint rule) to LLM, let LLM directly return task graph describing task execution graph (node, edge, condition, parallel relationship). Here, the direct call to generate execution task flow interface corresponds to simple process below 3 steps.

[0091] (2) Generate execution task flow: parse the task graph returned by LLM, create tf::Task for each node using taskflow library, and establish dependency or branch according to edge;

[0092] (3) Execution graph: RunGraph(tf::Taskflow& g, Context& ctx) gives the Taskflow generated in the last step to tf::Executor for running; nodes dynamically call external tools through AI adapter during execution, and the results are written back to Context.

[0093] For AI adapter

[0094] (1) AI Client SDK is responsible for tasks with traditional agentless functions. It is equivalent to directly storing function registry, and directly calling OS functions without AI. The main function interface implementation is as follows:

[0095] RegisterNormalTask(const std::string& tool, const std::string& task, std::function<int(const Json::Value&, std::string&)> func) registers the task function pointer of type:normal into the global table, and directly uses the tool name and task name to get std::function from the table and execute it when calling;

[0096] ExecuteNormal(const Json::Value& cmd, std::string& result) formats the text command when type==normal, AI adapter directly calls this interface, internally looks up the table, calls the specific function, and writes the result back to result.

[0097] (2) AI Agentic SDK is responsible for tasks that require agent functions. It will first dlopen model plug-in, then call sensor engine to get data, send prompt words and data to model for inference, and finally update workflow. The main function interface implementation is as follows:

[0098] LoadPlugin(const std::string& plugin_path) original dlopen corresponding so.

[0099] CaptureSensor(const std::string& sensor_type, const std::string&task_hint) according to the task description, trigger the corresponding sensor (camera / microphone, etc.), return the sensor data or file path;

[0100] Infer(void* plugin_handle, const std::string& prompt, const Json::Value& data) put the prompt and data into the loaded model plug-in, return the model inference result;

[0101] UpdateWorkflow(const std::string& new_cmd_json) according to the inference result, generate the next one or more commands, directly update to the workflow queue of the agent engine.

[0102] As Figure 2 shown, in order to achieve the purpose of the present application, the method for realizing the AI processing capability of the operating system provided by the embodiment of the present application comprises:

[0103] S1, the captured user instruction is subjected to variable definition through an AI adapter, and the variable is sent to an intent engine in a logic processing engine;

[0104] S2, the intent engine of the logic processing is subjected to corresponding processing through a prompt engine in the logic processing engine to generate an inference flow, and the inference flow is sent to an AI model processing module;

[0105] S3, when the AI model processing executes the inference flow, a corresponding AI model is called to perform inference;

[0106] S4, a task flow is generated according to the inference result, and each execution flow in the task flow is executed through an AI adapter calling a corresponding task execution module.

[0107] Preferably, when the AI model processing executes the inference flow, the corresponding AI model is called to perform inference, and the method specifically comprises:

[0108] calling a packaging plug-in of the AI model processing, wherein the packaging plug-in comprises a session management interface and a model inference interface;

[0109] A model session is created through the session management, and a reasoning task is performed on an input reasoning flow through the model inference interface;

[0110] The reasoning result is converted to a prompt word engine of the logic processing engine.

[0111] Preferably, the performing of the reasoning task on the input reasoning flow through the model inference interface specifically comprises:

[0112] The reasoning flow is taken as an input of the AI model processing through a prompt word engine interface of the corresponding logic processing engine;

[0113] The AI model processing performs model inference, completes reasoning through a model inference interface, and generates a command;

[0114] The generated command is returned to an intent engine in the logic processing engine through the prompt word engine.

[0115] Preferably, the generating of the task flow according to the reasoning result, and the performing of each execution flow in the task flow through an AI adapter to call a corresponding task execution module specifically comprises:

[0116] The intent engine of the logic processing engine receives the command, judges whether the command is executable by comparing the command with a pre-registered support function whitelist, and if yes, passes the command as an action flow string to an agent engine of the logic processing engine to perform multi-task arrangement, the multi-task arrangement generates a serial graph of an execution node, each node is an execution flow, and each execution flow is executed through an AI adapter to call a corresponding module.

[0117] Preferably, the performing of each execution flow through the AI adapter to call the corresponding module and tool specifically comprises:

[0118] For a task requiring AI capability, the AI adapter calls an agent task processing;

[0119] The agent task processing calls a corresponding model of the AI model processing through a task prompt word passed to the prompt word engine, a model inference interface calls a required sensor engine, and the sensor engine calls an OS module to perform data collection of a sensor;

[0120] The model inference interface performs reasoning in combination with data collected by the sensor and a task prompt word of the prompt word engine, and obtains a reasoning result;

[0121] The reasoning result is returned to the agent engine in the logic processing engine, and a task execution flow is updated;

[0122] The updated task execution workflow is passed to the OS module and then returned to the agent task processing of the AI adapter, a registered fixed task interface is called, and the task execution is completed,

[0123] A flag indicating execution completion is returned through the agent task processing of the AI adapter.

[0124] Next, the present application combines specific application scenarios to as far as possible to show the mutual interaction process between the various modules of the present application. In fact, the technical solution of the present application has a wide range of applications, so this application scenario is only an example, which aims to illustrate the interaction between the various modules of the present application technical solution.

[0125] Scenario description: the user issues a voice instruction "please play suitable music according to the current environmental atmosphere", hoping to run the system equipped with the present application to realize the capture of images through the camera, the inference of the atmosphere label using the VLM model, and the search and play of music of the same label type. The specific process is as follows:

[0126] (1) Capture voice instruction: the system OS module calls the microphone hardware device to capture the voice input instruction "please play suitable music according to the current environmental atmosphere" through the microphone software library, such as through the ALSA (Advanced Linux Sound Architecture) audio API, and passes the audio stream to the intent engine as input in the form of defining audio variables in the AI adapter.

[0127] (2) Voice instruction processing flow: the intent engine provides a C++ interface to determine whether the task input is a segment of independent audio type data, and then trigger the pre-solidified inference flow of voice-to-text command, that is, through C++ code implementation, input the prompt word string of the inference flow of ASR+LLM model to the prompt word engine, and the obtained audio data, and then pass the above data to the AI model processing module for inference to execute the conversion of audio input to specified format text command.

[0128] (3) ASR voice-to-text:

[0129] When the AI model processing module executes the inference flow of the voice-to-text command, it receives the audio and prompt words such as "convert the audio stream to text" transmitted by the prompt word engine through C++ variables, calls the ASR model such as whisper model for inference;

[0130] The whisper inference is performed by calling the model encapsulation plug-in lib_whisper_plugin.so through dlopen. The lib_whisper_plugin.so is based on the inference engine such as whisper.cpp to perform API modification and finally provide external session management interfaces (such as createSession, releaseSession, etc.), plug-in management interfaces (such as createPlugin, releasePlugin, etc.), model inference interfaces (such as Create, Inference, InferenceAsync, etc.), and other APIs.

[0131] When the whisper model is used, the API is called to create a model session and load a whisper model plug-in, then the inference interface is called to input the obtained prompt word and audio stream to perform the inference task, and finally the inference result text "Please play suitable music according to the current environmental atmosphere" is returned to the prompt word engine.

[0132] (4) LLM text conversion format command: the prompt word engine takes the text returned by the ASR inference as input and transmits it to the AI model processing module to perform LLM such as Qwen model inference. Through the prompt word engine C++ interface, the text data "Please play suitable music according to the current environmental atmosphere" and the prompt word "text conversion format text command" are transmitted to the Qwen model for inference. Similarly, through dlopen, the C++ API in the model encapsulation plug-in lib_qwen_plugin.so is called to complete the inference of the text conversion format text command, such as "1. Take a photo VLM to identify the environmental atmosphere (such as xx); 2. Search for xx type music and play". The corresponding formatted text command is "[{id:1, type:ai, tool:vlm_model, task: scene_mood}, {id:2, type:normal, tool:music_player, task:play_music_{{step1.scene_mood}}}]". Then the formatted text command is returned to the intent engine through the prompt word engine.

[0133] (4) Text command intent judgment: the intent engine takes the formatted text command and judges whether it can be executed by comparing the current command task operation with the pre-registered support function whitelist. If it can be executed, the above formatted command is transmitted to the agent engine as an action flow string for multi-task scheduling.

[0134] (5) Task intelligent arrangement: draw a serial graph of execution nodes with Taskflow for formatted commands; each node is a call to an external tool, and each execution flow is executed by calling the corresponding module and tool through the AI adapter.

[0135] (6) Task execution: after the task table is arranged, the agent engine will pass it to the OS module as a workflow for execution. For tasks that do not require AI capabilities, call the AI Client SDK in the AI adapter for execution. Determine the type value of the formatted command. For tasks that do not require AI capabilities, call the AI Client SDK in the AI adapter for execution. For tasks that require AI capabilities, call the AI Agentic SDK in the AI adapter for execution.

[0136] (7) Execute task 1 - take a picture VLM to identify the environment atmosphere: the AI adapter passes the task "{id:1, type:ai, tool:vlm_model, task: scene_mood}" to the AI Agentic SDK for execution through the type value. The AI Agentic SDK passes the task prompt word to the prompt word engine, and then calls the VLM model of the AI model processing module, such as the MobileVLM model, in the same way as dlopen lib_mobilevlm_plugin.so. The model inference engine code calls the sensor engine through the OpenCV library, and finally calls the OS module to control the camera to take a picture and return the picture data. The MobileVLM model takes the picture and the prompt word "scene_mood" for inference, and finally outputs the current environment atmosphere, such as "relaxed" atmosphere. Then return to the agent engine, and update the task "2. Search for relaxed type music and play" formatted command "{id:2, type:normal, tool:music_player, task: play_music_relax}".

[0137] (8) Execute task 2 - search for relaxed type music and play: the updated task is passed to the OS module as an execution workflow, and then returned to the AI Client SDK of the AI adapter. Call the registered fixed task interface, such as the C++ API playMusic(relax). playMusic(relax) is to search for the corresponding type of playlist in the specified folder first, and then use system() to call mpg123 to play the music in the relaxed playlist.

[0138] (9) Execution completion: all tasks are executed, and the AI Client SDK of the AI adapter returns a flag indicating execution completion. At this point, the workflow execution ends.

[0139] The present application embeds a modular AI capability framework in a traditional operating system (OS), endowing various terminal devices (mobile phones, car machines, embedded devices, etc.) with intelligent task processing, natural interaction, and multi-model collaboration capabilities, while ensuring data privacy and system flexibility. Specifically, it includes:

[0140] 1. Cross-platform multi-system AI empowerment: Integrating AI capabilities for various operating systems (Linux / Android / QNX, etc.) on different hardware platforms (Qualcomm / MediaTek / Nvidia, etc.), breaking through device and OS limitations.

[0141] 2. Multi-modal AI capability integration and invocation: Supporting the invocation and management of various types of AI models, including LLM (Large Language Model) for complex language understanding and generation, VLM (Visual Language Model) for cross-modal understanding of images and text, ASR (Automatic Speech Recognition) for converting speech to text, TTS (Text-to-Speech) for converting text to natural speech, and RAG (Retrieval Augmented Generation) for enhancing answer accuracy by combining external knowledge.

[0142] 3. Plug-in model management and flexible adaptation: Adopting a plug-in model design, allowing developers or users to flexibly access, replace, or upgrade different AI models. The AI model processing module is responsible for creating model inference sessions and dynamically invoking corresponding model plugins for inference based on input (prompt words).

[0143] 4. End-side priority privacy protection inference: Mainly relying on local offline lightweight models to perform inference tasks, significantly reducing cloud data transmission, and protecting user sensitive data (such as images, audio, and conversation content) privacy from the source.

[0144] 5. Intelligent end-to-end task processing engine: Achieving intelligent analysis, orchestration, and execution of tasks through a logic processing module.

[0145] 6. Unified developer interface (AI adapter SDK): Providing standardized SDK function interfaces as a bridge between upper-layer applications or services and the AI framework of the present application, simplifying task submission and data transmission processes.

[0146] In a third aspect, the present application also provides a computing device, which includes a processor and a memory. The memory is configured to store a computer program, and the computer program includes program instructions. The processor or computer is configured to invoke the program instructions to execute the method as described above.

[0147] In a fourth aspect, the embodiments of the present application further provide a computer readable storage medium, which stores a computer program, the computer program comprising program instructions, the program instructions causing a processor or a computer to execute the method as described above when the processor or the computer executes the program instructions.

[0148] As shown in Figure 3 The computing device 1000 provided by the embodiments of the present application includes a processor or a computer (not shown in the figure) 1001 and a memory 1002, and the processor or the computer 1001 and the memory 1002 can be connected with each other through a communication bus 1003. The communication bus 1003 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus 1003 can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the memory 1002 is used for storing a computer program, and the computer program includes program instructions, and the processor 1001 is configured to invoke the program instructions, and the above program includes instructions for executing part or all of the steps of the method contained above.

[0149] The processor 1001 can be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the programs of the above solutions.

[0150] The memory 1002 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, and can also be an Electrically Erasable Programmable Read-Only Memory (EEPROM), a Compact Disc Read-Only Memory (CD-ROM) or other optical disk storage, an optical disk storage (including a compact optical disk, a laser disk, an optical disk, a digital versatile optical disk, a Blu-ray optical disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program codes in the form of instructions or data structures and capable of being accessed by a computer, but is not limited to this. The memory can exist independently and be connected with the processor through a bus. The memory can also be integrated with the processor.

[0151] The computing device 1000 can also include a communication module 1004 and a display 1005. The communication module 1004 can be in communication connection with the optical tracking device. The communication module 1004 can be a wireless communication module (such as a WiFi module, a Bluetooth module, etc.) or a wired communication module.

[0152] In addition, the computing device 1000 can also include a communication interface (such as a USB interface, a microphone interface, etc.), an antenna, and the like general components, which are not described here in detail.

[0153] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all described as a series of action combinations, but those skilled in the art should know that the present application is not limited to the order of the actions described, because according to the present application, certain steps can be performed in other order or at the same time. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0154] In the above embodiments, the description of each embodiment is focused on, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0155] In several embodiments provided by the present application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are only schematic. The division of the units is only a logical function division. There can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another apparatus, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, apparatuses or units, and can be electrical or other forms.

[0156] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0157] In addition, the function units in each embodiment of the application can be integrated in one processing unit, or each unit can be physically present alone, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of software program module.

[0158] The integrated unit, if implemented in the form of a software program module and sold or used as an independent product, can be stored in a computer readable memory. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a memory and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned memory includes: a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0159] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer readable memory, which can include a flash disk, a read-only memory, a random access memory, a magnetic disk or an optical disk, etc.

[0160] The embodiments of the present application are described in detail above, and the principles and implementation manners of the present application are described by applying specific examples in the embodiments of the present application. The above embodiment descriptions are only used to help understand the method of the present application and its core idea; meanwhile, for a person of ordinary skill in the art, according to the idea of the present application, the specific implementation manners and application ranges will be changed, and the above description of the present application should not be understood as a limitation.

[0161] It should be noted that the above embodiments can be freely combined according to needs. The above is only a preferred implementation manner of the present application, and it should be noted that, for a person of ordinary skill in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.< / toolmeta>

Claims

1. A method for implementing AI processing capability of an operating system, the method comprising: The method comprises: The captured user instruction is subjected to variable definition through an AI adapter, and the variable is sent to an intent engine in a logic processing engine; The intent engine in the logic processing engine is subjected to corresponding processing through a prompt word engine in the logic processing engine to generate an inference flow, and the inference flow is sent to AI model processing; When the AI model processing executes the inference flow, a corresponding AI model is called to perform inference; A task flow is generated according to the inference result, and each execution flow in the task flow is executed through an AI adapter calling a corresponding task execution module; The AI model processing executes the inference flow, and a corresponding AI model is called to perform inference, specifically comprising: An encapsulation plug-in of the AI model processing is called, and the encapsulation plug-in comprises a session management interface and a model inference interface; A model session is created through the session management, and an inference task is executed on the input inference flow through the model inference interface; The inference result is returned to the prompt word engine of the logic processing engine; The model inference interface is used to execute the inference task on the input inference flow, specifically comprising: The inference flow is taken as the input of the AI model processing through the prompt word engine interface of the corresponding logic processing engine; The AI model processing executes model inference, completes inference through the model inference interface, and generates a command; The generated command is returned to the intent engine in the logic processing engine through the prompt word engine; The logic processing engine receives the command, compares the current command with a pre-registered support function whitelist to determine whether the command is executable, and if so, takes the command as an action flow string to pass to the agent engine of the logic processing engine for multi-task arrangement, which generates a serial graph of an execution node, each node being an execution flow, and each execution flow being executed through an AI adapter calling a corresponding module and tool; The AI adapter calls the agent task processing for the task requiring AI capability; The agent task processing calls the corresponding model of the AI model processing through the task prompt word passed to the prompt word engine, the model inference interface calls the required sensor engine, and the sensor engine calls the OS module to execute data collection of the sensor; The model inference interface combines the data collected by the sensor and the task prompt word of the prompt word engine to perform inference and obtain an inference result; The inference result is returned to the agent engine in the logic processing engine, and the task execution workflow is updated; The updated task execution workflow is passed to the OS module and then returned to the agent task processing of the AI adapter, a registered fixed task interface is called, and the task execution is completed, The agent task processing of the AI adapter returns a flag indicating that the execution is completed. The system comprises: ​ 2. A system for implementing AI processing capability of an operating system, the system comprising: ​ An OS module, an AI model processing module, a logic processing module, and an AI adapter, wherein the OS module is connected with an agent engine submodule in the AI model processing module and the logic processing module, the AI model processing module is connected with a sensor engine submodule in the logic processing module and a prompt word engine submodule in the logic processing module, the prompt word engine submodule in the logic processing module is connected with an intent engine submodule in the logic processing module, and the AI adapter is connected with the sensor engine submodule in the logic processing module, the prompt word engine submodule in the logic processing module, and the intent engine submodule in the logic processing module, respectively; The AI model processing module supports an AI large model, includes various model inference engines, and is encapsulated into a model inference interface; The sensor engine submodule of the logic processing module is used for scheduling and managing sensors; The prompt word engine submodule of the logic processing module is used for prompt word caching, prompt word splicing, prompt word trigger inference, and returning a result returned by an inference interface of the AI model processing module; The intent engine submodule of the logic processing module is used for converting input into a formatted text command, i.e., feeding the input into a corresponding model plug-in and returning a formatted text command generated by the model; The agent engine submodule of the logic processing module is used for generating a serial graph of an execution node from the formatted text command, wherein each node in the graph is an execution flow, and each execution flow is executed by calling a corresponding module through the AI adapter; The AI adapter is used for executing a task of an agent function. 3.The system for realizing AI processing capability of an operating system according to claim 2, wherein, The intent engine submodule of the logic processing module is also used for comparing a pre-registered support function whitelist according to the input formatted text command, judging whether a task is executable, and if the task is executable, taking the formatted text command as an action flow string and passing the action flow string to the agent engine of the logic processing module for multi-task arrangement.

4. A computing device, comprising: The computing device includes a processor and a memory, the memory is used for storing a computer program, the computer program includes program instructions, and the processor is configured to invoke the program instructions and execute the method in claim 1.

5. A computer readable storage medium, characterized in that, The computer readable storage medium stores a computer program, the computer program includes program instructions, and the program instructions make the processor or calculator execute the method in claim 1 when executed by the processor.

Citation Information

Patent Citations

  • Large model intention reasoning method and device based on reasoning template, equipment and medium

    CN119005153A