Intelligent agent service system based on domestic operating system

By integrating operational behavior perception and reinforcement learning technologies on the domestic operating system, a closed loop of perception-modeling-optimization-explanation technology is formed, and the problems of poor behavioral perception sensitivity, lagging response to preference changes, and low decision interpretability in the domestic ecosystem are solved, and the agent's autonomous controllability and efficient reasoning are achieved.

CN120029517AActive Publication Date: 2025-05-23INSPUR SOFTWARE TECH CO LTD

Patent Information

Application Number
CN202510503462.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-23
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

In the domestic ecosystem, conventional agent systems have problems such as poor behavioral perception sensitivity, lagging response to preference changes, and low interpretation of decisions.

Method used

The intelligent service system based on the domestic operating system is adopted, and the operation behavior perception and reinforcement learning technology is integrated, and a complete technical closed loop of perception-modeling-optimization-explanation-based interpretation is formed through the operation perception engine, human feedback reinforcement learning center, MCP protocol adapter and thinking chain visual designer.

Benefits of technology

The paradigm transition from passive execution to active collaboration has been realized, the sensitivity of behavioral perception, response speed of preference change, and decision interpretability are improved, and the system's autonomous controllable and efficient reasoning capabilities are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029517A_ABST
    Figure CN120029517A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent agent service system based on a localized operating system, belongs to the technical field of fusion of operating systems and artificial intelligence, and aims to solve the technical problems of poor behavior perception sensitivity, lagging preference change response, low decision interpretability, high reliability and the like of a conventional intelligent agent system in localized ecology. According to the technical scheme, the system integrates operation behavior perception and reinforcement learning technologies, and forms a perception-modeling-optimization-explanation complete technology closed loop through an operation perception engine, a human feedback reinforcement learning center, an MCP protocol adapter and a thinking chain visual designer. Normal form transition of the intelligent agent from passive execution to active collaboration is realized; wherein user behavior data are captured and analyzed in real time through the operation perception engine, operation sequence data are obtained, the operation sequence data serve as features to be input into the reinforcement learning center, and the reinforcement learning center outputs an operation instruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of integration of operating system and artificial intelligence, and specifically to an intelligent body service system based on a domestic operating system. Background Art

[0002] As intelligent body technology accelerates its penetration into the desktop, the defects of the existing technology system in terms of localization adaptation, security and trustworthiness, and ecological collaboration have seriously restricted the process of my country's information technology independence, which is specifically manifested in three core bottlenecks: 1. Fragmentation of the protocol ecosystem: Currently, intelligent tools generally use private communication protocols (such as MCP's JSON-RPC and ANP's semantic network model), which leads to data silos and a surge in development costs. Taking enterprise ERP system integration as an example, the traditional solution requires the development of independent interfaces for each tool, resulting in an extension of the integration cycle by 3-8 times and an increase in maintenance costs by more than 4 times. This division is not only reflected in the technical architecture, but also extends to the security mechanism level - MCP's ACL whitelist and ANP's zero-knowledge proof are difficult to coordinate, which easily leads to the risk of privacy leakage.

[0003] 2. In-depth reconstruction of behavior perception: Traditional systems can only capture surface events such as clicks and scrolls, and the accuracy of understanding the semantics of GUI operations is insufficient. For example, in the WPS office scenario, the collaborative editing needs of users who frequently switch to "revision mode" are often ignored, and the correlation analysis capabilities of multi-window overlay operations (Excel and PPT linkage) are weak.

[0004] 3. Delayed response to preference changes: The conventional human feedback reinforcement learning (RLHF) cycle is too long. The traditional thinking chain relies on manually labeled data. The RLHF training cycle usually exceeds 72 hours and cannot respond to user preference changes in real time. When users adjust the document format standard, the static model needs to be retrained with all parameters, which can easily lead to catastrophic forgetting.

[0005] Therefore, how to overcome the defects of conventional intelligent systems in the domestic ecosystem, such as poor behavior perception sensitivity, delayed response to preference changes, and low decision-making interpretability, is a technical problem that needs to be solved urgently. Summary of the invention

[0006] The technical task of the present invention is to provide an intelligent agent service system based on a domestic operating system to solve the problems of poor behavior perception sensitivity, delayed response to preference changes, and low decision interpretability existing in conventional intelligent agent systems in a domestic ecosystem.

[0007] The technical task of the present invention is achieved in the following way: an intelligent agent service system based on a domestic operating system, which integrates operation behavior perception and reinforcement learning technology, forms a complete technical closed loop of perception-modeling-optimization-interpretation through an operation perception engine, a human feedback reinforcement learning center, an MCP protocol adapter, and a thinking chain visual designer, and realizes the paradigm transition of intelligent agents from passive execution to active collaboration; Among them, the operation perception engine captures and analyzes user behavior data in real time, obtains operation sequence data, and inputs the operation sequence data as features into the reinforcement learning center. The reinforcement learning center outputs operation instructions. The MCP protocol adapter interacts with different data sources and services through standardized interfaces for the operation instructions recommended by the reinforcement learning center to obtain the results corresponding to the operation instructions, and transmits feedback from external systems (such as user ratings and eye tracking data) back to the human feedback reinforcement learning center to continuously optimize the strategy network of the human feedback reinforcement learning center. The thinking chain visualization designer converts the complex decision-making process and data relationships of the human feedback reinforcement learning center into intuitive visualization views, helping users understand the decision logic and behavior patterns of intelligent behavior analysis devices, improving the system's interpretability and user trust, and guiding the adjustment and optimization of the human feedback reinforcement learning center to form a closed-loop optimization process.

[0008] Preferably, the operation perception engine includes: The operation sequence data acquisition module is used to capture user operation behaviors in real time based on the kernel-level hook mechanism of the domestic operating system to form operation sequence data; wherein, user operation behaviors include GUI operation events and file system access traces; GUI operation events include window focus switching, control clicks and shortcut key triggers; file system access traces include creation, reading, writing and deletion operations; The user multi-dimensional portrait construction module is used to construct a user multi-dimensional behavior portrait through event tracing technology and realize the contextual relationship analysis of operational semantics.

[0009] Preferably, the operation perception engine also has the following functions: ①Support multi-device synchronous perception: unified capture of mobile, desktop, and cloud operation behaviors; ②Add abnormal behavior detection function: real-time alarm for atypical operation modes; ③ Provide encrypted storage and transmission of behavioral data to ensure data security; ④Support plug-in extension, allowing third-party developers to access customized behavior perception rules.

[0010] Preferably, the reinforcement learning hub includes: The model training module is used to train the behavior-document multimodal joint probability model. Among them, the behavior-document multimodal joint probability model is a technology that combines behavior data and document data (such as text, images, etc.) and models the multimodal association through a probability framework. The behavior-document multimodal joint probability model realizes the joint modeling of multimodal data through a probabilistic graphical model (such as Bayesian network, Markov random field) or a deep learning framework (such as variational autoencoder, generative adversarial network). The optimization engine establishment module is used to establish a two-channel feedback-driven policy optimization engine to realize the optimization of the behavior-document multimodal joint probability model. The privacy protection and security module is used to realize privacy protection in the reinforcement learning center by using privacy protection and security mechanisms.

[0011] More preferably, the working process of the model training module is as follows: (1) Calculate the eigenvalue of three dimensions of the corresponding operation frequency, duration, and path complexity for each operation behavior. Among them, the operation frequency refers to the number of times the user executes each operation behavior within a specific time period, reflecting the usage frequency of different operation behaviors by the user; the duration refers to the time the user spends on each operation behavior, reflecting the degree of attention and the amount of time invested by the user in different operations; the path complexity refers to analyzing the complexity of the user's operation path, such as the directory depth of accessing files, the number of jumps, etc., measuring the complexity of the user's operation path. (2) Regard different operation behaviors of the user as words, regard a series of operation sequences of the user as a document, and use the TF-IDF algorithm to quantify the preference weight W of the user for different types of operation behaviors. The specific formula is as follows; ; Among them, represents the number of times the word t appears in the document Di; represents the total number of words in all words in the document ; N represents the total number of documents; represents whether the document contains the word t. If it contains, it is 1; if it does not contain, it is 0; Among them, TF-IDF (Term Frequency-Inverse Document Frequency) is a commonly used weighting technology for text retrieval and text mining. TF-IDF is a statistical method used to evaluate the importance of a word for a document in a document set or a corpus. The importance of a word increases in direct proportion to the number of times it appears in the document, but at the same time decreases in inverse proportion to the frequency of its appearance in the corpus, which can effectively avoid the influence of common words on keywords and improve the relevance between keywords and articles. (3) Using the user’s preference weights and three-dimensional feature values ​​for different types of operation behaviors, all operation sequence data are converted into structured feature vectors as input for the training of the behavior-file multimodal joint probability model; (4) The joint probability distribution of operation behavior and file access behavior is calculated through a dynamic Bayesian network, and the user's operation sequence and file access behavior at different time points are modeled as a conditional probability distribution to capture the causal relationship between operation behavior and file access; (5) The structured feature vector is input into a multi-layer neural network, and the output of the multi-layer neural network is the probability distribution of the operation suggestion. A supervised learning method is used to train the behavior-file joint probability model through historical behavior data, build the spatiotemporal association between operation sequence and file access, and dynamically update the probability distribution to reflect the temporal sequence and context dependency of user behavior. (6) Verify the performance of the behavior-document multimodal joint probability model through cross-validation and indicator evaluation (such as accuracy, recall, and F1 score) to ensure the generalization ability of the model.

[0012] Preferably, the working process of the optimization engine establishment module is as follows: (1) Design feedback channels, which include explicit feedback channels and implicit feedback channels. The explicit channel receives the user's star rating of the intelligent suggestions through the user interface, with a rating range of 1-5. The implicit channel records the user's gaze point, scanning path, and pupil changes during the operation through eye tracking, and records the user's dwell time on specific operations or interface elements. The cognitive load indicators, such as the number of gazes and average dwell time, are calculated based on the eye movement and dwell time data. The ratings and cognitive load indicators are then converted into numerical feedback signals as input to the reward function of reinforcement learning. (2) Designing a reward function by combining explicit and implicit feedback signals. Explicit feedback is used directly as the reward value, while implicit feedback indirectly affects the reward through cognitive load indicators. (3) The PPO algorithm is used to calculate the gradient update of the behavior-file multimodal joint probability model policy network, and the stability of training is ensured by truncating the policy update. The PPO algorithm adopts multi-objective optimization to optimize the user operation path by minimizing the path entropy, improve the operation efficiency, and maximize the confusion of sensitive operations to improve the unrecognizability of sensitive operations and protect user privacy.

[0013] Preferably, the working process of the privacy protection and security module is as follows: (1) During the back-propagation process, gradient masking technology is used to mask gradients involving sensitive data to ensure that private data is not leaked; (2) By adding noise or transforming feature vectors, obfuscate the behavioral characteristics of sensitive operations and reduce the identifiability of sensitive operations; (3) Regularly evaluate the privacy protection effect of the system to ensure the effectiveness of the privacy protection mechanism.

[0014] Preferably, the MCP protocol adapter includes a protocol gateway deployment module and a dynamic service discovery module; Among them, the protocol network management deployment module is used to deploy the MCP protocol gateway using a client-server architecture; the client is used to receive the operation policy generated by the reinforcement learning center, the server side is an external data source and tool, and functional negotiation is carried out between the client and the server to determine the functions and services provided by the client and the server to each other; the protocol network management deployment module integrates the JSON-RPC 2.0 standard protocol and supports two communication modes, specifically as follows: ① Local pipe (stdio) mode: Achieve <10ms low-latency response, suitable for processing local operation behavior data; ② Network flow (SSE) mode: Support high-concurrency calls, suitable for processing behavior data in distributed systems; The dynamic service discovery module is used to identify available MCP servers through an automatic scanning mechanism. The available MCP servers include local IDE plugins, enterprise ERP system interfaces, and cloud AI services (such as the Claude inference engine); the client sends requests to the server according to user requests or the needs of the AI model, the server processes the user requests and may interact with local or remote resources; after the operation is completed, the server returns the processing result to the client, and the client then passes the information back to the host application; then use the URI dynamic template to implement parameterized resource location, support the JSON-RPC 2.0 standard protocol, and ensure the flexibility and dynamicity of service discovery.

[0015] More preferably, the MCP protocol adapter has the following functions: ① Support cross-platform compatibility: Windows, Linux, macOS, mobile; ② Support protocol version management function, support seamless switching between different versions of the MCP protocol; ③ Provide service health monitoring function, and monitor the availability of MCP servers in real time; ④ Support service circuit breaker mechanism to avoid system crashes caused by single-point failures.

[0016] Preferably, the thought chain visualizer includes: The decision tracing model building module is used to build a decision tracing model based on the multi-head attention mechanism, integrate the time series behavior data and system state characteristics, track and record the formation process of each decision generated by the reinforcement learning center, and build a complete decision chain; the formation process of each decision generated by the reinforcement learning center includes the key influencing factors of the decision and the contextual information at the time of decision-making; The decision path reconstruction module is used to reconstruct the decision path using an LSTM network weighted by a time decay factor, model and analyze the time series data in the decision process, highlight the decision trends and patterns that change over time, and help understand the evolution of decisions; The causal relationship analysis module is used to combine knowledge graph technology to generate an explainable view with causal relationships, associate various events and operations in the decision-making process with their results, form a causal relationship diagram, and reveal the logic and motivation behind the decision; Visual output module, used to obtain behavior heat map, file relationship network and strategy evolution timeline; among them, behavior heat map is used to present the distribution of operation modes, showing the operation frequency and mode of users at different times and in different scenarios, helping to identify high-frequency operation areas and user behavior habits; file association network is used to reveal implicit knowledge structure, showing the direct reference, indirect association, and content similarity between files, helping to discover potential knowledge structure and information flow; strategy evolution timeline is used to show the learning process, present the evolution and optimization process of reinforcement learning strategy over time, and help evaluate the learning effect and strategy convergence; the evolution and optimization process of reinforcement learning strategy over time includes the key nodes of strategy adjustment and the changing trend of performance indicators; The feedback and optimization module is used to feed back the visualization results to the reinforcement learning center to provide a reference for further optimization of the reinforcement learning strategy. By analyzing the visualization output, potential problems and room for improvement in the reinforcement learning strategy are discovered, guiding the adjustment and optimization of the reinforcement learning algorithm to form a closed-loop optimization process.

[0017] The intelligent agent service system based on the domestic operating system of the present invention has the following advantages: (I) This invention innovatively integrates MCP protocol access, dynamic behavior modeling and human-machine co-evolution mechanism through multimodal behavior perception, behavior-file multimodal reinforcement learning optimization and lightweight thinking chain generation technology, solves the problems of poor behavior perception sensitivity, delayed response to preference changes, and low decision interpretability of conventional intelligent agent systems in the localized ecosystem, and realizes autonomous control and efficient reasoning of intelligent agent services in the localized environment; (ii) The present invention realizes dynamic access to cross-platform intelligent agent tools through the contextual protocol (MCP), combines operation behavior perception, file semantic understanding and human feedback reinforcement learning algorithm to build a user-personalized thinking chain system; (III) The present invention breaks through the core bottleneck of the application of intelligent technology in domestic operating systems through the three major technical paths of localization adaptation, security enhancement and ecological synergy; (IV) The present invention aims to construct an intelligent agent service system and device based on a domestic operating system, and realize the paradigm shift of intelligent agents from passive execution to active collaboration by building a technical closed loop of "operation-feedback-optimization". Specifically, the unified tool access framework is constructed using the MCP protocol, and the seamless integration of heterogeneous tools is realized based on the client-server architecture. 2.0 protocol encapsulates tool call requests, and the dynamic service discovery mechanism can automatically match local tools or server APIs, combined with fine-grained permission control to ensure that only verified devices and users can access system resources; establishes a three-dimensional behavior modeling system of user operation-file-feedback, analyzes window focus trajectory and gesture operation sequence through bidirectional LSTM, and designs a multimodal feedback interface to convert user ratings into incentive signals for reinforcement learning; deploys a thinking chain optimization engine driven by a hierarchical PPO algorithm, in which the meta-strategy network integrates operation timing features and file TF-IDF vector optimization long-term goals, and the task strategy network achieves rapid decision updates through real-time data pipelines; the present invention constructs a new method of intelligent collaborative paradigm of "operation is training, feedback is evolution" through protocol ecological standardization, in-depth behavior modeling and dynamic optimization of thinking chains, providing a safe and reliable digital foundation for the integration of AI technology into human work; (V) The present invention ensures data sovereignty through domestic kernel-level monitoring and uses national secret algorithms to achieve transmission encryption; and constructs a dynamic association model between behavior and files to break through the static limitations of traditional log analysis; at the same time, it innovatively integrates human feedback mechanisms and explainable AI technology to enable the intelligent agent decision-making process to have both evolutionary capabilities and transparency; experiments show that the present invention can improve the operating efficiency of common office scenarios by 37% and reduce the error rate by 62%, while providing privacy protection capabilities that meet the GB / T 35273 standard. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The present invention is further described below in conjunction with the accompanying drawings.

[0019] Attached Figure 1 This is the structural block diagram of the intelligent agent service system based on the domestic operating system. DETAILED DESCRIPTION

[0020] The following is a detailed description of an intelligent body service system based on a domestic operating system of the present invention with reference to the drawings and specific embodiments of the specification.

[0021] Example: As attached Figure 1As shown, this embodiment provides an intelligent agent service system based on a domestic operating system. The system integrates operation behavior perception and reinforcement learning technology, and forms a complete technical closed loop of perception-modeling-optimization-interpretation through an operation perception engine, a human feedback reinforcement learning center, an MCP protocol adapter, and a thinking chain visual designer, thereby realizing the paradigm transition of intelligent agents from passive execution to active collaboration. Among them, the operation perception engine captures and analyzes user behavior data in real time, obtains operation sequence data, and inputs the operation sequence data as features into the reinforcement learning center. The reinforcement learning center outputs operation instructions. The MCP protocol adapter interacts with different data sources and services through standardized interfaces for the operation instructions recommended by the reinforcement learning center to obtain the results corresponding to the operation instructions, and transmits feedback from external systems (such as user ratings and eye tracking data) back to the human feedback reinforcement learning center to continuously optimize the strategy network of the human feedback reinforcement learning center. The thinking chain visualization designer converts the complex decision-making process and data relationships of the human feedback reinforcement learning center into intuitive visualization views, helping users understand the decision logic and behavior patterns of intelligent behavior analysis devices, improving the system's interpretability and user trust, and guiding the adjustment and optimization of the human feedback reinforcement learning center to form a closed-loop optimization process.

[0022] The operation perception engine in this embodiment includes: The operation sequence data acquisition module is used to capture user operation behaviors in real time based on the kernel-level hook mechanism of the domestic operating system to form operation sequence data; wherein, user operation behaviors include GUI operation events and file system access traces; GUI operation events include window focus switching, control clicks and shortcut key triggers; file system access traces include creation, reading, writing and deletion operations; The user multi-dimensional portrait construction module is used to construct a user multi-dimensional behavior portrait through event tracing technology and realize the contextual relationship analysis of operational semantics.

[0023] The operation perception engine in this embodiment also has the following functions: ①Support multi-device synchronous perception: unified capture of mobile, desktop, and cloud operation behaviors; ②Add abnormal behavior detection function: real-time alarm for atypical operation modes; ③ Provide encrypted storage and transmission of behavioral data to ensure data security; ④Support plug-in extension, allowing third-party developers to access customized behavior perception rules.

[0024] The reinforcement learning hub in this embodiment includes: Model training module, used for training behavior-file multimodal joint probability model; the behavior-file multimodal joint probability model is a technology that combines behavior data with file data (such as text, images, etc.) to model multimodal associations through a probabilistic framework; the behavior-file multimodal joint probability model achieves joint modeling of multimodal data through probabilistic graph models (such as Bayesian networks, Markov random fields) or deep learning frameworks (such as variational autoencoders, generative adversarial networks); The optimization engine building module is used to build a dual-channel feedback-driven policy optimization engine to achieve the optimization of the behavior-file multimodal joint probability model; The privacy protection and security module is used to implement privacy protection in the reinforcement learning hub using the privacy protection and security mechanism.

[0025] The working process of the model training module in this embodiment is as follows: (1) For each operation behavior, the corresponding characteristic values ​​of operation frequency, duration and path complexity are calculated respectively; among them, operation frequency refers to the number of times a user performs each operation behavior in a specific time period, reflecting the frequency of users' use of different operation behaviors; duration refers to the time a user spends on each operation behavior, reflecting the degree of attention and time invested by users in different operations; path complexity refers to analyzing the complexity of the path taken by users when performing operations, such as the directory depth of the accessed file, the number of jumps, etc., to measure the complexity of the user's operation path; (2) Different user operation behaviors are regarded as words, and a series of user operation sequences are regarded as documents. The TF-IDF algorithm is used to quantify the user's preference weight W for different types of operation behaviors. The specific formula is as follows; ; in, Represents the number of times word t appears in document Di; Representation Document The number of words in all the vocabulary; N represents the total number of documents; Representation Document Whether the word t is included, if included, it is 1, if not included, it is 0; (3) Using the user’s preference weights and three-dimensional feature values ​​for different types of operation behaviors, all operation sequence data are converted into structured feature vectors as input for the training of the behavior-file multimodal joint probability model; (4) The joint probability distribution of operation behavior and file access behavior is calculated through a dynamic Bayesian network, and the user's operation sequence and file access behavior at different time points are modeled as a conditional probability distribution to capture the causal relationship between operation behavior and file access; (5) The structured feature vector is input into a multi-layer neural network, and the output of the multi-layer neural network is the probability distribution of the operation suggestion. A supervised learning method is used to train the behavior-file joint probability model through historical behavior data, build the spatiotemporal association between operation sequence and file access, and dynamically update the probability distribution to reflect the temporal sequence and context dependency of user behavior. (6) Verify the performance of the behavior-document multimodal joint probability model through cross-validation and indicator evaluation (such as accuracy, recall, and F1 score) to ensure the generalization ability of the model.

[0026] The working process of the optimization engine establishment module in this embodiment is as follows: (1) Design feedback channels, which include explicit feedback channels and implicit feedback channels. The explicit channel receives the user's star rating of the intelligent suggestions through the user interface, with a rating range of 1-5. The implicit channel records the user's gaze point, scanning path, and pupil changes during the operation through eye tracking, and records the user's dwell time on specific operations or interface elements. The cognitive load indicators, such as the number of gazes and average dwell time, are calculated based on the eye movement and dwell time data. The ratings and cognitive load indicators are then converted into numerical feedback signals as input to the reward function of reinforcement learning. (2) Designing a reward function by combining explicit and implicit feedback signals. Explicit feedback is used directly as the reward value, while implicit feedback indirectly affects the reward through cognitive load indicators. (3) The PPO algorithm is used to calculate the gradient update of the behavior-file multimodal joint probability model policy network, and the stability of training is ensured by truncating the policy update. The PPO algorithm adopts multi-objective optimization to optimize the user operation path by minimizing the path entropy, improve the operation efficiency, and maximize the confusion of sensitive operations to improve the unrecognizability of sensitive operations and protect user privacy.

[0027] The working process of the privacy protection and security module in this embodiment is as follows: (1) During the back-propagation process, gradient masking technology is used to mask gradients involving sensitive data to ensure that private data is not leaked; (2) By adding noise or transforming feature vectors, the behavioral characteristics of sensitive operations are obfuscated, thereby reducing the identifiability of sensitive operations; (3) Regularly evaluate the privacy protection effect of the system to ensure the effectiveness of the privacy protection mechanism.

[0028] The MCP protocol adapter in this embodiment includes a protocol gateway deployment module and a dynamic service discovery module; Among them, the protocol network management deployment module is used to deploy the MCP protocol gateway using a client-server architecture; the client is used to receive the operation strategy generated by the reinforcement learning center, and the server is an external data source and tool. The client and the server negotiate functions to determine the functions and services provided by the client and the server to each other; the protocol network management deployment module integrates the JSON-RPC 2.0 standard protocol and supports two communication modes, as follows: ① Local pipeline (stdio) mode: achieves a low latency response of <10ms, suitable for processing local operation behavior data; ② Network Stream (SSE) mode: supports high-concurrency calls and is suitable for processing behavioral data in distributed systems; The dynamic service discovery module is used to identify available MCP servers through an automatic scanning mechanism. Available MCP servers include local IDE plug-ins, enterprise ERP system interfaces, and cloud AI services (such as the Claude reasoning engine). The client sends a request to the server based on the user request or the needs of the AI ​​model. The server processes the user request and may interact with local or remote resources. After the operation is completed, the server returns the processing result to the client, and the client passes the information back to the host application. The URI dynamic template is then used to implement parameterized resource positioning, supporting the JSON-RPC 2.0 standard protocol to ensure the flexibility and dynamism of service discovery.

[0029] The MCP protocol adapter in this embodiment has the following functions: ①Support cross-platform compatibility: Windows, Linux, macOS, and mobile terminals; ②Support protocol version management function and support seamless switching of different versions of MCP protocols; ③ Provide service health monitoring function to monitor the availability of MCP servers in real time; ④Support service circuit breaker mechanism to avoid system crash due to single point failure.

[0030] The thinking chain visual designer in this embodiment includes: The decision tracing model building module is used to build a decision tracing model based on the multi-head attention mechanism, integrate the time series behavior data and system state characteristics, track and record the formation process of each decision generated by the reinforcement learning center, and build a complete decision chain; the formation process of each decision generated by the reinforcement learning center includes the key influencing factors of the decision and the contextual information at the time of decision-making; The decision path reconstruction module is used to reconstruct the decision path using an LSTM network weighted by a time decay factor, model and analyze the time series data in the decision process, highlight the decision trends and patterns that change over time, and help understand the evolution of decisions; The causal relationship analysis module is used to combine knowledge graph technology to generate an explainable view with causal relationships, associate various events and operations in the decision-making process with their results, form a causal relationship diagram, and reveal the logic and motivation behind the decision; Visual output module, used to obtain behavior heat map, file relationship network and strategy evolution timeline; among them, behavior heat map is used to present the distribution of operation modes, showing the operation frequency and mode of users at different times and in different scenarios, helping to identify high-frequency operation areas and user behavior habits; file association network is used to reveal implicit knowledge structure, showing the direct reference, indirect association, and content similarity between files, helping to discover potential knowledge structure and information flow; strategy evolution timeline is used to show the learning process, present the evolution and optimization process of reinforcement learning strategy over time, and help evaluate the learning effect and strategy convergence; the evolution and optimization process of reinforcement learning strategy over time includes the key nodes of strategy adjustment and the changing trend of performance indicators; The feedback and optimization module is used to feed back the visualization results to the reinforcement learning center to provide a reference for further optimization of the reinforcement learning strategy. By analyzing the visualization output, potential problems and room for improvement in the reinforcement learning strategy are discovered, guiding the adjustment and optimization of the reinforcement learning algorithm to form a closed-loop optimization process.

[0031] The working process of this embodiment is as follows: S1. Based on the operation perception engine, the user's operation behavior is captured in real time; the user's operation behavior is the input basis of subsequent modules (such as the reinforcement learning center and the MCP protocol adapter): Based on the kernel-level hook mechanism of the domestic operating system, the user's operation behavior is captured in real time, including GUI operation events (such as window focus switching, control clicks, shortcut key triggers) and file system access traces (such as create / read / write / delete operations) to form operation sequence data; S2. The user operation behavior data captured by the operation perception engine is used as a feature input into the reinforcement learning center to train the behavior-file multimodal joint probability model. The behavior-file multimodal joint probability model is used to receive the user's task instructions and output the probability distribution of the recommended operation instructions, as follows: S201, for each operation behavior, respectively calculate the characteristic values ​​of the three dimensions of operation frequency, duration and path complexity; wherein, operation frequency refers to the number of times a user performs each operation behavior within a specific time period, reflecting the frequency of use of different operation behaviors by users; duration refers to the time spent by users on each operation behavior, reflecting the degree of attention and time invested by users on different operations; path complexity refers to analyzing the complexity of the path when users perform operations, such as the directory depth of accessed files, the number of jumps, etc., to measure the complexity of the user's operation path; S202, different operation behaviors of the user are regarded as "vocabulary", and a series of operation sequences of the user are regarded as "documents", and the TF-IDF algorithm is used to quantify the preference weight W of the user for different types of operation behaviors. The specific formula is as follows; ; in, Represents the number of times word t appears in document Di; Representation Document The number of words in all the vocabulary; N represents the total number of documents; Representation Document Whether the word t is included, if included, it is 1, if not included, it is 0; S203, using the calculated preference weights and the extracted three-dimensional features, converting all operation sequence data into structured feature vectors as input for model training; S204, calculating the joint probability distribution of the operation behavior and the file access behavior through a dynamic Bayesian network, modeling the user's operation sequence and file access behavior at different time points as a conditional probability distribution, and capturing the causal relationship between the operation behavior and the file access; S205, construct a multi-layer neural network, with the input being a structured feature vector and the output being a probability distribution of operation suggestions, using a supervised learning method to train a behavior-file joint probability model through historical behavior data, constructing a spatiotemporal association between operation sequences and file accesses, and dynamically updating the probability distribution to reflect the temporal sequence and context dependency of user behaviors; S206. Verify model performance through cross-validation and indicator evaluation (such as accuracy, recall, F1 score) to ensure the generalization ability of the model; S3. Receive the user's task instructions, obtain the optimal operation instructions based on the trained behavior-file multimodal joint probability model, call external data sources and services through the standardized interface provided by the MCP protocol adapter and execute the corresponding instructions to obtain the operation results corresponding to the operation instructions; among them, the operation strategy generated by the reinforcement learning center is passed to the external system (such as ERP, AI service) through the MCP protocol adapter to automatically execute the corresponding operation; the construction of the MCP protocol adaptation framework includes two parts: protocol gateway deployment and dynamic service discovery. The protocol network management deployment adopts the client-server architecture to deploy the MCP protocol gateway. The client is used to receive the operation strategy generated by the reinforcement learning center, and the server is an external data source and tool. The client and the server negotiate functions to determine what functions and services they can provide to each other; integrate the JSON-RPC 2.0 standard protocol and support two communication modes: ① The local pipeline (stdio) mode achieves a low latency response of <10ms, which is suitable for processing local operation behavior data; ② The network stream (SSE) mode supports high-concurrency calls and is suitable for processing behavioral data in distributed systems; Dynamic service discovery uses an automatic scanning mechanism to identify available MCP servers, including but not limited to local IDE plug-ins, enterprise ERP system interfaces, and cloud AI services (such as the Claude reasoning engine). The client sends requests to the server based on user requests or the needs of the AI ​​model. The server processes these requests and may interact with local or remote resources. After the operation is completed, the server returns the processing results to the client, and the client passes the information back to the host application. URI dynamic templates are used to implement parameterized resource positioning, and the JSON-RPC 2.0 standard protocol is supported to ensure the flexibility and dynamism of service discovery. S4. Based on the dual-channel feedback-driven strategy optimization engine in the reinforcement learning hub, collect user feedback on operation results (such as user ratings, eye tracking data), further optimize the strategy, and improve the model operation efficiency and privacy protection capabilities; the details are as follows: S401. Design feedback channels, including explicit feedback channels and implicit feedback channels. The explicit channel receives the user's star rating for intelligent suggestions through the user interface, with a rating range of 1-5. The implicit channel records the user's gaze point, scanning path, and pupil changes during the operation through eye tracking, and records the user's stay time on specific operations or interface elements. The cognitive load index, such as the number of fixations and average stay time, is calculated through eye movement and stay time data; the rating and cognitive load index are converted into numerical feedback signals as input to the reward function of reinforcement learning; S402. Design a reward function by combining explicit and implicit feedback signals, where explicit feedback is directly used as the reward value, and implicit feedback indirectly affects the reward through the cognitive load indicator; S403, using the PPO algorithm to calculate the gradient update of the behavior-file multimodal joint probability model policy network, and ensuring the stability of the training by truncating the policy update; wherein the PPO algorithm adopts multi-objective optimization, optimizes the user operation path by minimizing the path entropy, improves the operation efficiency, and maximizes the confusion of sensitive operations to improve the unrecognizable nature of sensitive operations and protect user privacy; S5. Use the Thinking Chain Visual Designer to transform the decision-making process and data relationship of the behavior-document multimodal joint probability model into an intuitive visualization view to help users understand the decision logic and behavior pattern of the system, improve the interpretability of the system and user trust, and guide the adjustment and optimization of the reinforcement learning algorithm to form a closed-loop optimization process. The details are as follows: S501. Decision Traceability Model Construction: Construct a decision traceability model based on the multi-head attention mechanism, integrate temporal behavior data and system state features, and track and record the formation process of each decision generated by the reinforcement learning center, including the key influencing factors of the decision, the context information at the time of decision-making, etc., to construct a complete decision chain; S502. Decision Path Reconstruction: Reconstruct the decision path using an LSTM network weighted by a time decay factor, model and analyze the temporal data in the decision-making process, and highlight the decision trends and patterns that change over time to help understand the evolution process of the decision; S503. Causal Association Analysis: Combine knowledge graph technology to generate an interpretable view with causal associations, associate each event and operation in the decision-making process with its resulting outcomes to form a causal relationship diagram, and reveal the logic and motivation behind the decision; S504. Visualization Output, specifically as follows: S50401. Behavior Heat Map: Present the distribution of operation patterns, display the operation frequencies and patterns of users at different times and in different scenarios, and help identify high-frequency operation areas and user behavior habits; S50402. File Association Network: Reveal the implicit knowledge structure, display the association relationships between files, including direct references, indirect associations, content similarities, etc., and help discover potential knowledge structures and information flows; S50403. Policy Evolution Timeline: Display the learning process, present the evolution and optimization process of the reinforcement learning policy over time, including the key nodes of policy adjustment, the change trends of performance indicators, etc., and help evaluate the learning effect and the convergence of the policy; S505. Feedback and Optimization: Feed the visualization results back to the reinforcement learning center to provide a reference for further optimization of the policy; by analyzing the visualization output, discover potential problems and improvement spaces in the policy, guide the adjustment and optimization of the reinforcement learning algorithm, and form a closed-loop optimization process.

[0032] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An intelligent agent service system based on a domestic operating system, characterized in that: The system integrates operational behavior perception and reinforcement learning technology. Through the operational perception engine, human feedback reinforcement learning center, MCP protocol adapter and thinking chain visual designer, it forms a complete technical closed loop of perception-modeling-optimization-interpretation, realizing the paradigm shift of intelligent agents from passive execution to active collaboration. Among them, the operation perception engine captures and analyzes user behavior data in real time, obtains operation sequence data, and inputs the operation sequence data as features into the reinforcement learning center. The reinforcement learning center outputs operation instructions. The MCP protocol adapter interacts with different data sources and services through standardized interfaces for the operation instructions recommended by the reinforcement learning center, obtains the results corresponding to the operation instructions, and transmits the feedback of the external system back to the human feedback reinforcement learning center, continuously optimizing the strategy network of the human feedback reinforcement learning center; the thinking chain visualization designer converts the complex decision-making process and data relationship of the human feedback reinforcement learning center into an intuitive visualization view, and guides the adjustment and optimization of the human feedback reinforcement learning center to form a closed-loop optimization process.

2. The intelligent agent service system based on the domestic operating system according to claim 1 is characterized in that: The Operation Awareness Engine includes: The operation sequence data acquisition module is used to capture user operation behaviors in real time based on the kernel-level hook mechanism of the domestic operating system to form operation sequence data; wherein, user operation behaviors include GUI operation events and file system access traces; GUI operation events include window focus switching, control clicks and shortcut key triggers; file system access traces include creation, reading, writing and deletion operations; The user multi-dimensional portrait construction module is used to construct a user multi-dimensional behavior portrait through event tracing technology and realize the contextual relationship analysis of operational semantics.

3. The intelligent agent service system based on the domestic operating system according to claim 1 or 2, characterized in that: The operation perception engine also has the following functions: ①Support multi-device synchronous perception: unified capture of mobile, desktop, and cloud operation behaviors; ②Add abnormal behavior detection function: real-time alarm for atypical operation modes; ③ Provide encrypted storage and transmission of behavioral data to ensure data security; ④Support plug-in extension, allowing third-party developers to access customized behavior perception rules.

4. The intelligent agent service system based on the domestic operating system according to claim 1 is characterized in that: The reinforcement learning hub includes: Model training module, used for behavior-document multimodal joint probability model training; The optimization engine building module is used to build a dual-channel feedback-driven policy optimization engine to achieve the optimization of the behavior-file multimodal joint probability model; The privacy protection and security module is used to implement privacy protection in the reinforcement learning hub using the privacy protection and security mechanism.

5. The intelligent agent service system based on the domestic operating system according to claim 4 is characterized in that: The working process of the model training module is as follows: (1) For each operation behavior, the corresponding characteristic values ​​of operation frequency, duration and path complexity are calculated respectively; among them, operation frequency refers to the number of times a user performs each operation behavior in a specific time period, reflecting the frequency of users' use of different operation behaviors; duration refers to the time a user spends on each operation behavior, reflecting the degree of attention and time invested by users in different operations; path complexity refers to analyzing the complexity of the path taken by users when performing operations, measuring the complexity of the user's operation path; (2) Different user operation behaviors are regarded as words, and a series of user operation sequences are regarded as documents. The TF-IDF algorithm is used to quantify the user's preference weight W for different types of operation behaviors. The specific formula is as follows; ; in, Represents the number of times word t appears in document Di; Representation Document The number of words in all the vocabulary; N represents the total number of documents; Representation Document Whether the word t is included, if included, it is 1, if not included, it is 0; (3) Using the user’s preference weights and three-dimensional feature values ​​for different types of operation behaviors, all operation sequence data are converted into structured feature vectors as input for the training of the behavior-file multimodal joint probability model; (4) The joint probability distribution of operation behavior and file access behavior is calculated through a dynamic Bayesian network, and the user's operation sequence and file access behavior at different time points are modeled as a conditional probability distribution to capture the causal relationship between operation behavior and file access; (5) The structured feature vector is input into a multi-layer neural network, and the output of the multi-layer neural network is the probability distribution of the operation suggestion. A supervised learning method is used to train the behavior-file joint probability model through historical behavior data, build the spatiotemporal association between operation sequence and file access, and dynamically update the probability distribution to reflect the temporal sequence and context dependency of user behavior. (6) The performance of the behavior-document multimodal joint probability model is verified through cross-validation and indicator evaluation to ensure the generalization ability of the model.

6. The intelligent agent service system based on the domestic operating system according to claim 4 is characterized in that: The working process of the optimization engine building module is as follows: (1) Design feedback channels, which include explicit feedback channels and implicit feedback channels. The explicit channel receives the user's star rating of the intelligent suggestions through the user interface, with a rating range of 1-5. The implicit channel records the user's gaze point, scanning path, and pupil changes during the operation through eye tracking, and records the user's dwell time on specific operations or interface elements. The cognitive load index is calculated based on the eye movement and dwell time data. The rating and cognitive load index are then converted into numerical feedback signals as input to the reward function of reinforcement learning. (2) Designing a reward function by combining explicit and implicit feedback signals. Explicit feedback is used directly as the reward value, while implicit feedback indirectly affects the reward through cognitive load indicators. (3) The PPO algorithm is used to calculate the gradient update of the behavior-file multimodal joint probability model policy network, and the stability of training is ensured by truncating the policy update. The PPO algorithm adopts multi-objective optimization to optimize the user operation path by minimizing the path entropy, and to improve the unrecognizability of sensitive operations by maximizing the confusion of sensitive operations, thereby protecting user privacy.

7. The intelligent agent service system based on the domestic operating system according to claim 4 is characterized in that: The working process of the privacy protection and security module is as follows: (1) During the back-propagation process, gradient masking technology is used to mask gradients involving sensitive data to ensure that private data is not leaked; (2) By adding noise or transforming feature vectors, the behavioral characteristics of sensitive operations are obfuscated, thereby reducing the identifiability of sensitive operations; (3) Regularly evaluate the privacy protection effect of the system to ensure the effectiveness of the privacy protection mechanism.

8. The intelligent agent service system based on the domestic operating system according to claim 1 is characterized in that: The MCP protocol adapter includes a protocol gateway deployment module and a dynamic service discovery module; Among them, the protocol network management deployment module is used to deploy the MCP protocol gateway using a client-server architecture; the client is used to receive the operation strategy generated by the reinforcement learning center, and the server is an external data source and tool. The client and the server negotiate functions to determine the functions and services provided by the client and the server to each other; the protocol network management deployment module integrates the JSON-RPC 2.0 standard protocol and supports two communication modes, as follows: ① Local pipeline mode: achieves low latency response of <10ms, suitable for processing local operation behavior data; ②Network flow mode: supports high concurrent calls and is suitable for processing behavioral data in distributed systems; The dynamic service discovery module is used to identify available MCP servers through an automatic scanning mechanism. Available MCP servers include local IDE plug-ins, enterprise ERP system interfaces, and cloud AI services. The client sends a request to the server based on the user request or the needs of the AI ​​model. The server processes the user request and may interact with local or remote resources. After the operation is completed, the server returns the processing result to the client, and the client passes the information back to the host application. The URI dynamic template is then used to implement parameterized resource positioning, supporting the JSON-RPC 2.0 standard protocol to ensure the flexibility and dynamism of service discovery.

9. The intelligent agent service system based on the domestic operating system according to claim 1 or 8, characterized in that: The MCP protocol adapter has the following functions: ①Support cross-platform compatibility: Windows, Linux, macOS, and mobile terminals; ②Support protocol version management function and support seamless switching of different versions of MCP protocols; ③ Provide service health monitoring function to monitor the availability of MCP servers in real time; ④Support service circuit breaker mechanism to avoid system crash due to single point failure.

10. The intelligent agent service system based on the domestic operating system according to claim 1 is characterized in that: The ThoughtChain Visual Designer includes: The decision tracing model building module is used to build a decision tracing model based on the multi-head attention mechanism, integrate the time series behavior data and system state characteristics, track and record the formation process of each decision generated by the reinforcement learning center, and build a complete decision chain; the formation process of each decision generated by the reinforcement learning center includes the key influencing factors of the decision and the contextual information at the time of decision-making; The decision path reconstruction module is used to reconstruct the decision path using an LSTM network weighted by a time decay factor, model and analyze the time series data in the decision process, highlight the decision trends and patterns that change over time, and help understand the evolution of decisions; The causal relationship analysis module is used to combine knowledge graph technology to generate an explainable view with causal relationships, associate various events and operations in the decision-making process with their results, form a causal relationship diagram, and reveal the logic and motivation behind the decision; Visual output module, used to obtain behavior heat map, file relationship network and strategy evolution timeline; among them, behavior heat map is used to present the distribution of operation modes, showing the operation frequency and mode of users at different times and in different scenarios, helping to identify high-frequency operation areas and user behavior habits; file association network is used to reveal implicit knowledge structure, showing the direct reference, indirect association, and content similarity between files, helping to discover potential knowledge structure and information flow; strategy evolution timeline is used to show the learning process, present the evolution and optimization process of reinforcement learning strategy over time, and help evaluate the learning effect and strategy convergence; the evolution and optimization process of reinforcement learning strategy over time includes the key nodes of strategy adjustment and the changing trend of performance indicators; The feedback and optimization module is used to feed back the visualization results to the reinforcement learning center to provide a reference for further optimization of the reinforcement learning strategy. By analyzing the visualization output, potential problems and room for improvement in the reinforcement learning strategy are discovered, guiding the adjustment and optimization of the reinforcement learning algorithm to form a closed-loop optimization process.

Citation Information

Patent Citations

  • User experience optimization system and method based on machine learning

    CN117632087A

  • Intelligent emotion intervention and personalized recommendation method and system based on large model in medical industry

    CN119007942A

  • Human-machine cooperation intelligent control method and system based on AIGC and storage medium

    CN119024723A

  • Digital human intelligent recommendation and decision-making system based on user behavior and context awareness

    CN119311943A

  • Standard electronic archive management method and system based on artificial intelligence

    CN119377997A

Cited By

  • Biological information MCP service calling method, system, equipment and medium

    CN120336048A

  • Urban rail transit intelligent operation and maintenance service construction method based on MCP

    CN120338288A

  • IFC model cognitive agent based on MCP protocol and working method thereof

    CN121433759A

  • MCP protocol-based IFC model cognitive agent and working method thereof

    CN121433759B

  • Insecure model context protocol server remediation for artificial intelligence agents

    US12549574B1