An office demand solving method and system based on an AI agent

By using AI-based multimodal semantic analysis and dynamic access control, the static lag problem of sensitive information identification and access control in enterprise office scenarios has been solved. This has enabled accurate labeling of sensitive information and forward-looking prediction of leakage risks, thereby improving information security and the intelligence level of the office system.

CN121009549BActive Publication Date: 2026-04-07HANGZHOU YIGE CLOUD TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies cannot effectively identify complex and diverse sensitive information in enterprise office scenarios, and the access control methods are static and lagging, making it difficult to dynamically respond to the needs of employees and external collaboration, resulting in low security and efficiency.

Method used

We employ an AI agent-based approach, generating dynamic access control strategies through multimodal semantic analysis, dynamic weighting of contextual information, and behavioral path reasoning, and combining these with an improved XGBoost model for fine-grained control.

Benefits of technology

It enables accurate labeling of sensitive information and forward-looking prediction of leakage risks, improving information security and the intelligence level of office systems, and reducing false alarm rates and the frequency of blocking by access control policies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009549B_ABST
    Figure CN121009549B_ABST
Patent Text Reader

Abstract

This invention relates to the field of enterprise office information security technology, and discloses a solution for office needs based on AI intelligent agents, including: context-sensitive semantic analysis, dynamic graph risk reasoning, multi-stage feature generation and dynamic learning, multi-level security policy generation, and feedback closed-loop optimization. Compared with existing technologies that rely solely on static keyword matching, lack dynamic risk reasoning and flexible permission adjustment, especially in handling complex conditions such as cross-network segment sharing, code names and coded messages, and abnormal operations, this invention addresses the technical problems of accurately identifying access to sensitive information and performing corresponding multi-level processing on sensitive access by using multi-stage feature modeling, path segmentation risk coding, context weighting, and improved XGBoost dynamic learning with AI intelligent technology. This achieves accurate labeling of sensitive information, forward prediction of leakage risks, and intelligent identification and control of permissions, thereby improving information security and the level of office intelligence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of enterprise office information security technology, and in particular to a solution and system for office needs based on AI intelligent agents. Background Technology

[0002] Currently, the demand for information security, sensitive information protection, and dynamic access control in enterprise office scenarios continues to grow. With the popularization of cloud-based office work, remote collaboration, and multi-terminal access, the amount of sensitive information, such as core business data, technical solutions, financial information, and strategic plans, involved in office documents and chat systems is increasing dramatically. Existing technologies, commonly used data loss prevention (DLP) solutions mainly rely on static keyword matching, fixed regular expressions, or preset rule bases to detect sensitive content. While these methods can identify obvious confidential information to a certain extent, they cannot effectively handle complex and diverse contextual expressions. For example, common project codes (such as "Document No. 1"), industry-specific terms, abbreviations, and metaphorical descriptions are easily overlooked by traditional methods, leading to missed detections of sensitive information.

[0003] Furthermore, traditional access control methods often employ Access Control Lists (ACLs) or Role-Based Access Control (RBAC) mechanisms. These mechanisms result in static permission allocation and delayed updates, making it difficult to dynamically respond to changes in employee roles, external collaboration needs, or abnormal user behavior. This often leads to permissions being either too broad or too strict, impacting security and work efficiency. Simultaneously, existing technologies lack the ability to perform multi-hop reasoning and leak path prediction based on user behavior. When sensitive information flows across network segments, is forwarded through multiple groups, or is sent via email, the system cannot detect and block potential leak chains in advance, posing significant security risks.

[0004] Therefore, there is an urgent need for a new technical solution that, in complex and ever-changing office scenarios, combines multimodal semantic analysis, dynamic weighting of contextual information, behavioral path reasoning, and adaptive permission learning mechanisms to still achieve accurate labeling of sensitive information, forward-looking prediction of leakage risks, and fine-grained dynamic permission control, so as to improve the overall information security, compliance, and intelligence level of the office system of enterprises. Summary of the Invention

[0005] To address the aforementioned technical shortcomings, the purpose of this invention is to propose a solution for office needs based on AI intelligent agents. This solution aims to resolve the technical problems of existing technologies that rely solely on static keyword matching, lack dynamic risk reasoning and flexible permission adjustment, and are particularly incapable of accurately identifying access to sensitive information and performing corresponding multi-level processing of sensitive access under complex conditions such as cross-network segment sharing, code names and coded language, and abnormal operations.

[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: The present invention provides a solution for office needs based on AI intelligent agents.

[0007] The AI-based solutions for office needs include:

[0008] Step S10: Collect raw text content from the internal office document system and office chat system using an AI agent, process the raw text content based on a preset multimodal text analysis network, and extract the context semantic vector. Based on a pre-defined set of sensitive topic vectors and context semantic vector By combining additional contextual information with dynamic contextual weighting and multi-level semantic matching, a sensitivity score is obtained. and output sensitivity label features. ;

[0009] Step S20: Obtain sensitivity score Greater than the preset sensitivity threshold Corresponding user operation events and resource change behavior Automatically generate user resource behavior graphs Based on sensitivity scoring User resource behavior graph After applying sensitivity weighting to the node weights, the user resource behavior graph is calculated. Each path in Leakage probability score and output behavioral map features. ;

[0010] Step S30: Based on the sensitivity annotation features and behavioral map features Coarse-grained importance discrimination and behavior path segment refinement encoding are performed to form multi-stage progressive feature vectors. The multi-stage progressive feature vectors are then input into a pre-trained improved XGBoost model to output the dynamic permission control factor for user resources.

[0011] Step S40: Generate a multi-level security policy set for specific user resources based on dynamic access control factors;

[0012] Step S50: Collect feedback information after the execution of the multi-level security policy set in real time, generate a security feedback matrix, and dynamically update the feature splitting condition and leaf node gain threshold inside the improved XGBoost model in step S30 based on the security feedback matrix.

[0013] Preferably, in step S10, the additional context information includes historical access context, historical user behavior information, user operation time and user operation device information, historical confidential information, historical financial terms, historical R&D project codes, and historical customer lists.

[0014] Preferably, in step S20, the user resource behavior graph... The nodes in the graph include user nodes, resource nodes, and external entity nodes; the user resource behavior graph The edges in the data include access behavior edges, download and copy edges, share and forward edges, edit and modify edges, and screenshot or screenshot edges. The attributes of each edge include behavior type, occurrence timestamp, operating device type, whether it crosses network segment or security domain, and permission context.

[0015] Preferably, in step S40, the multi-level security policy set includes direct access denial, triggering manual approval, allowing only read-only access, adding multi-factor authentication, outbound blocking, real-time monitoring prompts, and pop-up warnings; the logic of the multi-level security policy set includes: when the value corresponding to the dynamic permission control factor is lower than the first threshold, directly denying access and recording an alarm log; when the value corresponding to the dynamic permission control factor is between the first threshold and the second approval threshold, automatically triggering the approval process or forcibly executing read-only access; when the value corresponding to the dynamic permission control factor is higher than the third threshold, allowing access but adding multi-factor authentication and real-time behavior monitoring.

[0016] Preferably, in step S10, the sensitive topic vector set is based on a preset set. and context semantic vector Sensitivity scores are obtained by performing context-based dynamic weighting and multi-level semantic matching. and output sensitivity label features. The steps specifically include:

[0017] Extract additional contextual information from the current access request, including historical access context, historical user behavior information, user operation time, and user operation device information, and generate additional contextual features; apply the additional contextual features to a pre-defined set of sensitive topic vectors. The weights of each topic vector are dynamically adjusted to generate an intermediate set of sensitive topic vectors.

[0018] Using a dynamically weighted set of intermediate sensitive topic vectors, a context-weighted semantic filtering mechanism is employed to filter the context semantic vectors. First, perform a first-level rapid semantic filtering to select candidate text fragments that are initially similar to sensitive topics, forming an initial sensitive candidate set;

[0019] A contextual reasoning alignment and comparison mechanism is used to perform deep implicit semantic comparison on the initial sensitive candidate set, including reasoning on codes, suggestive expressions, and non-direct descriptions, to further form highly sensitive text fragments;

[0020] Sensitivity scores are calculated by combining the initial set of sensitive candidate texts and the set of highly sensitive text fragments with the intermediate set of sensitive topic vectors, based on similarity matching. and output sensitivity label features. .

[0021] Preferably, in step S20, the leakage probability score is... The calculation uses a continuous product method where the probability of each edge is not disclosed. The formula is as follows: ,in, For path A single edge on the top is used to represent a single operation. For the current edge The conditional probability of being exploited and leaked; For the current user The behavioral risk score is determined based on access frequency, time anomalies, and device trustworthiness. These are contextual factors, determined by the network state; It is the product of the probabilities of no leakage under all one-sided conditions on path p, reflecting the overall probability of no leakage on the path, and is used to calculate the final leakage probability score.

[0022] Preferably, in step S30, based on the sensitivity annotation features... and behavioral map features The steps include: performing coarse-grained importance discrimination and refined encoding of behavior path segments to form multi-stage progressive feature vectors; inputting these multi-stage progressive feature vectors into a pre-trained improved XGBoost model to output a dynamic access control factor for user resources; and specifically including:

[0023] Step S301: Analyze and summarize the sensitivity annotation features and behavioral graph features to form an initial feature set;

[0024] Step S302: Use the importance weight selection library in Python to perform coarse-grained importance discrimination on the initial feature set, retaining the core features;

[0025] Step S303: Use the networkx graph analysis library in Python to decompose the behavioral path into sub-segments and perform detailed coding to generate path segment risk factors;

[0026] Step S304: Integrate core features and path segmentation risk factors to generate multi-stage progressive feature vectors;

[0027] Step S305: Input the multi-stage progressive feature vector into the improved XGBoost model, which has been continuously optimized and trained through context feedback, and output a dynamic permission control factor. The improved XGBoost model includes an input layer for receiving multi-stage progressive feature vectors and automatically supporting grouped features at different stages; a dynamic feature weighting layer for weighted discrimination within the model; a staged residual deepening layer for progressively stacking multi-stage features and using residual learning to strengthen the contribution of important fine-grained features at different stages, improving adaptability to dynamic behavior sequences; a context feedback adaptive adjustment layer for real-time optimization of the model structure, including splitting conditions and leaf node weights, combined with policy execution feedback, forming a closed-loop learning mechanism; and a final decision output layer for outputting the dynamic permission control factor for user resources.

[0028] This invention also provides an office needs resolution system based on AI intelligent agents, comprising:

[0029] The context-sensitive semantic analysis module is used to collect raw text content from internal office document systems and office chat systems, process the raw text content based on a pre-set multimodal text analysis network, and extract context semantic vectors. Based on a pre-defined set of sensitive topic vectors and context semantic vector By combining additional contextual information with dynamic contextual weighting and multi-level semantic matching, a sensitivity score is obtained. and output sensitivity label features. ;

[0030] The dynamic graph risk reasoning module is used to obtain sensitivity scores. Greater than the preset sensitivity threshold Corresponding user operation events and resource change behavior Automatically generate user resource behavior graphs Based on sensitivity scoring User resource behavior graph After applying sensitivity weighting to the node weights, the user resource behavior graph is calculated. Each path in Leakage probability score and output behavioral map features. ;

[0031] A multi-stage feature generation and dynamic learning module is used to generate features based on sensitivity annotations. and behavioral map features Coarse-grained importance discrimination and behavior path segment refinement encoding are performed to form multi-stage progressive feature vectors. The multi-stage progressive feature vectors are then input into a pre-trained improved XGBoost model to output the dynamic permission control factor for user resources.

[0032] The multi-level security policy generation module is used to generate a set of multi-level security policies for specific user resources based on dynamic permission control factors.

[0033] The strategy feedback closed-loop optimization module is used to collect feedback information after the execution of multi-level security policy sets in real time, generate a security feedback matrix, and dynamically update the feature splitting conditions and leaf node gain thresholds inside the improved XGBoost model in step S30 based on the security feedback matrix.

[0034] The present invention also provides an office demand resolution device based on AI intelligent agents, comprising: a memory, a processor, and an office demand resolution program based on AI intelligent agents stored in the memory and executable on the processor. When the office demand resolution program based on AI intelligent agents is executed by the processor, it implements an office demand resolution method based on AI intelligent agents.

[0035] The present invention also provides a computer program product, including an AI-based office demand resolution program, which, when executed by a processor, implements the AI-based office demand resolution method.

[0036] The beneficial effects of this invention are as follows: Compared with the existing technology, which relies solely on static keyword matching, lacks dynamic risk reasoning and flexible permission adjustment, and especially fails to accurately identify access to sensitive information and perform corresponding multi-level response processing under complex conditions such as cross-network segment sharing, code names and coded messages, and abnormal operations, this invention achieves accurate labeling of sensitive information, forward prediction of leakage risks, and intelligent identification and control of permissions through multi-stage feature modeling, path segmentation risk coding, context weighting, and improved XGBoost dynamic learning with the introduction of AI technology, thereby improving information security and the level of office intelligence. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1This is a flowchart illustrating the first embodiment of an AI-based solution for office needs according to the present invention.

[0039] Figure 2 This is a schematic diagram of a device for solving office needs based on AI intelligent agents according to the present invention. Detailed Implementation

[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] Example 1: As Figure 1 The diagram shown is a flowchart of the first embodiment of the office demand solution based on AI intelligent agents of the present invention, which presents the first embodiment of the office demand solution based on AI intelligent agents of the present invention.

[0042] In the first embodiment, the solution to office needs based on AI intelligent agents includes:

[0043] Step S10: Collect raw text content from the internal office document system and office chat system using an AI agent, process the raw text content based on a preset multimodal text analysis network, and extract the context semantic vector. Based on a pre-defined set of sensitive topic vectors and context semantic vector By combining additional contextual information with dynamic contextual weighting and multi-level semantic matching, a sensitivity score is obtained. and output sensitivity label features. ;

[0044] It should be noted that this is based on a pre-defined set of sensitive topic vectors. and context semantic vector Sensitivity scores are obtained by performing context-based dynamic weighting and multi-level semantic matching. and output sensitivity label features. The steps specifically include: extracting additional context information from the current access request, including historical access context, historical user behavior information, user operation time, and user operation device information, and generating additional context features; and applying the additional context features to a preset set of sensitive topic vectors. The weights of each topic vector are dynamically adjusted to generate an intermediate sensitive topic vector set. Using this dynamically weighted intermediate sensitive topic vector set, a context-weighted semantic filtering mechanism is employed to filter the context semantic vectors. First, a rapid semantic filter is performed to identify candidate text fragments that are initially similar to sensitive topics, forming a preliminary sensitive candidate set. Then, a contextual reasoning alignment mechanism is used to perform a deep implicit semantic comparison on this preliminary sensitive candidate set, including inferences about codes, suggestive expressions, and indirect descriptions, further forming highly sensitive text fragments. Finally, a similarity matching calculation is performed on the preliminary sensitive candidate set and the highly sensitive text fragments, combined with an intermediate sensitive topic vector set, to obtain a sensitivity score. and output sensitivity label features. .

[0045] Understandably, compared to traditional sensitive information detection methods based solely on static keywords or pattern matching, this invention innovatively introduces a dynamic contextual weighting mechanism and a multi-level semantic reasoning alignment mechanism. Technically, it deeply couples visitor behavioral characteristics (such as operation time, historical access frequency, and device trustworthiness) with the dynamic weight calculation of sensitive topic vectors. This allows feature discrimination to not only rely on static content but also incorporate real-time scenario risk factors. Specifically, traditional technologies rely on fixed rule bases (such as keywords = "confidential" or "financial") and cannot perceive hidden codes, abbreviations, or cross-sentence expressions (such as "Plan A" or "File F"). By constructing an additional contextual feature matrix (such as behavioral risk factor vectors and access context factors) and dynamically adjusting the weights of sensitive topics during the feature fusion stage, potential sensitive topics can be proactively perceived and reinforced during the detection process, thereby achieving dynamic and fine-grained semantic judgment capabilities.

[0046] It should be understood that the context-based dynamic weighted processing and multi-level semantic matching process, through joint modeling of user behavior graph features and document content features, can adaptively adjust the thresholds and inference logic of sensitive topic vectors according to different security policies. For example, the model can adjust the priority of code-type, policy-type, and R&D-type sub-vectors in sensitive topic vectors based on historical policy feedback matrices, thereby dynamically modifying the overall detection sensitivity and triggering conditions. This adaptive mechanism enables the system to automatically tighten or loosen detection policies based on different stages of business risk levels (such as financial audit periods or product release periods) or attack trends, significantly improving the resilience and forward-looking nature of security policies and solving the problem that traditional fixed threshold solutions cannot adjust in real time according to different scenarios.

[0047] For example, in an internal enterprise verification experiment, for R&D documents containing multi-layered implicit expressions such as "Plan A" and "F File," the experiment set up an external VPN late-night access scenario. The device trust score was 0.45 (low trust), and the historical access anomaly score was 0.3 (high). Through a context-based dynamic weighting mechanism, the weight of code-type topics was automatically increased by 1.5 times, and the weight of topics with non-explicit technical expressions was increased by 1.3 times. Ultimately, the overall semantic similarity score of the context improved from 0.58 in the original static model to 0.91, and dynamic permission factor adjustment was triggered, so that the access process was assigned to a multi-level approval and two-factor authentication strategy path. In a total of 1,500 documents and 3,000 access requests, the average sensitive information detection accuracy improved by 23.7%, the false positive rate decreased by 17.5%, and the approval strategy triggering rate became more accurate (reducing business interruption by about 12%), verifying the effectiveness and advancement of the present invention's context-based multi-dimensional dynamic discrimination and path reasoning capabilities in a real environment.

[0048] Step S20: Obtain sensitivity score Greater than the preset sensitivity threshold Corresponding user operation events and resource change behavior Automatically generate user resource behavior graphs Based on sensitivity scoring User resource behavior graph After applying sensitivity weighting to the node weights, the user resource behavior graph is calculated. Each path in Leakage probability score and output behavioral map features. ;

[0049] It should be noted that the user resource behavior graph The nodes in the graph include user nodes, resource nodes, and external entity nodes; the user resource behavior graph The edges in the data include access behavior edges, download and copy edges, share and forward edges, edit and modify edges, and screenshot or screenshot edges. The attributes of each edge include behavior type, occurrence timestamp, operating device type, whether it crosses a network segment or security domain, and permission context; leakage probability score. The calculation uses a continuous product method where the probability of each edge is not disclosed. The formula is as follows: ,in, For path A single edge on the top is used to represent a single operation. For the current edge The conditional probability of being exploited and leaked; For the current user The behavioral risk score is determined based on access frequency, time anomalies, and device trustworthiness. These are contextual factors, determined by the network state; It is the product of the probabilities of no leakage under all one-sided conditions on path p, reflecting the overall probability of no leakage on the path, and is used to calculate the final leakage probability score.

[0050] Understandably, compared to traditional methods that rely solely on static permissions or event detection based on single access actions, this invention constructs a user resource behavior graph with multiple nodes and various edge types. This allows for a systematic and comprehensive analysis of a user's historical operation links, resource interaction paths, and external propagation pathways. User nodes reflect individual access risk characteristics, resource nodes reflect the sensitivity of files or information, and external entity nodes model potential leakage targets. This graph's structured multi-level nodes and fine-grained behavioral edge information enable the system not only to identify direct operational risks but also to accurately infer potential multi-hop leakage links, thereby improving overall security awareness.

[0051] It should be understood that the path leakage probability scoring method based on edge conditional probability considers the independent risk of each operation as well as the cumulative effect across steps. By multiplying the non-leakage probabilities of all operation edges in the path, it can truly reflect the overall risk accumulation characteristics of the path, overcoming the shortcomings of traditional methods that only approximate path risk estimation based on single-point probabilities or maximum edge risks. Especially in complex paths involving cross-segment sharing, multi-person collaborative editing, and external group forwarding, this method can accurately assess the impact of each node in the complete link on the final leakage risk, achieving path-level quantitative security assessment and dynamic early warning.

[0052] For example, in an internal enterprise security test, a marketing manager accessed a strategic planning document with a sensitivity score of 0.82 via a laptop (device trustworthiness 0.6) outside of work hours and forwarded it to an external consultant via a group chat. The external consultant then took a screenshot and shared it on a partner social media platform. In the corresponding user resource behavior graph, path p contains three key edges: internal download edge, group forwarding edge, and screenshot edge. In this scenario, the conditional probabilities of the edges are 0.5, 0.7, and 0.8, respectively, and the user behavior risk score R(u) = 0.3, with a contextual factor C = 0.4 (VPN abnormal network environment). According to the formula... The final leakage probability score was 0.93, and the security policy engine successfully triggered multi-level approval, mandatory watermarking and full-process auditing strategies, which significantly reduced the risk of sensitive information leakage.

[0053] Step S30: Based on the sensitivity annotation features and behavioral map features Coarse-grained importance discrimination and behavior path segment refinement encoding are performed to form multi-stage progressive feature vectors. The multi-stage progressive feature vectors are then input into a pre-trained improved XGBoost model to output the dynamic permission control factor for user resources.

[0054] It should be noted that, based on the sensitivity labeling features and behavioral map features The process involves several steps, including: performing coarse-grained importance judgment and refining the encoding of behavioral path segments to form a multi-stage progressive feature vector; inputting this multi-stage progressive feature vector into a pre-trained improved XGBoost model to output a dynamic access control factor for user resources; specifically, parsing and summarizing sensitivity annotation features and behavioral graph features to form an initial feature set; using an importance weight selection library in Python to perform coarse-grained importance judgment on the initial feature set, retaining core features; using the networkx graph analysis library in Python to decompose the behavioral path into segments and perform refining encoding to generate path segmentation risk factors; combining core features and path segmentation risk factors to generate a multi-stage progressive feature vector; and inputting the multi-stage progressive feature vector into a pre-trained improved XGBoost model to output a dynamic access control factor for user resources. The vector input is fed into an improved XGBoost model that has been continuously optimized and trained through context feedback, and the output is a dynamic permission control factor. The improved XGBoost model includes an input layer for receiving multi-stage progressive feature vectors and automatically supporting features grouped at different stages; a feature dynamic weighting layer for weighted discrimination within the model; a staged residual deepening layer for progressively stacking multi-stage features and using residual learning to strengthen the contribution of important fine-grained features at different stages, thereby improving adaptability to dynamic behavior sequences; a context feedback adaptive adjustment layer for real-time optimization of the model structure, including splitting conditions and leaf node weights, in conjunction with policy execution feedback, forming a closed-loop learning mechanism; and a final decision output layer for outputting the dynamic permission control factor for user resources.

[0055] Understandably, compared to traditional static permission policies or access control schemes based on single-dimensional static scoring, this invention combines sensitivity labeling features and behavioral graph features, utilizing Python's importance weight selection library and NetworkX graph analysis library to form multi-stage progressive feature vectors. This allows for the full integration of dynamic contextual information, path segmentation risk information, and staged feature evolution patterns in the overall modeling process. This dynamic feature fusion mechanism not only improves the model's adaptability to complex multi-stage access behaviors but also significantly enhances its ability to identify unknown risk patterns, overcoming the shortcomings of traditional models that are insensitive to cross-stage risk evolution and lack path-level refinement capabilities. The feature dynamic weighting layer and staged residual deepening layer set in the improved XGBoost model can be used to adjust feature contributions in different security scenarios (such as audit cycles, outsourced project access periods, and cross-departmental joint analysis periods), and the residual enhancement mechanism can be used to improve the robustness of identifying sudden abnormal behavior sequences. Furthermore, the context feedback adaptive adjustment layer automatically adjusts the model's splitting conditions and leaf node weights by combining the real-time feedback matrix of policy execution (such as approval delay, blocking hit rate, and false alarm rate), achieving closed-loop self-optimization. Compared with the traditional XGBoost static tree model, this invention can dynamically update the internal tree structure and policy rules, significantly improving the fine-grained control effect on sensitivity labeling and path feature interaction, and reducing the operational burden caused by overly broad permissions or overly strict policies.

[0056] For example, in a large-scale simulation experiment involving 5,000 internal enterprise documents containing highly sensitive R&D code names, internal financial forecasts, and cross-departmental collaborative sharing behaviors, the system, based on initial sensitivity labeling features (e.g., 0.83), refined coding of behavioral graph paths (e.g., path risk factors of 0.7, 0.85, and 0.9), and context feedback weight correction, generated multi-stage progressive feature vectors and input them into an improved XGBoost model. The final dynamic access control factor output value reached 0.95 (high-risk access), automatically triggering the strictest approval and access audit strategy. Compared to the traditional static feature linear weighting scheme, which only outputs an access factor of approximately 0.72 and often suffers from a high false positive rate (>15%) due to a fixed threshold, the experiment demonstrates that the adaptive model combining multi-stage features and a closed-loop context feedback improves the overall detection accuracy by approximately 18.6% and the refinement of the access control strategy by approximately 21.3%, demonstrating superior performance under dynamic protection requirements across multiple business stages.

[0057] Step S40: Generate a multi-level security policy set for specific user resources based on dynamic access control factors;

[0058] It should be noted that the multi-level security policy set includes direct access denial, triggering manual approval, allowing only read-only access, adding multi-factor authentication, outbound blocking, real-time monitoring prompts, and pop-up warnings. The logic of the multi-level security policy set includes: when the value of the dynamic access control factor is lower than the first threshold, direct access denial and alarm log recording are recorded; when the value of the dynamic access control factor is between the first threshold and the second approval threshold, the approval process is automatically triggered or read-only access is enforced; when the value of the dynamic access control factor is higher than the third threshold, access is allowed but multi-factor authentication and real-time behavior monitoring are added. The first threshold refers to the minimum risk tolerance threshold, used to identify immediate blocking decisions when there is a significant abnormal risk in access behavior; it includes a risk value calculated based on sensitivity score (e.g., higher than 0.85), probability of leakage of behavioral graph path (e.g., higher than 0.8), user's historical abnormal frequency (e.g., more than three high-sensitivity behaviors in the past 30 days), and current access time (e.g., 22:00–06:00 at night). The second approval threshold refers to the medium-to-high-risk access range, used for fine-grained control of operations requiring additional approval or read-only policies. It includes a weighted assessment of factors such as comprehensive sensitivity score (e.g., 0.6–0.85), behavioral graph path risk (e.g., 0.5–0.8), device trustworthiness below a certain threshold (e.g., less than 0.7), and network status (e.g., using VPN or accessing across security domains). The third threshold refers to a relatively controllable low-risk access range; it includes a comprehensive sensitivity score below 0.6, a path leakage probability below 0.5, no abnormal records in user history, and a device trustworthiness above 0.9. It is mainly used to allow access but with additional multi-factor authentication, real-time monitoring, and watermarking tracing mechanisms.

[0059] Understandably, traditional solutions typically employ static, globally uniform thresholds (such as a single sensitivity score), which cannot adapt to dynamic changes in different scenarios and among different users. In contrast, this solution innovatively introduces a context-based dynamic weighting factor mechanism, which automatically fine-tunes the threshold range based on a real-time feedback matrix. For example, it automatically and adaptively adjusts the threshold based on historical data such as the approval rejection rate, policy trigger rate, and actual business interruption rate over the past week, achieving dynamic convergence and personalized refinement, ensuring both security and business continuity.

[0060] It should be understood that, compared to the traditional binary permission determination method that directly rejects or approves based solely on static threshold scores, this invention introduces a multi-level security policy set and a context-based dynamic threshold mechanism. This allows for a comprehensive assessment of user access risks based on multiple dimensions, including sensitivity scores, path risks, and historical behavior, enabling more granular dynamic adaptation of multiple policies. This not only reduces the problem of excessive blocking caused by misjudgments but also significantly improves policy flexibility and scenario adaptability.

[0061] Step S50: Collect feedback information after the execution of the multi-level security policy set in real time, generate a security feedback matrix, and dynamically update the feature splitting condition and leaf node gain threshold inside the improved XGBoost model in step S30 based on the security feedback matrix.

[0062] It should be noted that the security feedback matrix refers to the collection, statistics, and matrix modeling of multi-dimensional feedback data after the execution of a multi-level security policy set. This data includes policy triggering frequency, approval rejection rate, false alarm rate, actual business blockage rate, user appeal rate, and subsequent compliance audit results. The matrix forms a two-dimensional matrix containing row dimensions (feedback indicators) and column dimensions (policy number or execution round), used to dynamically adjust the model's internal parameter configuration. This matrix is ​​not only used for policy optimization but also for automatically generating comparison factors to drive the improved XGBoost model's internal feature splitting conditions, leaf node gain thresholds, and dynamic weight adjustments, further enhancing the model's real-time adaptive capability to access risks.

[0063] Understandably, compared to traditional access control or sensitivity detection models that rely solely on static offline training or fixed threshold updates, this invention introduces a security feedback matrix. This matrix enables the acquisition of real feedback results after each round of policy execution, automatically fine-tuning the model's splitting conditions (such as feature partitioning thresholds and information gain standards) and leaf node gain thresholds (which affect the final prediction score of leaf nodes). This ensures that the model can continuously optimize based on the actual risk situation of the enterprise during long-term operation, achieving high-frequency, fine-grained closed-loop self-evolution, thereby balancing security, business continuity, and policy accuracy.

[0064] Example 2: Furthermore, the present invention provides an AI-based office demand resolution system that employs an AI-based office demand resolution method as described in the above embodiments, thereby solving a technical problem related to AI-based office demand resolution. Compared with the prior art, the beneficial effects of the AI-based office demand resolution system provided by the present invention are the same as those of the AI-based office demand resolution method provided in the above embodiments, and other technical features of the AI-based office demand resolution system are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0065] Example 3: This invention provides an office needs-solving device based on an AI intelligent agent. Please refer to... Figure 2An AI-based office needs resolution device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to perform an AI-based office needs resolution method as described in Embodiment 1 above. An AI-based office needs resolution device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. An AI-based office needs resolution device is merely an example and should not limit the functionality or scope of the embodiments of this invention. An AI-based office needs resolution device may include a processing unit 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. Random access memory 1004 also stores various programs and data required for the operation of an AI-based office needs resolution device. Processing unit 1001, read-only memory 1002, and random access memory 1004 are interconnected via bus 1005. I / O interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003; and communication devices 1009. Communication device 1009 allows an AI-based office needs resolution device to communicate wirelessly or wiredly with other devices to exchange data. Although an AI-based office needs resolution device with various systems is shown in the figure, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems may be implemented alternatively.

[0066] Example 4: This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps described above for a solution to office needs based on an AI intelligent agent. The computer program product provided by this invention can solve a technical problem related to solving office needs based on an AI intelligent agent. Compared with the prior art, the beneficial effects of the computer program product provided by this invention are the same as those of the solution to office needs based on an AI intelligent agent provided in the above embodiments, and will not be repeated here.

[0067] In particular, according to the embodiments disclosed in this invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this invention.

[0068] It should be understood that the various parts disclosed in this invention can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics may be combined in any suitable manner in one or more embodiments or examples.

[0069] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A solution to office needs based on AI intelligent agents, characterized in that, The methods include: Step S10: Collect raw text content from the internal office document system and office chat system using an AI agent, process the raw text content based on a preset multimodal text analysis network, and extract the context semantic vector. Based on a pre-defined set of sensitive topic vectors and context semantic vector By combining additional contextual information with dynamic contextual weighting and multi-level semantic matching, a sensitivity score is obtained. and output sensitivity label features. ; Among them, based on a preset set of sensitive topic vectors and context semantic vector Sensitivity scores are obtained by performing context-based dynamic weighting and multi-level semantic matching. and output sensitivity label features. The steps specifically include: Extract additional contextual information from the current access request, including historical access context, historical user behavior information, user operation time, and user operation device information, and generate additional contextual features; apply the additional contextual features to a pre-defined set of sensitive topic vectors. The weights of each topic vector are dynamically adjusted to generate an intermediate set of sensitive topic vectors. Using a dynamically weighted set of intermediate sensitive topic vectors, a context-weighted semantic filtering mechanism is employed to filter the context semantic vectors. First, perform a first-level rapid semantic filtering to select candidate text fragments that are initially similar to sensitive topics, forming an initial sensitive candidate set; A contextual reasoning alignment and comparison mechanism is used to perform deep implicit semantic comparison on the initial sensitive candidate set, including reasoning on codes, suggestive expressions, and non-direct descriptions, to further form highly sensitive text fragments; Sensitivity scores are calculated by combining the initial set of sensitive candidate texts and the set of highly sensitive text fragments with the intermediate set of sensitive topic vectors, based on similarity matching. and output sensitivity label features. ; Step S20: Obtain sensitivity score Greater than the preset sensitivity threshold Corresponding user operation events and resource change behavior Automatically generate user resource behavior graphs Based on sensitivity scoring User resource behavior graph After applying sensitivity weighting to the node weights, the user resource behavior graph is calculated. Each path in Leakage probability score and output behavioral map features. ; Among them, the leakage probability score The calculation uses a continuous product method where the probability of each edge is not disclosed. The formula is as follows: ,in, For path A single edge on the top is used to represent a single operation. For the current edge The conditional probability of being exploited and leaked; For the current user The behavioral risk score is determined based on access frequency, time anomalies, and device trustworthiness. These are contextual factors, determined by the network state; It is the product of the probabilities of no leakage of all one-sided conditions on path p, reflecting the overall probability of no leakage of the path, and is used to calculate the final leakage probability score; Step S30: Based on the sensitivity annotation features and behavioral map features Coarse-grained importance discrimination and behavior path segment refinement encoding are performed to form multi-stage progressive feature vectors. The multi-stage progressive feature vectors are then input into a pre-trained improved XGBoost model to output the dynamic permission control factor for user resources. Step S40: Generate a multi-level security policy set for specific user resources based on dynamic access control factors; Step S50: Collect feedback information after the execution of the multi-level security policy set in real time, generate a security feedback matrix, and dynamically update the feature splitting condition and leaf node gain threshold inside the improved XGBoost model in step S30 based on the security feedback matrix.

2. The solution to office needs based on AI intelligent agents as described in claim 1, characterized in that, In step S10, the additional context information includes historical access context, historical user behavior information, user operation time and user operation device information, historical confidential information, historical financial terms, historical R&D project codes and historical customer lists.

3. The solution to office needs based on AI intelligent agents as described in claim 1, characterized in that, In step S20, the user resource behavior graph is generated. The nodes in the graph include user nodes, resource nodes, and external entity nodes; the user resource behavior graph The edges in the data include access behavior edges, download and copy edges, share and forward edges, edit and modify edges, and screenshot or screenshot edges. The attributes of each edge include behavior type, occurrence timestamp, operating device type, whether it crosses network segment or security domain, and permission context.

4. The solution to office needs based on AI intelligent agents as described in claim 1, characterized in that, In step S40, the multi-level security policy set includes direct access denial, triggering manual approval, allowing only read-only access, adding multi-factor authentication, outbound blocking, real-time monitoring prompts, and pop-up warnings. The logic of the multi-level security policy set includes: when the value corresponding to the dynamic permission control factor is lower than the first threshold, direct access denial and recording alarm logs; when the value corresponding to the dynamic permission control factor is between the first threshold and the second approval threshold, automatically triggering the approval process or forcibly executing read-only access; when the value corresponding to the dynamic permission control factor is higher than the third threshold, allowing access but adding multi-factor authentication and real-time behavior monitoring.

5. The solution to office needs based on AI intelligent agents as described in claim 1, characterized in that, In step S30, based on the sensitivity annotation features and behavioral map features The steps include: performing coarse-grained importance discrimination and refined encoding of behavior path segments to form multi-stage progressive feature vectors; inputting these multi-stage progressive feature vectors into a pre-trained improved XGBoost model to output a dynamic access control factor for user resources; and specifically including: Step S301: Analyze and summarize the sensitivity annotation features and behavioral graph features to form an initial feature set; Step S302: Use the importance weight selection library in Python to perform coarse-grained importance discrimination on the initial feature set, retaining the core features; Step S303: Use the networkx graph analysis library in Python to decompose the behavioral path into sub-segments and perform detailed coding to generate path segment risk factors; Step S304: Integrate core features and path segmentation risk factors to generate multi-stage progressive feature vectors; Step S305: Input the multi-stage progressive feature vector into the improved XGBoost model, which has been continuously optimized and trained through context feedback, and output a dynamic permission control factor. The improved XGBoost model includes an input layer for receiving multi-stage progressive feature vectors and automatically supporting grouped features at different stages; a dynamic feature weighting layer for weighted discrimination within the model; a staged residual deepening layer for progressively stacking multi-stage features and using residual learning to strengthen the contribution of important fine-grained features at different stages, improving adaptability to dynamic behavior sequences; a context feedback adaptive adjustment layer for real-time optimization of the model structure, including splitting conditions and leaf node weights, combined with policy execution feedback, forming a closed-loop learning mechanism; and a final decision output layer for outputting the dynamic permission control factor for user resources.

6. An office demand resolution system based on AI intelligent agents, applied to the office demand resolution system based on AI intelligent agents as described in any one of claims 1 to 5, characterized in that, The AI-based intelligent agent-based office needs resolution system includes: The context-sensitive semantic analysis module is used to collect raw text content from internal office document systems and office chat systems using an AI agent. Based on a pre-set multimodal text analysis network, the module processes the raw text content to extract context semantic vectors. Based on a pre-defined set of sensitive topic vectors and context semantic vector By combining additional contextual information with dynamic contextual weighting and multi-level semantic matching, a sensitivity score is obtained. and output sensitivity label features. ; Among them, based on a preset set of sensitive topic vectors and context semantic vector Sensitivity scores are obtained by performing context-based dynamic weighting and multi-level semantic matching. and output sensitivity label features. The steps specifically include: Extract additional contextual information from the current access request, including historical access context, historical user behavior information, user operation time, and user operation device information, and generate additional contextual features; apply the additional contextual features to a pre-defined set of sensitive topic vectors. The weights of each topic vector are dynamically adjusted to generate an intermediate set of sensitive topic vectors. Using a dynamically weighted set of intermediate sensitive topic vectors, a context-weighted semantic filtering mechanism is employed to filter the context semantic vectors. First, perform a first-level rapid semantic filtering to select candidate text fragments that are initially similar to sensitive topics, forming an initial sensitive candidate set; A contextual reasoning alignment and comparison mechanism is used to perform deep implicit semantic comparison on the initial sensitive candidate set, including reasoning on codes, suggestive expressions, and non-direct descriptions, to further form highly sensitive text fragments; Sensitivity scores are calculated by combining the initial set of sensitive candidate texts and the set of highly sensitive text fragments with the intermediate set of sensitive topic vectors, based on similarity matching. and output sensitivity label features. ; The dynamic graph risk reasoning module is used to obtain sensitivity scores. Greater than the preset sensitivity threshold Corresponding user operation events and resource change behavior Automatically generate user resource behavior graphs Based on sensitivity scoring User resource behavior graph After applying sensitivity weighting to the node weights, the user resource behavior graph is calculated. Each path in Leakage probability score and output behavioral map features. ; Among them, the leakage probability score The calculation uses a continuous product method where the probability of each edge is not disclosed. The formula is as follows: ,in, For path A single edge on the top is used to represent a single operation. For the current edge The conditional probability of being exploited and leaked; For the current user The behavioral risk score is determined based on access frequency, time anomalies, and device trustworthiness. These are contextual factors, determined by the network state; It is the product of the probabilities of no leakage of all one-sided conditions on path p, reflecting the overall probability of no leakage of the path, and is used to calculate the final leakage probability score; A multi-stage feature generation and dynamic learning module is used to generate features based on sensitivity annotations. and behavioral map features Coarse-grained importance discrimination and behavior path segment refinement encoding are performed to form multi-stage progressive feature vectors. The multi-stage progressive feature vectors are then input into a pre-trained improved XGBoost model to output the dynamic permission control factor for user resources. The multi-level security policy generation module is used to generate a set of multi-level security policies for specific user resources based on dynamic permission control factors. The strategy feedback closed-loop optimization module is used to collect feedback information after the execution of multi-level security policy sets in real time, generate a security feedback matrix, and dynamically update the feature splitting conditions and leaf node gain thresholds inside the improved XGBoost model in the multi-stage feature generation and dynamic learning module based on the security feedback matrix.

7. An office needs-solving device based on an AI intelligent agent, characterized in that, The AI-based office demand resolution device includes: a memory, a processor, and an AI-based office demand resolution program stored in the memory and executable on the processor. When the AI-based office demand resolution program is executed by the processor, it implements an AI-based office demand resolution method according to any one of claims 1 to 5.

8. A computer program product, characterized in that, The computer program product includes an AI-based office needs resolution program, which, when executed by a processor, implements an AI-based office needs resolution method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Structured sensitive data identification method and device

    CN116451072A

  • Government affair data sharing system based on data security law risk control mode

    CN119989417A