Artificial intelligence task sandbox policy automatic derivation method and system

CN122528140BActive Publication Date: 2026-09-29BEIJING CAPITEK +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611022623.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-10
Publication Date
2026-09-29
Estimated Expiration
2046-07-10

AI Technical Summary

Technical Problem

其一,容器化隔离方案能够提供进程级隔离,但其安全配置项需要由具备专业背景的运维人员手工编写,且配置粒度通常为容器或工作负载级别,与细粒度任务执行单元的隔离需求不匹配,容器启动开销对细粒度执行单元也难以承受

Benefits of technology

通过将结构化配置描述解析为抽象语法树并提取、量化四类风险特征,实现了从业务配置描述到安全隔离策略的自动推导,配置编写方无需配置任何安全参数,消除了人为疏忽导致的安全隐患,也解决了安全策略与业务配置不一致的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122528140B_ABST
    Figure CN122528140B_ABST
Patent Text Reader

Abstract

The application discloses an artificial intelligence task sandbox strategy automatic derivation method and system, and relates to the technical field of computers.The method comprises the following steps: parsing a structured configuration description of an artificial intelligence task into an abstract syntax tree, extracting four types of risk features and quantifying the risk features into a risk feature vector; performing correlation analysis, determining a combined correction value according to a combined risk superposition rule, and constructing a data flow graph based on a data dependency relationship between steps, and applying a conduction coefficient to determine a conduction correction value; calculating a comprehensive risk score according to the above steps, and determining a target sandbox strategy from a plurality of sandbox strategies with different isolation intensities; injecting a sandbox boundary in the form of a node attribute into an executable workflow definition at a compiling link; performing execution according to the sandbox boundary at runtime, and performing versioned hot updating when the strategy is changed.The application realizes automatic derivation of a sandbox strategy, improves risk identification accuracy, avoids performance loss caused by indiscriminate isolation, and does not interrupt business when the strategy is updated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to an automatic inference method and system for sandbox strategies of artificial intelligence tasks based on configuration description semantic risk quantification. Background Technology

[0002] In general AI agent development frameworks, the business side typically defines AI tasks (hereinafter referred to as task configuration descriptions or skill configurations) using structured configuration description languages ​​such as YAML and JSON, which are then parsed and executed by the execution platform. Since these tasks can call external tools and process sensitive data (such as authentication credential data, network channel data, and precise location data), their security isolation is a core issue in the framework design.

[0003] Existing security isolation methods mainly include the following categories: Firstly, containerized isolation solutions can provide process-level isolation, but their security configuration items need to be manually written by operations and maintenance personnel with professional backgrounds, and the configuration granularity is usually at the container or workload level, which does not match the isolation requirements of fine-grained task execution units. The container startup overhead is also difficult for fine-grained execution units to bear.

[0004] Secondly, static code analysis tools can detect code-level vulnerabilities at the source code level. Their analysis object is the source code, and the detection target is code defects such as injection classes, outputting a vulnerability list. However, in the context of artificial intelligence tasks, risks also include business semantic risks. For example, a task may call a fund transfer interface. The code itself may not have vulnerabilities, but it is a high-risk operation at the business semantic level. Such semantic risks cannot be identified by static analysis tools, and static analysis tools cannot deduce runtime isolation strategies based on this.

[0005] Third, while serverless computing simplifies deployment, it has limited security control capabilities and struggles to meet the requirements of state management and hot updates.

[0006] Fourth, while existing process-level sandbox tools have isolation capabilities, they all require explicit manual configuration of isolation parameters, which is separate from business logic. After changes in business configuration, the sandbox configuration needs to be manually modified synchronously, making it difficult to guarantee the consistency between security policies and business configurations.

[0007] In summary, existing technologies, when addressing the question of "how to determine security policies," either rely on manual configuration or employ indiscriminate default policies, neither of which can automatically derive security isolation policies from business configuration descriptions. Furthermore, these policies are mostly statically configured, requiring service restarts for changes, which cannot accommodate long-running execution instances; indiscriminate strict isolation also causes significant performance penalties. Therefore, it is necessary to propose a technical solution that can automatically derive sandbox isolation policies from task configuration descriptions and supports uninterrupted policy updates. Summary of the Invention

[0008] The present invention aims to provide an automatic derivation method and system for sandbox strategies in artificial intelligence tasks, so as to overcome the shortcomings of the existing technology. The technical problem to be solved by the present invention is achieved through the following technical solution.

[0009] According to a first aspect of this application, an automatic inference method for sandbox strategies in artificial intelligence tasks is provided, the method comprising: Obtain a structured configuration description of the artificial intelligence task, parse the structured configuration description into an abstract syntax tree, traverse the abstract syntax tree, extract tool call features, data flow features, control flow features, and human intervention features, and quantize each extracted feature into a risk feature vector according to a preset quantization rule; A correlation analysis is performed on the risk feature vector to obtain the correlation analysis results; wherein, the correlation analysis includes: determining the combination correction value according to the combination risk superposition rule when high-risk features and manual intervention features appear in combination or not; and traversing the data dependencies between each step in the abstract syntax tree, constructing a data flow graph between steps, applying a transmission coefficient to the data flow edges of the data flow graph, and determining the transmission correction value. A comprehensive risk score is calculated based on the risk feature vector and the correlation analysis results. Then, based on a comparison of the comprehensive risk score with a preset threshold, a target sandbox strategy is determined from multiple sandbox strategies with different isolation strengths. The comprehensive risk score is the sum of a weighted average, the combined correction value, and the transmission correction value. During the compilation of the structured configuration description into an executable workflow definition, the sandbox boundary defined by the target sandbox strategy is injected into the nodes of the executable workflow definition in the form of node attributes. This injection does not change the structure and execution semantics of the executable workflow definition. When executing the executable workflow definition, the corresponding node is executed in the isolated environment according to the injected sandbox boundary; and when the sandbox policy changes, a new policy version is generated, the newly started execution instance uses the new policy version, and the execution instance being executed locks the policy version it was created until execution is completed.

[0010] Preferably, the preset quantification rules include: the call type risk value in the tool call feature is assigned according to the tool risk weight dimension table, and different risk categories of tools correspond to different call type risk values; the feature value corresponding to the approval node in the manual intervention feature is assigned a negative value, and the higher the approval level, the larger the absolute value of the negative value.

[0011] Preferably, the weights of each feature in the preset quantification rule are determined based on the analytic hierarchy process (AHP), including: constructing a three-layer feature hierarchy structure, wherein the target layer is the risk score, the criterion layer is the tool call feature, the data flow feature, the control flow feature, and the human intervention feature, and the scheme layer is each specific feature item; by constructing a judgment matrix, each feature is compared pairwise, the weight vector is calculated, and a consistency check is performed, and the consistency ratio is less than 0.1.

[0012] Preferably, the combined risk superposition rule includes: when a high-risk feature does not occur in combination with a feature requiring manual intervention, multiplying the risk value corresponding to the high-risk feature by an amplification factor, wherein the amplification factor ranges from 1.1 to 1.5;

[0013] When a combination of high-risk characteristics and manual intervention characteristics occurs, the risk value corresponding to the high-risk characteristic is multiplied by a mitigation coefficient. The mitigation coefficient is stratified according to the matching degree between the approval level and the operational risk. The mitigation coefficient corresponding to the first approval level ranges from 0.4 to 0.8, and the mitigation coefficient corresponding to the second approval level, which is lower than the first approval level, ranges from 0.7 to 0.9.

[0014] Preferably, the construction of the data flow graph between steps includes: traversing the data dependencies between the output storage fields and input parameter fields of each step in the abstract syntax tree, and constructing the data flow graph with steps as nodes and data dependencies as edges; the transmission coefficient ranges from 0.8 to 1.5, and the transmission coefficient is greater than 1 when data flows from a lower-risk step to a higher-risk step.

[0015] Preferably, the plurality of sandbox strategies include a first strategy, a second strategy, a third strategy, and a fourth strategy, and the preset threshold includes a first threshold and a second threshold, wherein the first threshold is greater than the second threshold;

[0016] The step of determining the target sandbox strategy from multiple sandbox strategies with different isolation strengths includes: when the comprehensive risk score is greater than or equal to the first threshold, determining the first strategy as the target sandbox strategy, wherein the first strategy includes process isolation, network isolation, file system isolation, and detailed level auditing;

[0017] When the comprehensive risk score is greater than or equal to the second threshold and less than the first threshold, the second strategy is determined to be the target sandbox strategy. The second strategy includes resource restrictions and disabling direct network access and forwarding it through an internal proxy.

[0018] When the overall risk score is less than the second threshold and the artificial intelligence task requires network access, the third strategy is determined to be the target sandbox strategy, which includes controlled network access; otherwise, the fourth strategy is determined to be the target sandbox strategy, which includes basic resource restrictions.

[0019] Preferably, injecting the sandbox boundary defined by the target sandbox strategy into the nodes of the executable workflow definition in the form of node attributes includes: for tool invocation nodes, injecting a complete sandbox boundary including entry resource checks, exit audit points, and exception handling points;

[0020] For conditional branch nodes, inject branch path audit points; for manual approval nodes, inject wait timeout policies and approval status verification.

[0021] For each response node, inject a security checkpoint into the output content;

[0022] The injection granularity is set to step-level injection by default. When there are steps with different risk levels in the same artificial intelligence task, sub-step-level injection is used so that steps with higher risk are injected into sandbox boundaries with higher isolation strength, and steps with lower risk are injected into sandbox boundaries with lower isolation strength.

[0023] Preferably, the step of executing the corresponding node in the isolated environment according to the injected sandbox boundary includes: detecting whether the operating system kernel-level isolation mechanism is available when the system starts; if the kernel-level isolation mechanism is available, entering enhanced mode, and implementing process isolation, network isolation and file system isolation through the kernel-level isolation mechanism; if the kernel-level isolation mechanism is unavailable, entering degraded mode, softly limiting processor time and memory through the operating system resource limiting mechanism, and outputting an alarm log indicating that the current mode is degraded; wherein, the enhanced mode and the degraded mode are automatically switched.

[0024] Preferably, the method further includes: recording feedback data of sandbox execution, the feedback data including false positive events where legitimate operations are rejected due to overly strict sandbox policies and false negative events where security incidents are missed due to overly lenient sandbox policies; periodically calculating the false positive rate and false negative rate based on the accumulated feedback data, and generating adjustment suggestions for the weights and preset thresholds in the preset quantization rules when the false positive rate is greater than a first proportional threshold or the false negative rate is greater than a second proportional threshold; wherein each adjustment of the weights and preset thresholds generates a new parameter version, and historical parameter versions are retained to support rollback.

[0025] According to a second aspect of this application, an automatic strategy derivation device for artificial intelligence task sandboxes is provided, the device comprising:

[0026] The semantic parsing module is configured to obtain the structured configuration description of the artificial intelligence task, parse the structured configuration description into an abstract syntax tree, traverse the abstract syntax tree, extract tool call features, data flow features, control flow features, and human intervention features, and quantize each extracted feature into a risk feature vector according to a preset quantization rule;

[0027] The correlation analysis module is configured to perform correlation analysis on the risk feature vector to obtain correlation analysis results; wherein, the correlation analysis includes: determining a combination correction value according to the combination risk superposition rule when high-risk features and manual intervention features appear in combination or not; and traversing the data dependencies between each step in the abstract syntax tree, constructing a data flow graph between steps, applying a transmission coefficient to the data flow edges of the data flow graph, and determining a transmission correction value.

[0028] The strategy decision module is configured to calculate a comprehensive risk score based on the risk feature vector and the correlation analysis results, and determine a target sandbox strategy from multiple sandbox strategies with different isolation strengths based on the comparison result of the comprehensive risk score and a preset threshold, wherein the comprehensive risk score is the sum of the basic weighted sum, the combined correction value, and the transmission correction value;

[0029] The strategy injection module is configured to inject the sandbox boundary defined by the target sandbox strategy into the nodes of the executable workflow definition as node attributes during the process of compiling the structured configuration description into an executable workflow definition. The injection does not change the structure and execution semantics of the executable workflow definition.

[0030] The runtime management module is configured to execute the corresponding node in the isolated environment according to the injected sandbox boundary when executing the executable workflow definition, and to generate a new policy version when the sandbox policy changes. Newly started execution instances use the new policy version, and the execution instance currently executing locks the policy version it was created with until execution is complete.

[0031] The embodiments of the present invention have the following advantages: By parsing the structured configuration description into an abstract syntax tree and extracting and quantifying four types of risk characteristics, the automatic derivation from business configuration description to security isolation policy is realized. The configuration writer does not need to configure any security parameters, eliminating security risks caused by human negligence and solving the problem of inconsistency between security policy and business configuration.

[0032] The correlation analysis uses a combination of risk superposition and risk transmission calculation, rather than a simple feature weighted summation, to make the risk assessment more consistent with the actual risk situation. In one embodiment, the high-risk identification accuracy reaches more than 95%, the false alarm rate is controlled within 8%, and can be further reduced to 3% after calibration.

[0033] The comprehensive risk score is mapped to multiple sandbox strategies with different isolation strengths, achieving a precise match between isolation strength and task risk. This avoids the performance loss caused by indiscriminate strict isolation. In one embodiment, compared to the approximately 30% performance loss of globally enabling the strictest strategy, this solution eliminates the overhead for most low-risk tasks.

[0034] Sandbox boundaries are injected at compile time in the form of node attributes, ensuring that the sandbox boundaries of each step are deterministic and auditable, without changing the workflow structure and execution semantics, and the binding relationship between strategies and business logic is deterministic and traceable.

[0035] The policy versioning hot update mechanism enables policy changes to take effect within minutes without interrupting the running instances. The running instances (including instances that have been running for a long time while awaiting manual approval) are strictly locked to the policy version at the time of their creation, and reference count management and one-click rollback are supported.

[0036] The sandbox execution engine supports automatic switching between enhanced and degraded modes. Even in environments lacking kernel-level isolation mechanisms, it can still provide soft resource limits and explicit alerts, improving the system's environmental adaptability.

[0037] Furthermore, the method of this invention can also be applied to the security management of AI model training and inference tasks (such as multimodal channel dynamic compression feedback) in intelligent mobile communication scenarios. While ensuring the security of sensitive data such as channel data, it avoids unnecessary losses to model training and inference performance caused by indiscriminate isolation, and has good application scenario scalability. Attached Figure Description

[0038] Figure 1 This is a flowchart of the steps of an automatic inference method for artificial intelligence task sandbox according to the present invention; Figure 2 This is a schematic diagram of the structure of an AI task sandbox strategy automatic derivation system according to the present invention. Detailed Implementation

[0039] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0040] It should be noted that the above detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0041] The system in this embodiment of the invention is deployed on a server and logically consists of four layers: The first layer is a task compiler, responsible for parsing the AI ​​task configuration description written in structured configuration description languages ​​such as YAML and JSON into an Abstract Syntax Tree (AST); the second layer is a semantic analysis and policy decision layer, where the semantic analysis part receives the AST, extracts risk features, and performs correlation analysis and risk propagation calculation between features; the policy decision part calculates a comprehensive risk score based on the risk feature vector and correlation analysis results and selects a sandbox policy; the policy decision part also embeds a scoring calibrator to dynamically adjust weights and thresholds based on runtime data feedback; the third layer is a sandbox execution engine, responsible for executing tasks in an isolated environment according to the sandbox boundaries, supporting enhancement and degradation modes; the fourth layer is a hot update manager, responsible for managing the entire lifecycle of the sandbox policy, including policy version generation, locking, archiving, and rollback. After the sandbox policy is determined, the sandbox boundaries are injected into the executable workflow definition during the compilation phase, a process that is completely transparent to the configuration writer.

[0042] In the following embodiments, YAML is used as a typical example of a structured configuration description language. The embodiments of the present invention are also applicable to other structured configuration description languages ​​such as JSON. Artificial intelligence tasks are also referred to as Skills in this document, and the two have the same meaning.

[0043] Example 1

[0044] This embodiment uses a refund processing task in an e-commerce scenario as an example to fully illustrate the execution process of the AI ​​task sandbox strategy automatic inference method. The refund processing task is written in YAML format by the configuration writer. Its business logic is as follows: First, query order data based on the order identifier and save the query result as an order variable; then, determine whether the order amount is greater than 5000 yuan or whether the corresponding membership level is high. If so, it first undergoes manual approval at the financial administrator level (waiting timeout set to 3600 seconds). After approval, the bank transfer refund tool is called, with transfer parameters taken from the bank account field and amount field in the order variable, and the execution result is saved as a result variable. If not, the original payment refund tool is called to process the refund and the execution result is saved as a result variable; finally, a response containing the status field in the result variable is returned to the requester. The specific steps are as follows: Step S1: Task semantic parsing and risk feature extraction.

[0045] After receiving the above YAML configuration description, the server first parses it into an abstract syntax tree, and then traverses each node of the abstract syntax tree to extract four types of risk characteristics: tool call characteristics, including call type, parameter source and whether it is a sensitive operation; control flow characteristics, including conditional branching, looping and parallelism; data flow characteristics, including input source, processing process and output target; and manual intervention characteristics, including approval nodes and their approval levels.

[0046] The extracted features are quantized into standardized risk feature vectors according to preset quantization rules. In this embodiment, the preset quantization rules adopt the following feature dimension quantization table:

[0047] The weight ratios in the aforementioned quantification rules are determined based on the Analytic Hierarchy Process (AHP) and calibrated using historical task execution data. Specifically, a three-layer feature hierarchy is constructed: the target layer is risk scoring, the criterion layer comprises four types of features: tool calls, data flow, control flow, and human intervention, and the solution layer consists of each specific feature item. A judgment matrix is ​​constructed to perform pairwise comparisons (e.g., comparing the relative importance of call type risk value and call quantity), the weight vector is calculated, and a consistency check is performed. The consistency ratio (CR) is less than 0.1 to ensure the rationality of the weight allocation.

[0048] The maintenance of the tool risk weight dimension table adopts a configurable mechanism: the system provides a default weight table, and the business side can customize and override it through configuration files; when adding a new tool, an initial weight is automatically assigned according to the category of the service to which the tool belongs (financial, data, communication, internal), and can be manually adjusted later.

[0049] For the refund processing task in this embodiment, the risk characteristics extracted and quantified are as follows: In terms of call type risk value, querying an order is worth 10 points, bank transfer refunds are assigned 35 points as financial tools, and refunds to the original payment method are worth 15 points; the number of calls is 3 tool calls, each worth 2 points, for a total of 6 points; bank transfer refunds contain sensitive operation identifiers, worth 20 points; there is order identifier input from the user terminal, external data input is worth 10 points; the refund result contains financial information, sensitive data output is worth 15 points; it contains conditional branches, conditional branch complexity is worth 3 points; it contains approval nodes at the financial administrator level, approval nodes are worth -20 points; this approval level is not at the director level, and there is no additional reduction for the approval level item, worth 0 points.

[0050] Step S2: Correlation analysis and comprehensive risk score calculation.

[0051] After feature extraction is completed, correlation analysis between features is performed. It should be noted that the comprehensive risk score is not a simple weighted sum of the features, but also needs to consider the combined effect between features and the transmission of risk along the data flow, specifically including the following three types of nonlinear correlation calculations.

[0052] First, there is the combination of risks. When high-risk characteristics and manual intervention characteristics do not occur together, i.e., when there is a high-risk operation but no approval node, the risk value corresponding to the high-risk characteristic is multiplied by an amplification factor for non-linear amplification, with the amplification factor ranging from 1.1 to 1.5. When high-risk characteristics and manual intervention characteristics occur together, the risk value corresponding to the high-risk characteristic is multiplied by a mitigation factor. The mitigation factor is stratified according to the matching degree between the approval level and the operational risk: the mitigation factor for the first approval level (e.g., department director level) ranges from 0.4 to 0.8, and the mitigation factor for the second approval level (e.g., team leader level, manager level) ranges from 0.7 to 0.9. The higher the approval level, the more significant the risk mitigation effect. For example, if a task includes a bank transfer operation with a risk value of 40 points but no approval step, a magnification factor of 1.25 is used. The combined risk of this operation is 40 multiplied by 1.25, which equals 50 points, not 40 points. If a task includes both a bank transfer operation with a risk value of 40 points and a director-level approval step, a mitigation factor of 0.6 is used. The combined risk of this operation is 40 multiplied by 0.6, which equals 24 points, not the 20 points obtained by simple subtraction. This non-linear combination calculation better reflects the actual risk situation—the risk difference between a high-risk operation with approval and a high-risk operation without approval is not a simple addition or subtraction relationship.

[0053] Second, risk propagation calculation. In multi-step tasks, when the output of a preceding step becomes the input of a subsequent step, risk propagates along the data flow. Specifically, the data dependencies between the output save field (represented as the `save_as` field in YAML configuration) and the input parameter field (represented as the `params` field) of each step in the abstract syntax tree are traversed. A data flow graph is constructed between steps, with steps as nodes and data dependencies as edges. A propagation coefficient is applied to the data flow edges of the data flow graph to calculate the propagation risk. The propagation coefficient ranges from 0.8 to 1.5: when data flows from a lower-risk step to a higher-risk step, the risk is amplified, and the propagation coefficient is greater than 1; when data remains within steps of the same level, the propagation coefficient is approximately equal to 1. For example, step A queries sensitive data with a risk of 20 points. Step B sends the output of step A via email with its own risk of 20 points. Since the sensitive data flows from the query step to the outward sending step, a propagation coefficient of 1.3 is taken. Therefore, the propagation risk of step B is 20 multiplied by 1.3, which equals 26 points, not 20 points.

[0054] Third, handling feature conflicts. When high-risk operations and approval nodes coexist, the mitigation effect of approval depends on the matching degree between the approval level and the operational risk. Lower-level approvals have a weaker mitigation effect on high-risk operations, while higher-level approvals have a stronger mitigation effect. This is the stratified value of the mitigation coefficient mentioned above.

[0055] The initial values ​​of the aforementioned amplification coefficient, mitigation coefficient, and transmission coefficient are determined as follows: The amplification coefficient is based on the principle of strict risk management for high-risk operations, meaning that when there are no corresponding mitigation measures for a high-risk feature, the risk should be higher than the single feature value. The initial amplification coefficient is determined by the importance ratio of high-risk without mitigation to a single high-risk feature in the analytic hierarchy process matrix. The mitigation coefficient is based on the qualitative requirement that access control measures effectively reduce risk, and is stratified according to the matching degree between the approval level and the operational risk level. The transmission coefficient is determined based on the attenuation or enhancement law of data sensitivity along the data flow. All correlation coefficients can be dynamically adjusted by the scoring calibrator based on operational data feedback. The adjustment process generates version records, and historical versions are retained for traceability.

[0056] For the refund processing task in this embodiment, the correlation analysis results are as follows: First, a combined feature is identified: bank transfer refund (35 points) and the approval node coexist. The approval level is the financial administrator level, which belongs to the second approval level. Taking the mitigation coefficient of 0.8, the combined risk of bank transfer refund is 35 multiplied by 0.8, which equals 28 points, replacing the original 35 points. Second, the risk transmission path is identified: the output order variable of the query order step is used as the parameter of the bank transfer refund step. There is data dependency transmission. The two belong to the same level transmission within the same process. Taking the transmission coefficient of 1.0, the transmission risk increment is 0 points.

[0057] The comprehensive risk score is calculated by summing the basic weighted sum, combined corrections, and transmission corrections. In this embodiment, for multiple tool calls, the risk of the main tool is taken according to the principle of maximum risk, and the risks of the remaining tools are calculated at 50% discount: the risk of the main tool is 28 points after combining bank transfer refunds, and the risk of the remaining tools is 10 points for order inquiries and 15 points for original payment refunds, which is 10 multiplied by 0.5 plus 15 multiplied by 0.5 equals 12.5 points, and the total score for tool call items is approximately 40 points. The comprehensive risk score equals 40 points for tool call items, plus 6 points for the number of calls, plus 20 points for sensitive operation identifiers, plus 10 points for external data input, plus 15 points for sensitive data output, plus 3 points for conditional branches, minus 20 points for approval nodes, totaling 74 points. Furthermore, considering the mitigation effect of approval nodes on the overall risk of the task (the entire task has an approval fallback), an overall mitigation coefficient of 0.88 is applied to the score, and the final comprehensive risk score is approximately 74 multiplied by 0.88, which is approximately 65 points.

[0058] Step S3: Dynamic strategy selection and generation.

[0059] The comprehensive risk score is compared with a preset threshold to determine the target sandbox strategy from multiple sandbox strategies with different isolation strengths. In this embodiment, the multiple sandbox strategies include four levels of strategies with decreasing isolation strength: The first strategy (strict strategy) is selected when the comprehensive risk score is greater than or equal to the first threshold. This includes process isolation, network isolation, file system isolation, and detailed auditing, suitable for high-risk operations and external interaction scenarios. The second strategy (standard strategy) is selected when the comprehensive risk score is greater than or equal to the second threshold but less than the first threshold. This includes resource restrictions and disabling direct network access and forwarding via an internal proxy, suitable for sensitive operations and routine business data processing scenarios. The third strategy (controlled network strategy) is selected when the comprehensive risk score is less than the second threshold and the task requires network access. This includes controlled network access, suitable for low-risk scenarios requiring network access. Otherwise, the fourth strategy (relaxed strategy) is selected, which includes basic resource restrictions, suitable for trusted internal tool invocation scenarios. In this embodiment, the first threshold is 70, and the second threshold is 40.

[0060] The determination of the first and second thresholds is based on a dual perspective: On the one hand, it is deduced from the hierarchical security requirements that high-risk operations must have complete isolation capabilities, including process isolation, network isolation, file system isolation, and detailed auditing, corresponding to the first strategy; while regular operations must have basic isolation capabilities, including resource restrictions, network control, and audit logs, corresponding to the second strategy. On the other hand, it is verified by statistical analysis of historical task operation data. Risk scores were calculated and manual security assessments were performed on more than 120 registered tasks on the platform. The statistical results show that the scores of tasks manually marked as high-risk all fall between 65 and 95, the scores of regular-risk tasks fall between 35 and 65, and the scores of low-risk tasks fall between 5 and 35. When the thresholds are set to 70 and 40, the accuracy rate of high-risk identification reaches more than 95%, and the false alarm rate is controlled within 8%.

[0061] For the refund processing task in this embodiment, the overall risk score is approximately 65 points, falling into the medium-to-high risk range, which maps to the second strategy, the standard strategy. The generated target sandbox strategy includes: resource restrictions, with a processor time limit of 30 seconds, a memory limit of 256 MB, and a maximum number of processes of 10; network policy, since the task involves bank operations, direct network access is disabled, the list of allowed hosts is empty, and access to the bank interface is forwarded through an internal proxy; file system policy, the configured directory is a read-only path, a temporary directory isolated by the task identifier is used as the read-write path, and a private temporary directory and a private device directory are enabled; the audit level is detailed. There is a design trade-off here: the bank interface access is disabled via direct network and forwarded through an internal proxy because if network access were directly opened, processes within the sandbox could theoretically send requests to any address, leading to uncontrollable risks, while forwarding through a proxy allows for whitelist filtering at the proxy layer.

[0062] It should be noted that there may be overlapping scenarios between the second and third strategies. For example, if a task needs to call an external interface and involves sensitive data processing, the second strategy should be used to isolate the network first, and external access should be forwarded through an internal proxy.

[0063] Step S4, transparent injection at compile time.

[0064] When compiling the structured configuration description of a task into an executable workflow definition, the sandbox boundaries defined by the target sandbox strategy are directly embedded into each node of the executable workflow definition as node attributes. The injection point selection rules are as follows: for tool call type nodes, inject the complete sandbox boundaries, including entry resource checks, exit audit points, and exception handling points; for conditional branch type nodes, inject branch path audit points to record branch paths; for manual approval type nodes, inject wait timeout policies and approval status verification; for response type nodes, inject output content security check points.

[0065] The default injection granularity is step-level injection, which means that each tool call step is wrapped with a sandbox execution boundary. When there are steps with different risk levels in the same task (such as a low-risk query step and a high-risk transfer step), sub-step-level injection is used to give the higher-risk step a sandbox boundary with higher isolation strength, and the lower-risk step a sandbox boundary with lower isolation strength, thus avoiding indiscriminate isolation.

[0066] The integration of injection and workflow graph is as follows: sandbox boundaries are embedded into each node of the workflow definition as node attributes, without changing the structure and execution semantics of the workflow graph; the injected sandbox information participates in workflow scheduling. The scheduler reads the sandbox attributes of the node before execution, creates the corresponding isolation environment, and reclaims the isolation environment after the node execution is completed; the injected workflow definition ensures that the execution semantics remain unchanged, the sandbox boundaries are appended attributes, do not modify the business logic definition of the node, and the control flow and data flow of the workflow remain unchanged. The business logic of the task itself is not modified in any way.

[0067] This invention chooses compile-time injection instead of runtime dynamic judgment because: compile-time injection ensures that the sandbox boundaries of each step are deterministic and auditable; while runtime dynamic judgment is more flexible, it has poor traceability and is difficult to troubleshoot when problems occur. The compile-time injection mechanism ensures that the binding relationship between the policy and business logic is deterministic and auditable, and the security policy is automatically updated synchronously after the task configuration changes, avoiding the inconsistency of the sandbox configuration remaining the old version even though the task has changed.

[0068] For the refund processing task in this embodiment, the compilation stage embeds the standard strategy into the workflow definition: the sandbox attributes of the query order node include the strategy name as standard strategy, timeout duration of 30 s, and memory limit of 256 MB; the sandbox attributes of the bank transfer node also include a flag that disables network access.

[0069] Step S5: Runtime isolation execution and hot update management.

[0070] When executing the executable workflow definition, each node runs in an isolated environment according to the injected sandbox boundary. The execution process is as follows: First, it checks whether the operating system kernel-level isolation mechanism is available. If it is available, it builds a sandbox to execute commands in enhanced mode; otherwise, it sets resource limits in degraded mode. Then, it starts a child process to execute the node logic while monitoring resource usage. If a timeout or limit exceedance occurs, the child process is automatically terminated. After execution, the execution results and audit logs are collected.

[0071] After a task configuration change, a recompilation is triggered, generating a new policy version. Newly started execution instances use the new policy version, while currently executing instances continue to use the policy version they were created with until completion, at which point the old version is archived. Long-running instances (e.g., those awaiting manual approval) strictly maintain a locked policy version. Audit logs record the entire change process, including the subject of the change, the content of the change, and the effective time. The system supports policy rollback; problematic policy versions can be rolled back with a single click, and newly started instances automatically use the rolled-back version.

[0072] In this embodiment, when a requester triggers a refund, the system executes the task in a sandbox: first, it queries the order, determines whether approval is required based on the amount, and then executes the refund operation in an isolated environment, recording audit logs throughout the process. No security parameters were configured by the system's developer; high-risk operations automatically received appropriate isolation, approval nodes had their risk scores appropriately reduced, and the security control requirements for financial transactions were met. Resource restrictions, access control, and audit logs are all built into the system and do not rely on external components.

[0073] Example 2

[0074] This embodiment illustrates the execution process of hot updating of sandbox policies. The difference between this embodiment and Embodiment 1 is that it focuses on describing policy change scenarios.

[0075] The scenario is as follows: Security management discovers a data breach risk with an order query tool and needs to tighten the sandbox policy for all tasks using this tool. The handling process is as follows.

[0076] First, the system scans all registered tasks and identifies those involving order query tools, including refund processing tasks, order status query tasks, and customer profile analysis tasks.

[0077] Subsequently, the security management team increased the risk value of the order query tool in the tool risk weight dimension table from 10 to 25. The system automatically triggered the recompilation and re-scoring of the above three tasks: the comprehensive risk score of the refund processing task increased from 65 to 75, exceeding the first threshold of 70, and the sandbox strategy was upgraded from the second strategy to the first strategy; the score of the order status query task increased from 40 to 60, the strategy remained the second strategy but the parameters were tightened; the score of the customer profile analysis task increased from 55 to 70, and the strategy was also upgraded to the first strategy.

[0078] The hot update takes effect as follows: newly received refund requests use the first strategy, with resource limits tightened to 10 seconds of processor time and 128 MB of memory; requests currently being processed continue to use the original second strategy until completion. Specifically, execution instances in the refund processing task awaiting manual approval (potentially waiting for more than 2 hours) strictly maintain their original second strategy version lock, unaffected by the new first strategy. Each strategy version maintains a reference count, representing the number of execution instances currently using that version. Once the reference count reaches zero, the version is delayed in archiving rather than immediately deleted to prevent instance reference loss in edge cases. The audit log fully records this change.

[0079] Furthermore, this embodiment also includes a policy rollback scenario: After the first policy is launched, the false alarm rate increases from 5% to 20%, and a large number of legitimate refund operations are interrupted due to resource limitations. The security manager triggers a one-click rollback, and the system reverts to the previous policy version. Newly launched refund instances resume using the second policy. The rollback operation is also recorded in the audit log, and new instances automatically use the rolled-back version.

[0080] The effect of this embodiment is that the time from the security manager adjusting the weights to the new policy taking effect is in the minute range, and it does not affect the business being executed (including long-running instances); the change history is traceable and meets audit requirements.

[0081] Example 3

[0082] This embodiment illustrates the operation process of the sandbox execution engine in downgrade mode. The difference from Embodiment 1 is that the operating environment lacks a kernel-level isolation mechanism.

[0083] The scenario is as follows: In development environments or certain restricted servers where the operating system kernel-level isolation mechanism is not deployed, the sandbox execution engine needs to run in degraded mode. The processing procedure is as follows.

[0084] When the system starts up, it automatically detects whether the operating system supports kernel-level isolation mechanisms such as process namespace isolation and network namespace isolation; in this embodiment, the detection result is that it does not support them, and the system enters downgrade mode.

[0085] The sandbox execution engine switches to soft-limited mode, employing the operating system's native resource limitation mechanism. Before creating and executing child processes, resource limits are set: the maximum memory address space for the child process is set to 256 MB, and the maximum processor time is set to 30 seconds. The child process execution node logic is then started under these soft limits. Simultaneously, the system outputs alarm logs, explicitly indicating that the current kernel-level isolation mechanism is unavailable, that soft-limited resource mode is in use, and that network isolation is unavailable in this mode. It also recommends enabling kernel-level isolation support to enhance security.

[0086] The effect of this embodiment is that the system can still run normally in the absence of a kernel-level isolation mechanism; processor and memory limits take effect normally, but network isolation is unavailable; the logs clearly indicate that the current mode is degraded, allowing the operations team to decide whether to enable kernel-level isolation support. It should be noted that degraded mode is only applicable to development and testing environments; production environments must enable enhanced mode to meet the network isolation requirements of tiered security management. Furthermore, the system in this embodiment does not silently degrade but instead explicitly alerts the system to the degraded status.

[0087] It should be further noted that, in one implementation, the kernel-level isolation mechanism can be carried out by a process-level sandbox tool that encapsulates the namespace isolation capabilities of the operating system to provide process-level isolation. In terms of system architecture, the sandbox execution engine is decoupled from the specific isolation tool. This process-level sandbox tool is only one optional implementation method. It can also be replaced by directly calling the namespace separation system call provided by the operating system, or by other lightweight virtualization solutions that provide kernel-level isolation capabilities.

[0088] Example 4

[0089] This embodiment illustrates the process by which the scoring calibrator performs scoring calibration based on operational data feedback.

[0090] The scenario is as follows: After the system has been running for 3 months, the statistics show that about 8% of the tasks were assigned overly strict policies and frequently timed out, which are false alarms; at the same time, two low-risk tasks experienced resource contention events due to overly lenient policies, which are missed alarms. The handling process is as follows.

[0091] The scoring calibrator continuously records feedback data from the sandbox execution during operation: a false positive event is recorded when a legitimate operation is rejected due to overly strict policies, and a false negative event is recorded when a security incident occurs due to overly lenient policies. Based on the accumulated feedback data, the scoring calibrator automatically calculates the false positive rate and false negative rate periodically (e.g., weekly). When the false positive rate exceeds a first proportional threshold (e.g., 15%) or the false negative rate exceeds a second proportional threshold (e.g., 5%), it triggers suggestions for adjusting the weights and thresholds.

[0092] After summarizing three months of operational data, the scoring calibrator analysis revealed the following: Regarding false alarms, approximately 70% of the 8% false alarms were caused by excessively high weighting (20 points) of sensitive operation identifier features. Although the relevant tasks contained sensitive identifiers, their actual execution did not result in high-risk behavior. Regarding false negatives, both false negative tasks involved a combination of internal tool calls and external data inputs. The current scoring did not fully consider the risk superposition effect of such combinations.

[0093] Based on this, the scoring calibrator generates the following adjustment suggestions: First, reduce the weight of the sensitive operation identification feature from 20 points to 15 points; second, add a combination rule to increase the risk value of external data input from 10 points to 15 points when internal tool calls and external data inputs are combined; third, increase the processor time limit of the second strategy from 30 seconds to 45 seconds to reduce false alarms due to timeout.

[0094] After the security management reviews and approves the above adjustments, the system generates a second-parameter version of the weights and thresholds, while retaining the first-parameter version for backtracking; newly compiled tasks use the second-parameter version to calculate risk scores. This forms a closed loop of operation, feedback, calibration, new version, and re-operation, enabling the scoring system to be continuously optimized with actual operation.

[0095] The effect of this embodiment is that the false alarm rate was reduced from 8% to 3% after adjustment, the false negative rate remained unchanged at 0%, and the overall scoring accuracy was significantly improved.

[0096] Comparative Example 1

[0097] In contrast, this comparative example adopts the existing approach of using a default strategy without differentiation: without semantic risk quantification and strategy derivation, it uniformly applies the strictest isolation strategy, equivalent to the first strategy, to all tasks. Testing showed that enabling the strictest strategy globally resulted in a performance loss of approximately 30%, while the risk level of most tasks on the platform did not require such a strict isolation level, causing a large number of low-risk tasks to bear unnecessary isolation overhead. Furthermore, this approach lacks the ability to match strategy with task risk and cannot adjust the strategy without restarting the service.

[0098] The key effects of each embodiment and the comparative example are compared in the table below:

[0099] It is evident that this solution has the following significant advantages over existing technologies: By automatically deriving sandbox strategies through semantic risk quantification, the accuracy rate of high-risk identification reaches over 95%, avoiding indiscriminate strict isolation. The initial false positive rate is controlled below 8%, and can be further reduced to 3% after feedback calibration, significantly reducing the probability of legitimate operations being falsely blocked. Strict isolation is applied only to high-risk tasks, while low-risk tasks do not incur additional overhead, avoiding the approximately 30% performance loss caused by the strictest global strategy in Comparative Example 1. Policy changes support versioned hot updates without requiring a service restart and do not affect currently running instances, whereas existing technologies require a restart to take effect.

[0100] In summary, this solution achieves a precise match between security isolation strength and task risk level, ensuring high recognition rate and low false alarm rate while significantly reducing performance overhead and supporting seamless business policy changes.

[0101] Example 5

[0102] Corresponding to the above method embodiments, this invention also provides an automatic strategy derivation device for artificial intelligence task sandboxes, which includes a semantic parsing module, an association analysis module, a strategy decision-making module, a strategy injection module, and an operation management module.

[0103] The semantic parsing module is configured to obtain the structured configuration description of the artificial intelligence task, parse the structured configuration description into an abstract syntax tree, traverse the abstract syntax tree, extract tool call features, data flow features, control flow features, and human intervention features, and quantize each extracted feature into a risk feature vector according to a preset quantization rule;

[0104] The correlation analysis module is configured to perform correlation analysis on the risk feature vector to obtain correlation analysis results, including determining the combined correction value according to the combined risk superposition rule, and constructing a data flow graph based on the data dependency relationship between steps and applying a transmission coefficient to determine the transmission correction value;

[0105] The strategy decision module is configured to calculate a comprehensive risk score based on the risk feature vector and the correlation analysis results, and determine a target sandbox strategy from multiple sandbox strategies with different isolation strengths based on the comparison results of the comprehensive risk score and a preset threshold.

[0106] The strategy injection module is configured to inject the sandbox boundaries defined by the target sandbox strategy into the nodes of the executable workflow definition as node attributes during the process of compiling the structured configuration description into an executable workflow definition;

[0107] The runtime management module is configured to execute the corresponding node in the isolated environment according to the injected sandbox boundary when executing the executable workflow definition, and to perform versioned hot updates when the sandbox policy changes.

[0108] The specific details of the functions implemented by each module are the same as the corresponding steps in the aforementioned method embodiments, and will not be repeated here. Those skilled in the art will clearly understand that the module division of the above device is only a logical functional division, and there may be other division methods in actual implementation. Each module can be integrated into a processing unit, or each module can exist physically separately.

[0109] Example 6

[0110] This embodiment illustrates the application process of the method of the present invention in a mobile communication intelligent scenario. The difference from Embodiment 1 is that the type of artificial intelligence task is a multimodal channel dynamic compression feedback model training task based on AI in a 6G network intelligent platform. This task is written in YAML format by the configuration writer, and its business logic is as follows: First, the channel data acquisition interface is called to read multimodal channel sample data (including different polarization states, spatial angles, delay spread, etc.), and the acquisition results are saved as channel sample variables; then, the compression model training tool is called, using the channel sample variables as input to learn key features of the multimodal channel and train the compression model. The training objective is to reduce feedback overhead by more than 50% while ensuring feedback accuracy loss is less than 5%, and the trained model is saved as model variables; subsequently, it is determined whether the feedback accuracy loss of the model variables on the validation set is less than 5% and whether the feedback overhead compression ratio reaches more than 50%. If so, the model return tool is called to return the compressed model parameters via the base station northbound interface for deployment, in order to reduce the signaling interaction overhead between the base station and the terminal; if not, the internal log tool is called to record the training indicators and end the process; finally, a response containing the training status is returned to the requester. The specific steps are as follows: The process of task semantic parsing and risk feature extraction is the same as in Example 1. Specifically, the tool risk weight dimension table automatically assigns initial weights to newly added tools in the intelligent communication scenario according to their service category: the channel data acquisition interface belongs to the data category, with a call type risk value of 15 points; the compressed model training tool belongs to the internal calculation category, with 5 points; the model feedback tool involves write operations to the external base station system, belonging to the communication category, with 20 points; the internal log tool has 2 points. The remaining extracted and quantified risk features are: 4 tool calls, each worth 2 points, totaling 8 points; no explicit marking of sensitive operations in the task configuration, with sensitive operation identification scoring 0 points; the sensitivity of channel data is reflected in the sensitive data output items and transmission analysis described later; channel sample data comes from terminal reporting and base station measurements, belonging to external data input, scoring 10 points; information such as spatial angles in the returned compressed channel features can infer the user's location, belonging to sensitive data output, scoring 15 points; it includes conditional branches, with conditional branch complexity scoring 3 points; there are no approval nodes, with manual intervention scoring 0 points.

[0111] The results of the correlation analysis are as follows: First, regarding the combined risk, the model feedback operation outputs sensitive channel characteristics and the task has no approval node, meaning that high-risk characteristics do not appear in combination with human intervention characteristics. Taking an amplification factor of 1.15, the risk value of the model feedback is corrected from 20 points to 23 points, and the combined correction value is 3 points. Second, regarding the risk transmission, the channel sample variables flow from the self-channel data acquisition step to the compressed model training step, which is a flow within the same domain. The transmission coefficient is taken as 1.0, with no transmission increment. The model variables carry compressed channel characteristics and flow from the self-training step to the external feedback step. Sensitive data flows from the internal processing domain to the external output domain. The transmission coefficient is taken as 1.2, and the transmission correction value is 20 multiplied by 0.2, which equals 4 points. The basic weighted sum is as follows: According to the principle of maximum risk, the main tool is model feedback, which accounts for 20 points. The other tools (15 points, 5 points, and 2 points) are calculated at 50% and account for 11 points. The tool call item accounts for 31 points. Adding the number of calls (8 points), external data input (10 points), sensitive data output (15 points), and conditional branches (3 points), the basic total is 67 points. The overall risk score is equal to the sum of the basic weighted sum of 67 points, the combined correction value of 3 points, and the transmission correction value of 4 points, for a total of 74 points.

[0112] The overall risk score of 74 points is greater than the first threshold of 70, so the target sandbox strategy is determined to be the first strategy, i.e., the strict strategy, which includes process isolation, network isolation, file system isolation, and detailed-level auditing. Specifically, the northbound interface access to the base station required for model backhaul is not directly accessible under network isolation; instead, it is forwarded to a whitelisted address via an internal proxy to implement whitelist filtering at the proxy layer. Furthermore, due to the significant differences in risk levels among the various steps within this task, the model backhaul step, after combination and propagation correction, has the highest risk, while the compressed model training step has the lowest risk. The compilation stage employs sub-step-level injection: for the channel data acquisition and model backhaul steps, a complete and strict sandbox boundary is injected, including process isolation, network isolation allowing forwarding only to the whitelisted addresses of the base station's northbound interface via internal agents, file system isolation, and detailed auditing, with resource restrictions of 60 s processor time and 512 MB memory; for the compressed model training step, a sandbox boundary with lower isolation strength is injected, with resource restrictions of 7200 s processor time and 8 GB memory, allowing the use of model training acceleration resources and disabling network access to avoid unnecessary performance loss to model training throughput caused by strict isolation; branch path audit points are injected into conditional branch nodes, and output content security checkpoints are injected into response nodes.

[0113] The effects of this embodiment are as follows: No security parameters are configured by the configuration developer; multimodal channel data containing inferable user location information is processed entirely in an isolated environment; external transmission can only reach the base station's northbound interface via a whitelisted proxy; the compressed model training step does not incur indiscriminate strict isolation overhead, model training completes normally, and meets the compression goals of less than 5% feedback accuracy loss and more than 50% reduction in feedback overhead. This embodiment demonstrates that the method of this invention is also applicable to the security management of AI model training and inference tasks in intelligent mobile communication scenarios, achieving a balance between channel data security and task execution performance.

[0114] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0115] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that the embodiments of this application described herein can be implemented in sequences other than those illustrated or described herein.

[0116] Furthermore, the terms “comprising” and “having”, and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product, or apparatus.

[0117] For ease of description, spatial relative terms such as "above," "on top of," "on the upper surface of," "above," etc., are used herein to describe the spatial positional relationship of a device or feature as shown in the figures to other devices or features. It should be understood that spatial relative terms are intended to encompass different orientations in use or operation beyond the orientation of the device as described in the figures. For example, if the device in the figures were inverted, a device described as "above" or "on top of" other devices or structures would subsequently be positioned as "below" or "under" other devices or structures. Thus, the exemplary term "above" can include both "above" and "below." The device may also be positioned in other different ways, such as rotated 90 degrees or in other orientations, and the spatial relative descriptions used herein will be interpreted accordingly.

[0118] In the detailed description above, reference has been made to the accompanying drawings, which form part of this document. In the drawings, similar symbols typically identify similar parts unless the context otherwise indicates otherwise. The illustrated embodiments described in the detailed specification, drawings, and claims are not intended to be limiting. Other embodiments may be used and other changes may be made without departing from the spirit or scope of the subject matter presented herein.

[0119] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for automatically deriving sandbox strategies for artificial intelligence tasks, characterized in that, The method includes: Obtain a structured configuration description of the artificial intelligence task, parse the structured configuration description into an abstract syntax tree, traverse the abstract syntax tree, extract tool call features, data flow features, control flow features, and human intervention features, and quantize each extracted feature into a risk feature vector according to a preset quantization rule; A correlation analysis is performed on the risk feature vector to obtain the correlation analysis results; wherein, the correlation analysis includes: determining the combination correction value according to the combination risk superposition rule when high-risk features and manual intervention features appear in combination or not; and traversing the data dependencies between each step in the abstract syntax tree, constructing a data flow graph between steps, applying a transmission coefficient to the data flow edges of the data flow graph, and determining the transmission correction value. A comprehensive risk score is calculated based on the risk feature vector and the correlation analysis results. Then, based on a comparison of the comprehensive risk score with a preset threshold, a target sandbox strategy is determined from multiple sandbox strategies with different isolation strengths. The comprehensive risk score is the sum of a weighted average, the combined correction value, and the transmission correction value. During the compilation of the structured configuration description into an executable workflow definition, the sandbox boundary defined by the target sandbox strategy is injected into the nodes of the executable workflow definition in the form of node attributes. This injection does not change the structure and execution semantics of the executable workflow definition. When executing the executable workflow definition, the corresponding node is executed in the isolated environment according to the injected sandbox boundary; and when the sandbox policy changes, a new policy version is generated, the newly started execution instance uses the new policy version, and the execution instance being executed locks the policy version it was created until execution is completed.

2. The method for automatically deriving sandbox strategies for artificial intelligence tasks according to claim 1, characterized in that... The preset quantification rules include: the call type risk value in the tool call feature is assigned according to the tool risk weight dimension table, and different risk categories of tools correspond to different call type risk values; the feature value corresponding to the approval node in the manual intervention feature is assigned a negative value, and the higher the approval level, the larger the absolute value of the negative value.

3. The method for automatically deriving sandbox strategies for artificial intelligence tasks according to claim 2, characterized in that... The weights of each feature in the preset quantification rules are determined based on the analytic hierarchy process (AHP), including: constructing a three-layer feature hierarchy structure, wherein the target layer is the risk score, the criterion layer is the tool call feature, the data flow feature, the control flow feature and the human intervention feature, and the scheme layer is each specific feature item; by constructing a judgment matrix, each feature is compared pairwise, the weight vector is calculated and a consistency check is performed, and the consistency ratio is less than 0.

1.

4. The method for automatically deriving sandbox strategies for artificial intelligence tasks according to claim 1, characterized in that... The combined risk superposition rule includes: when a high-risk feature does not appear in combination with a feature requiring manual intervention, the risk value corresponding to the high-risk feature is multiplied by an amplification factor, wherein the amplification factor ranges from 1.1 to 1.5; When a combination of high-risk characteristics and manual intervention characteristics occurs, the risk value corresponding to the high-risk characteristic is multiplied by a mitigation coefficient. The mitigation coefficient is stratified according to the matching degree between the approval level and the operational risk. The mitigation coefficient corresponding to the first approval level ranges from 0.4 to 0.8, and the mitigation coefficient corresponding to the second approval level, which is lower than the first approval level, ranges from 0.7 to 0.

9.

5. The method for automatically deriving sandbox strategies for artificial intelligence tasks according to claim 1, characterized in that... The construction of the data flow graph between the steps includes: traversing the data dependencies between the output storage fields and input parameter fields of each step in the abstract syntax tree, and constructing the data flow graph with steps as nodes and data dependencies as edges; the transmission coefficient ranges from 0.8 to 1.5, and the transmission coefficient is greater than 1 when data flows from a lower-risk step to a higher-risk step.

6. The method for automatically deriving sandbox strategies for artificial intelligence tasks according to claim 1, characterized in that... The multiple sandbox strategies include a first strategy, a second strategy, a third strategy, and a fourth strategy; the preset threshold includes a first threshold and a second threshold, wherein the first threshold is greater than the second threshold. The step of determining the target sandbox strategy from multiple sandbox strategies with different isolation strengths includes: when the comprehensive risk score is greater than or equal to the first threshold, determining the first strategy as the target sandbox strategy, wherein the first strategy includes process isolation, network isolation, file system isolation, and detailed level auditing; When the comprehensive risk score is greater than or equal to the second threshold and less than the first threshold, the second strategy is determined to be the target sandbox strategy. The second strategy includes resource restrictions and disabling direct network access and forwarding it through an internal proxy. When the overall risk score is less than the second threshold and the artificial intelligence task requires network access, the third strategy is determined to be the target sandbox strategy, which includes controlled network access; otherwise, the fourth strategy is determined to be the target sandbox strategy, which includes basic resource restrictions.

7. The method for automatically deriving sandbox strategies for artificial intelligence tasks according to claim 1, characterized in that... The step of injecting the sandbox boundary defined by the target sandbox strategy into the nodes of the executable workflow definition in the form of node attributes includes: for tool invocation nodes, injecting the complete sandbox boundary including entry resource checks, exit audit points, and exception handling points; For conditional branch nodes, inject branch path audit points; for manual approval nodes, inject wait timeout policies and approval status verification. For each response node, inject a security checkpoint into the output content; The injection granularity is set to step-level injection by default. When there are steps with different risk levels in the same artificial intelligence task, sub-step-level injection is used so that steps with higher risk are injected into sandbox boundaries with higher isolation strength, and steps with lower risk are injected into sandbox boundaries with lower isolation strength.

8. The method for automatically deriving sandbox strategies for artificial intelligence tasks according to claim 1, characterized in that... The step of executing the corresponding node in the isolated environment according to the injected sandbox boundary includes: detecting whether the operating system kernel-level isolation mechanism is available when the system starts; if the kernel-level isolation mechanism is available, entering enhanced mode, and implementing process isolation, network isolation and file system isolation through the kernel-level isolation mechanism; if the kernel-level isolation mechanism is unavailable, entering degraded mode, softly limiting processor time and memory through the operating system resource limiting mechanism, and outputting an alarm log indicating that it is currently in degraded mode; wherein, the enhanced mode and the degraded mode are automatically switched.

9. The method for automatically deriving sandbox strategies for artificial intelligence tasks according to claim 1, characterized in that... The method further includes: recording feedback data of sandbox execution, the feedback data including false positive events where legitimate operations are rejected due to overly strict sandbox policies and false negative events where security events are missed due to overly lenient sandbox policies; periodically calculating the false positive rate and false negative rate based on the accumulated feedback data, and generating adjustment suggestions for the weights and preset thresholds in the preset quantization rules when the false positive rate is greater than a first proportional threshold or the false negative rate is greater than a second proportional threshold; wherein each adjustment of the weights and preset thresholds generates a new parameter version, and historical parameter versions are retained to support rollback.

10. An automatic strategy derivation device for an artificial intelligence task sandbox, characterized in that... The device includes: The semantic parsing module is configured to obtain the structured configuration description of the artificial intelligence task, parse the structured configuration description into an abstract syntax tree, traverse the abstract syntax tree, extract tool call features, data flow features, control flow features, and human intervention features, and quantize each extracted feature into a risk feature vector according to a preset quantization rule; The correlation analysis module is configured to perform correlation analysis on the risk feature vector to obtain correlation analysis results; wherein, the correlation analysis includes: determining a combination correction value according to the combination risk superposition rule when high-risk features and manual intervention features appear in combination or not; and traversing the data dependencies between each step in the abstract syntax tree, constructing a data flow graph between steps, applying a transmission coefficient to the data flow edges of the data flow graph, and determining a transmission correction value. The strategy decision module is configured to calculate a comprehensive risk score based on the risk feature vector and the correlation analysis results, and determine a target sandbox strategy from multiple sandbox strategies with different isolation strengths based on the comparison result of the comprehensive risk score and a preset threshold, wherein the comprehensive risk score is the sum of the basic weighted sum, the combined correction value, and the transmission correction value; The strategy injection module is configured to inject the sandbox boundary defined by the target sandbox strategy into the nodes of the executable workflow definition as node attributes during the process of compiling the structured configuration description into an executable workflow definition. The injection does not change the structure and execution semantics of the executable workflow definition. The runtime management module is configured to execute the corresponding node in the isolated environment according to the injected sandbox boundary when executing the executable workflow definition, and to generate a new policy version when the sandbox policy changes. Newly started execution instances use the new policy version, and the execution instance currently executing locks the policy version it was created with until execution is complete.

Citation Information

Patent Citations

  • Dynamic generation and control method and device for policy-driven data security sandbox

    CN121935904A

  • Audible workflow of audible network security agent and construction method of audible workflow

    CN122020638A