A sandbox resource processing method and system
Patent Information
- Application Number
- CN202611234285.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-14
- Publication Date
- 2026-09-25
AI Technical Summary
这样,当沙箱因创建失败、进程异常崩溃或网络分区等原因脱离正常管理流程时,会导致大量无主沙箱持续占用如CPU、内存、存储等物理资源,造成资源泄漏和系统稳定性下降
[0042]由此可见,本申请提出的沙箱资源处理方法及系统,通过从资源编排平台获取实际运行的沙箱实例的第一实例信息,与中央状态存储中记录的合法沙箱实例的第二实例信息进行跨层比对,不依赖沙箱自身上报,主动精准识别未被任何记录追踪的未记录沙箱实例(即物理上存在但逻辑上无合法记录的孤儿沙箱实例),解决了因创建失败、进程崩溃或网络分区等原因导致的无主沙箱实例难以被发现的问题。之后,本申请将对识别出的未记录沙箱实例执行清理操作,及时释放被占用的物理资源,避免资源持续泄漏,提升系统资源利用率,且本申请通过自动回收异常沙箱实例,防止残留沙箱干扰正常任务执行,降低因资源耗尽或状态混乱导致的系统崩溃风险,提高了系统资源的稳定性。本申请无需人工介入,可周期性或事件触发执行清理,显著降低运维成本,提高了管理效率。
Smart Images

Figure CN122816772A_ABST
Abstract
Description
Technical Field
[0001] This application mainly relates to the field of artificial intelligence technology, and more specifically to a sandbox resource processing method and system. Background Technology
[0002] With the rapid development of AI (Artificial Intelligence) agent technology, it has become possible for agents to perform complex tasks such as code compilation, web page automation, and data analysis in isolated environments. Container technologies (such as Docker) and container orchestration technologies (such as Kubernetes) provide the foundation for building such isolated sandbox environments.
[0003] Currently, sandbox environments for intelligent agents are typically provided on an on-demand, real-time basis, with sandbox instances relying on themselves to register with the system and maintain their lifecycle state. This can lead to a situation where, when a sandbox fails to create, crashes abnormally, or is separated from the normal management process due to network partitions, a large number of unattended sandboxes continuously consume physical resources such as CPU, memory, and storage, resulting in resource leaks and decreased system stability. Summary of the Invention
[0004] To address the above problems, this application provides the following solution:
[0005] The first aspect of this application provides a sandbox resource processing method, comprising:
[0006] Get the first instance information of all sandbox instances currently running on the resource orchestration platform;
[0007] Obtain the second instance information of all legitimate sandbox instances recorded in the central state storage, wherein the legitimate sandbox instances include sandbox instances in the allocated state and sandbox instances in the unallocated state;
[0008] The first instance information is compared with the second instance information to obtain the unrecorded sandbox instances in the first instance information that are not recorded in the second instance information;
[0009] Perform a cleanup operation on the unrecorded sandbox instance.
[0010] Optionally, the method further includes:
[0011] When a sandbox instance is created, its identification information is recorded in the third instance information stored in the central state storage.
[0012] After the sandbox instance is created and enters the unassigned state, the identification information is removed from the third instance information and added to the second instance information.
[0013] Optionally, before performing the cleanup operation, the following may also be included:
[0014] Obtain the creation time information of the unrecorded sandbox instance; determine the creation duration of the unrecorded sandbox instance based on the creation time information;
[0015] If there are unrecorded sandbox instances whose creation time is less than a preset time threshold, no cleanup operation will be performed on the unrecorded sandbox instances.
[0016] Optionally, the method further includes:
[0017] In response to a business request initiated by an agent to a sandbox instance, the active time information of the corresponding sandbox instance in the central state storage is updated according to the task identifier corresponding to the business request.
[0018] Based on the active time information, determine the sandbox instances that have entered the inactive state from among the sandbox instances that were in the assigned state;
[0019] Clean up the sandbox instances that have entered an inactive state.
[0020] Optionally, the method further includes:
[0021] Obtain the number of sandbox instances in the unassigned state from the second instance information;
[0022] If the quantity is lower than the preset water level, create a new sandbox instance.
[0023] Optionally, the method further includes:
[0024] In response to the sandbox retrieval request, a sandbox instance in an unallocated state is retrieved from the second instance information;
[0025] The status of the acquired sandbox instance is updated to the assigned status, and the association information between the sandbox instance and the task identifier is recorded.
[0026] Optionally, the method further includes:
[0027] Obtain the network location information of the sandbox instance and write it into the location record of the central state storage;
[0028] If the association information recording fails or the location record writing fails, the identification information of the corresponding sandbox instance will be re-added to the second instance information;
[0029] If it fails to re-add the sandbox instance's identifier information to the second instance information, the application programming interface of the resource orchestration platform is invoked to destroy the sandbox instance.
[0030] Optionally, in response to a sandbox acquisition request, the method further includes: if there is no sandbox instance in the central state storage that is in an unallocated state, creating a corresponding sandbox instance for the task that initiated the sandbox acquisition request.
[0031] Optionally, the method further includes:
[0032] In response to an access request for a target sandbox instance, determine the task identifier carried in the access request;
[0033] Based on the allocation records in the central state storage, determine whether there is an association between the task identifier and the target sandbox instance;
[0034] If a relationship exists, the access request is permitted.
[0035] A second aspect of this application provides a sandbox resource processing system, the system comprising:
[0036] Processing unit, and central state storage;
[0037] The processing device is configured as follows:
[0038] Get the first instance information of all sandbox instances currently running on the resource orchestration platform;
[0039] Obtain the second instance information of all legitimate sandbox instances recorded in the central state storage, wherein the legitimate sandbox instances include sandbox instances in the allocated state and sandbox instances in the unallocated state;
[0040] The first instance information is compared with the second instance information to obtain the unrecorded sandbox instances in the first instance information that are not recorded in the second instance information;
[0041] Perform a cleanup operation on the unrecorded sandbox instance.
[0042] Therefore, the sandbox resource processing method and system proposed in this application, by obtaining the first instance information of the actually running sandbox instance from the resource orchestration platform and comparing it with the second instance information of the legitimate sandbox instances recorded in the central state storage, proactively and accurately identifies unrecorded sandbox instances (i.e., orphan sandbox instances that physically exist but logically lack legitimate records) without relying on the sandbox itself for reporting. This solves the problem of the difficulty in discovering ownerless sandbox instances caused by creation failures, process crashes, or network partitions. Subsequently, this application performs a cleanup operation on the identified unrecorded sandbox instances, promptly releasing occupied physical resources, preventing continuous resource leakage, and improving system resource utilization. Furthermore, by automatically reclaiming abnormal sandbox instances, this application prevents residual sandboxes from interfering with normal task execution, reduces the risk of system crashes due to resource exhaustion or state chaos, and improves system resource stability. This application requires no manual intervention and can perform cleanup periodically or event-triggered, significantly reducing operation and maintenance costs and improving management efficiency. Attached Figure Description
[0043] The above and other features, advantages, and aspects of the embodiments of this application will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0044] Figure 1 This is a flowchart illustrating the sandbox resource processing method proposed in Embodiment 1 of this application;
[0045] Figure 2 This is a flowchart illustrating the sandbox resource processing method proposed in Embodiment 2 of this application;
[0046] Figure 3 This is a flowchart illustrating the sandbox resource processing method proposed in Embodiment 3 of this application;
[0047] Figure 4 This is a flowchart illustrating the sandbox resource processing method proposed in Embodiment 4 of this application;
[0048] Figure 5 This is a flowchart illustrating the sandbox resource processing method proposed in Embodiment 5 of this application;
[0049] Figure 6 This is a schematic diagram of the structure of a sandbox resource processing device provided in an embodiment of this application;
[0050] Figure 7 This is a schematic diagram of the hardware structure of a sandbox resource processing system proposed in an embodiment of this application. Detailed Implementation
[0051] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is only for explaining specific embodiments and is not intended to limit the application. The embodiments of this application are described below with reference to the accompanying drawings. It will be understood by those skilled in the art that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0052] The terms “first,” “second,” etc., used throughout this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0053] It is understood that before using the technical solutions disclosed in the embodiments of this application, relevant rights holders should be notified and their authorization obtained in accordance with relevant laws and regulations through appropriate means. This application does not restrict the prompting information and authorization implementation methods. Furthermore, the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0054] Furthermore, the resource orchestration platform involved in this application can be a container orchestration platform such as Kubernetes (which can be abbreviated as K8s) or Docker Swarm (Docker's native container cluster orchestration tool), i.e., a system that manages the lifecycle of containerized applications and interacts through encapsulated specific operating system APIs (Application Programming Interfaces). The centralized state store can be a high-performance, centralized data storage system such as Redis (an open-source in-memory key-value database) or etcd (Extended Key-Value Store, a distributed reliable key-value repository), used to manage the state, allocation records, heartbeat information, and sandbox network location information for external systems to query, serving as the authoritative source of facts for the entire system. The sandbox pool is a pre-created collection of sandbox instances in a ready state. When an AI (Artificial Intelligence) agent needs to isolate its execution environment, it can directly obtain sandbox instances from the sandbox pool, thus avoiding the delay of immediate creation. Sandbox instances can be containerized, isolated execution environments that provide AI agents with the tools they need to perform tasks, such as file systems, shells, and browsers, while being isolated from the host system and other sandbox environments to ensure security.
[0055] Reference Figure 1 This is a flowchart illustrating the sandbox resource processing method proposed in Embodiment 1 of this application. This sandbox resource processing method is applicable to electronic devices with deployed sandbox management services, such as servers, terminal devices, or device nodes in a cloud platform. The management service runs on the processing device of the electronic device and can interact with a resource orchestration platform and a central state storage (this application uses Redis as an example) to achieve full lifecycle management of sandbox instances. Based on this, as... Figure 1 As shown, the sandbox resource processing method proposed in this embodiment may include, but is not limited to:
[0056] Step S11: Obtain the first instance information of all sandbox instances currently running on the resource orchestration platform.
[0057] In this embodiment, a sandbox instance refers to an isolated execution environment provided for the AI agent, which includes an operating system, file system, shell, browser, and various development tools and API services. The resource orchestration platform manages the lifecycle of these sandbox instances, such as creation, scheduling, running, and destruction. The processing device can obtain a list / manifestation of all sandbox instances currently running on the platform, i.e., the first instance information, by calling the API of the resource orchestration platform. Therefore, the first instance information records the sandbox instances that actually physically exist on the resource orchestration platform.
[0058] In one possible implementation, taking Kubernetes as an example of a resource orchestration platform, the processing device can periodically call the Kubernetes API to obtain a list / manifestation of all running instances in the current namespace, which serves as the first instance information. Therefore, the first instance information obtained in this application can include the sandbox instance's unique identifier, creation time, running status, etc. The identifier of each sandbox instance can be a Pod name or a UUID (Universally Unique Identifier), etc., without restriction.
[0059] Step S12: Obtain the second instance information of all legitimate sandbox instances recorded in the central state storage. Legitimate sandbox instances include sandbox instances in the allocated state and sandbox instances in the unallocated state.
[0060] As described above, the central state storage is a storage system independent of the resource orchestration platform, used to maintain the logical state information of sandbox instances. The processing device of the electronic device can obtain a list / manifestation of all sandbox instances recognized by the system as legitimate from the central state storage, i.e., the second instance information.
[0061] Legitimate sandbox instances refer to sandbox instances that are explicitly recorded in the central state storage. These can include at least: sandbox instances in an assigned state, i.e., sandbox instances bound to a specific task identifier (such as projectId), currently being used by a task executed by an AI agent. The association between the identifier and the task identifier / task executor (AI agent or user) is stored in the assignment record of the central state storage. Legitimate sandbox instances also include sandbox instances in an unassigned state, i.e., sandbox instances that have completed initialization, are in a ready state, and have not yet been assigned to any task. These sandbox instances have been warmed up, and their identifier information is usually stored in a sandbox pool for unified management. This application merges the records of these two types of sandbox instances to form the second instance information. Therefore, the second instance information records the logically existing sandbox instances in the central state storage.
[0062] It is understandable that as the tasks performed by the AI agent change, such as adding new tasks or completing existing tasks, the running status of the sandbox instances recorded in the first and second instance information will dynamically change. For example, if a sandbox instance changes from a running state to an unassigned state after being released or cleaned up, it will be deleted from the first instance information. After an unassigned sandbox instance is assigned to a task, it will automatically update to an assigned state, and its status information in the second instance information will change accordingly. Therefore, the sandbox instances and their status information contained in the first and second instance information of this application can be dynamically adjusted. The implementation process will not be detailed in this application.
[0063] Step S13: Compare the first instance information with the second instance information to obtain the unrecorded sandbox instances in the first instance information that are not recorded in the second instance information.
[0064] Step S14: Perform a cleanup operation on unrecorded sandbox instances.
[0065] Following the above analysis, this application identifies sandbox instances that exist in the physical layer (i.e., platform layer) of the resource orchestration platform but not in the logical layer (i.e., application layer) of the central state storage by cross-layer state comparison. That is, the first instance information is cross-compared with the second instance information to find those sandbox instances that are actually running on the resource orchestration platform but have no valid record in the central state storage (neither in the sandbox pool nor allocated). These sandbox instances are usually caused by exceptions during creation, failures during allocation, or cleanup exceptions, and can be called orphaned sandbox instances, that is, sandbox instances that are not recorded in the second instance information in the first instance information (they are called unrecorded sandbox instances).
[0066] Since the sandbox itself is a stateless execution unit and does not undertake self-deregistration tasks, the cleanup operations (i.e., deletion operations) performed on all unrecorded sandbox instances (orphan sandbox instances) identified in this application are unidirectionally initiated and driven by the sandbox management service (SandboxPoolService) running on the processing device. Optionally, after determining each sandbox instance that needs to be cleaned up, the corresponding cascading transaction chain is actively executed, that is, the resource orchestration platform API is called to initiate the corresponding physical destruction request, such as a request to delete physical resources such as Pods / Deployments. After the API call successfully responds, a synchronous deletion instruction is actively sent to the central state storage to delete the physical resources corresponding to these unrecorded sandbox instances and clean up any related data that may remain. For example, HDEL sandbox:assigned <projectid>(Remove allocation table), HDEL sandbox:heartbeats <projectid>(Remove heartbeat monitoring), HDEL sandbox_locations <deploymentid>(Removing network location routes), etc., this application does not restrict the implementation method of the cleanup operation of sandbox instances.
[0067] In some embodiments, this application may periodically or according to preset cleanup rules actively execute steps S11-S14 of this embodiment. For example, when a periodic task is started, the function sandboxService.findAllActualSandboxNames is called to obtain a list of all actually running sandbox deployments (first instance information) under the current namespace from the container orchestration platform, and to obtain the IDs of all sandbox instances that have been assigned (such as HKEYs sandbox:assigned) and the IDs of all sandbox instances in the sandbox pool (in the assigned state) from Redis, such as traversing LRANGEsandbox:pool and parsing JSON (JavaScript Object Notation) data, converting these identification information into sandbox deployment names, and constructing a set containing all legal sandbox instance names, i.e., second instance information.
[0068] Afterwards, you can iterate through each actual running sandbox instance in the first instance information and check whether its name is recorded in the second instance information. If the name of an actual running sandbox instance is not in the list of valid sandbox instances, the sandbox instance is determined to be an orphan sandbox instance. Then, you can directly call sandboxService.cleanupSandbox to perform the cleanup operation, or you can perform the cleanup operation again after iterating through all the orphan sandbox instances in the current stage, etc. There are no restrictions on this.
[0069] In summary, in AI sandbox scenarios, this application, through cross-layer state comparison, without relying on the sandbox itself to report, allows the processing device to proactively identify physically existing sandbox instances that lack logically valid records. This solves the problem of unclaimed sandbox instances being difficult to discover due to creation failures, process crashes, or network partitions. Subsequently, this application performs cleanup operations on the identified unrecorded sandbox instances, promptly releasing occupied physical resources, preventing continuous resource leakage, and improving system resource utilization. Furthermore, by automatically reclaiming abnormal sandbox instances, this application prevents residual sandboxes from interfering with normal task execution, reducing the risk of system crashes due to resource exhaustion or state chaos, and improving system resource stability. In addition, this application requires no manual intervention and can perform cleanup periodically or event-triggered, significantly reducing operational costs and improving management efficiency.
[0070] In practical applications, for example, a background creation thread might just start a Pod in Kubernetes, while an orphan cleanup thread starts at the same time. However, the sandbox instance hasn't yet entered the sandbox pool. If it hasn't been pushed into Redis, the sandbox instance will be classified as an orphan sandbox instance according to the method described in Embodiment 1. Therefore, to prevent the risk of sandbox instances in the creation process being mistakenly classified as unrecorded sandbox instances and cleaned up, this application can also propose a protection strategy for sandbox instances in the creation process. In this regard, the central state store also maintains third instance information to record sandbox instances in the creation process. This third instance information can be a record in the sandbox creation process of a Redis Set data structure, such as `sandbox:creating_set`, used to store the identification information of all sandbox instances in the creation process.
[0071] Based on this, when the processing device executes the orphan sandbox instance cleanup process of Embodiment 1, the second instance information it obtains includes not only sandbox instances in the allocated state and sandbox instances in the unallocated state, but also sandbox instances being created as recorded in the third instance information. In other words, sandbox instances being created (sandbox instances in the initialization process) are logically also considered legitimate sandbox instances, and the third instance information is part of the second instance information. Thus, when comparing the first instance information with the second instance information, sandbox instances being created will not be identified as unrecorded sandbox instances (orphan sandbox instances), thereby avoiding their cleanup.
[0072] Based on the above analysis, in this application, the sandbox instance's identification information can be recorded in the third instance information of the central state storage at the start of sandbox instance creation. For example, when the processing device starts creating a new sandbox instance and submits a sandbox instance creation request to the resource orchestration platform, it can execute the Redis command `SADD sandbox:creating`.<pod_name> The identifier information (such as the name) of the sandbox instance currently being created is added to the third instance information as a creation record. This record indicates that the sandbox instance is in the process of being created and is not yet ready.
[0073] Optionally, at the instant the createSingleSandboxAsync task is started, the management service running the processing device, in addition to increasing the creatingCounterKey counter (which records the count information of all created instances), will also register the identification information of the currently created sandbox instance (such as the deploymentId unique placeholder) in Redis to a temporary creation status / record set (i.e., third instance information), such as the Redis Set structure: sandbox:creating_set, and merge the identification information therein into the second instance information to participate in the comparison and identification of orphan sandbox instances, ensuring that it will not be identified as an orphan sandbox instance in the current state (creating).
[0074] Once a sandbox instance is created and enters an unassigned state (i.e., added to the sandbox pool and available for subsequent allocation), the processing device can perform a state transition operation to remove the sandbox instance's identifier information from the third-party instance information. This is done, for example, by executing the Redis command `SREM sandbox:creating`.<pod_name> This ensures that a sandbox instance being initialized is logically considered a valid sandbox instance. When the running state of this sandbox instance changes, i.e., from creation / initialization to ready (unallocated state), it is added to the unallocated instance information in the second instance information, for example, by executing the Redis command LPUSH sandbox:pool.<pod_name> This allows the newly created sandbox instance to be allocated and used by subsequent sandbox acquisition requests.
[0075] In another possible implementation, the above record and remove operations can be achieved through atomic Redis commands, such as executing `SADD sandbox:creating_set` at the start of creation. <deploymentid>After creation is complete and the sandbox is entered into the pool, execute SREM sandbox:creating_set <deploymentid>This ensures that the third-party instance information remains consistent with the actual physical state. The entire creation process is encapsulated in an exception handling block. Even if an exception occurs during creation, the processing unit will perform a cleanup operation in the finally block. This involves removing the identifier information of the currently created sandbox instance from the third-party instance information and calling the resource orchestration platform API to destroy the corresponding physical instance, preventing resource leaks.
[0076] Therefore, throughout the entire creation process of a sandbox instance, its identification information is always present in the third instance information. The sandbox instance in the third instance information is considered part of the second instance information, and thus will not be mistakenly identified as an unrecorded sandbox instance in the orphan sandbox instance cleanup process of Embodiment 1. Only after creation is completed and successfully added to the sandbox pool record (the instance information in the unallocated state in the second instance information) does the sandbox instance completely detach from the protection of the third instance information, and instead continue to be protected by its unallocated state identity in the sandbox pool record, still preventing it from being mistakenly identified as an unrecorded sandbox instance. Therefore, this embodiment of the application, by introducing third instance information, achieves precise protection for sandbox instances in the process of creation, thereby avoiding false positives caused by timing competition between orphan cleanup and sandbox creation, and improving the stability and reliability of the system.
[0077] To further enhance the security of orphan sandbox instance cleanup, such as to prevent the protection of sandbox instances under creation from failing due to extreme situations like central state storage write delays, this application also proposes a hard check of the sandbox instance's lifespan. That is, before identifying sandbox instances that are not recorded in the second instance information from the first instance information and determining them as orphan sandbox instances, or after identifying orphan sandbox instances but before performing cleanup operations on them, the resource orchestration platform API can be called to query the sandbox instance's creation timestamp (metadata.creationTimestamp), i.e., obtain the creation time information of sandbox instances that are not recorded, so as to determine whether the sandbox instance is exempt from cleanup based on the creation time.
[0078] Therefore, based on the creation time information of unrecorded sandbox instances, the creation duration of such instances is determined. For example, the difference between the current time and the creation time can be considered the sandbox instance's lifespan. A pre-configured safety grace period (i.e., a preset duration threshold) is set for unrecorded sandbox instances not to be cleaned up. If the lifespan of a sandbox instance is less than this grace period, the system can exempt it from orphan sandbox status, determining it to be in the normal cold start / image pulling phase and not performing cleanup. In other words, if an unrecorded sandbox instance exists with a creation duration less than the preset duration threshold, no cleanup operation is performed on it. Only containers with a creation duration greater than or equal to the preset duration threshold and no legitimate identity in the central state store are defined as orphan sandbox instances, and cleanup operations are then performed on them.
[0079] The preset time threshold can be determined comprehensively based on actual factors such as sandbox image size, network conditions, and internal service initialization time. It is usually set as the maximum cumulative time limit for image pulling and service initialization, such as 3 to 10 minutes. It may also be set to more or longer, without any restrictions.
[0080] In summary, the liveness hard verification method in this embodiment, as a supplement to the protection of third instance information, further reduces the risk of sandbox instances being mistakenly killed during cold start due to extreme timing issues (such as delays in updating third instance information), thus forming a two-layer security protection.
[0081] In some embodiments, in addition to the orphaned sandboxes cleanup method described in Embodiment 1 above, the automated cleanup method for the sandbox lifecycle can also monitor the business activity of allocated sandbox instances and reclaim resources abandoned after normal use, i.e., cleanup inactive sandboxes. In this way, through these two complementary and independently inspected automated cleanup mechanisms, the system is endowed with strong self-healing capabilities.
[0082] Based on this, refer to Figure 2 The flowchart illustrating the sandbox resource processing method proposed in Embodiment 2 of this application describes a possible implementation method for cleaning up inactive sandboxes. It reclaims allocated but long-term inactive sandbox instances. This method can be executed independently and in parallel with the orphan sandbox instance cleanup process described in Embodiment 1, thereby proactively identifying and repairing resource leaks and state inconsistencies caused by normal or abnormal conditions. Figure 2 As shown, the implementation method may include:
[0083] Step S21: In response to the business request initiated by the agent to the sandbox instance, update the active time information of the corresponding sandbox instance in the central state storage according to the task identifier corresponding to the business request.
[0084] Step S22: Based on the active time information, determine the sandbox instances that have entered the inactive state from among the sandbox instances that were in the assigned state.
[0085] Step S23: Perform a cleanup operation on sandbox instances that have entered an inactive state.
[0086] During task execution, the AI agent initiates business requests to the sandbox instance, such as executing shell commands or reading / writing files. Upon intercepting these requests, the processing device parses them to obtain the task identifier, such as projectId, and updates the active time information of the sandbox instance associated with that task identifier in the central state storage. This update process can be implemented based on a heartbeat mechanism. The heartbeat is proactively updated by the management service (McpContextInterceptor / SystemToolService) running on the processing device, eliminating the need for the AI agent to send meaningless Ping packets or for the management service to periodically poll and probe the sandbox. The cleanup process in this embodiment can be triggered based on interactive requests.
[0087] In one possible implementation, the processing device updates the current timestamp to the heartbeat record in the central state storage in an event-driven manner each time it responds to a business request. For example, whenever the AI agent initiates a business request to the sandbox, after the interceptor parses the task identifier (projectId), it can execute HSET sandbox:heartbeats once in the central state storage. <projectid> <timestamp>This updates the active time information of the corresponding sandbox instance in the central state storage, facilitating the subsequent use of the time difference between the current system time and the last heartbeat time to determine whether the corresponding sandbox instance has entered an inactive state. It is evident that the active time information update process seamlessly embeds the heartbeat refresh into the normal business flow, requiring no additional network overhead and avoiding wasted bandwidth resources.
[0088] The inactive sandbox cleanup process based on the heartbeat mechanism proposed in this embodiment can be implemented periodically. The update cycle can be determined based on the actual business interaction frequency of the AI agent. For example, the more frequently the agent calls the sandbox to perform tasks, the more frequently the heartbeat updates, that is, the shorter the update cycle. If the agent is in a thinking state and does not call the sandbox, the heartbeat update is not required.
[0089] Based on this, this application can retrieve all sandbox records (sandbox:assigned) and all heartbeat records (such as sandbox:heartbeats, which can be stored using a hash data structure, with the task identifier in the hash table as the key and the current system timestamp as the value, etc., but not limited to this) from the central state storage HGETALL when a periodic task starts. It then iterates through each sandbox instance in the assigned state, using its associated task identifier to find the last heartbeat time. If the time difference between the current system time and the last heartbeat time is greater than the inactivity threshold (which can be represented as inactiveThresholdMs, and can range from 5 to 15 minutes, but is not limited to this), the sandbox instance can be determined to be inactive, and a cleanup operation can be performed on it. For example, calling sandboxService.cleanupSandboxByProject to find the deploymentId of the sandbox instance and calling the underlying API to delete the sandbox instance. This application does not restrict the implementation method of the cleanup operation.
[0090] Understandably, while deleting inactive and orphan sandbox instances from sandbox records and heartbeat records, it is also possible to delete metadata such as network location records from the local deployment of sandbox instances (such as sandbox_locations) to ensure that external systems cannot query the cleaned-up sandbox instances.
[0091] Optionally, the aforementioned inactivity threshold can be determined based on factors such as the maximum single-step planning and thinking time limits of the large model (such as a large language model) upon which the AI agent relies. After the AI agent executes a task instruction in the invoked sandbox instance, it can send the execution result back to the large model. Before initiating the next sandbox call, the large model may experience delays of tens of seconds to several minutes due to complex reflection, multi-step reasoning, and prompt word concatenation (i.e., the CoT chain). To prevent the sandbox instance from being mistakenly judged as inactive and cleaned up by the management service during the large model's thinking period, the aforementioned inactivity threshold needs to be significantly greater than the maximum single-step reasoning time limit of the large model. For example, if the maximum time limit of the large language model is within 2 minutes, the inactivity threshold can be set at 5-10 minutes as an optimal safety buffer, but it is not limited to this.
[0092] As can be seen, this application, in response to business requests or preset heartbeat cycles, traverses the sandbox instances of business requests or the active time information of each existing sandbox instance in the central state storage, thereby timely and accurately identifying sandbox instances that have entered an inactive state and performing cleanup, thus promptly discovering and reclaiming sandbox instances that have been allocated but have not been used for a long time, avoiding resource idleness and waste.
[0093] Moreover, the heartbeat inactivation sandbox instance cleanup method in this embodiment complements the orphan sandbox instance cleanup method described in Embodiment 1. The former solves the waste of resources that exist logically but are not used for a long time, while the latter solves the abnormal residues that exist physically but do not exist logically. These two cleanup methods are not simply functional superpositions, but rather a whole that is functionally orthogonal and complementary, designed for different types of anomalies (normal but inactive, abnormal residues) that may occur at different life stages of the sandbox. Together, they ensure that the system has a strong self-healing capability. That is, through automated, multi-dimensional inspection and cleanup capabilities, it can proactively discover and repair resource leaks and state inconsistencies caused by normal or abnormal situations, ensuring the long-term stability of the system and a resource utilization rate of over 99%. It achieves a 100% recovery rate of transient resources, ensuring that idle resources caused by any reason can be automatically discovered and completely recovered by the system, greatly improving the robustness of the system and ensuring long-term operational stability and resource efficiency.
[0094] When AI agents replace or assist humans in completing complex digital world tasks, a sandbox environment similar to a human work environment is needed. This sandbox requires not only an operating system and command line (shell), but also tools such as a browser, file system, and code editor. This isolated environment can be built using container technologies (such as Docker) and container orchestration technologies (such as Kubernetes). Currently, however, a common approach is to use on-demand, real-time containerized sandbox systems. When a central service receives a request from an AI agent or user to create a sandbox environment, it can call the Kubernetes API, submitting a request to create a new Deployment or Pod. This request specifies a container image containing all the tools. The Kubernetes scheduler allocates resources on a worker node, and the Kubelet on the node begins pulling the container image. After the container starts, the internal startup scripts (such as supervisord) sequentially start Xvfb (virtual desktop), a VNC (Virtual Network Computing) server, a Chrome browser, and various Python API services for the agent to use. Once all services within the container have started and passed health checks, the container will notify the central service of its readiness through some means (such as a callback API) and report its assigned IP address. The central service then returns the sandbox information (IP address, access credentials, etc.) to the AI agent, allowing the agent to begin using the corresponding sandbox instance.
[0095] As can be seen, this passive response mode, which only begins creating the sandbox environment upon request from the agent, involves starting a virtual machine or container, pulling a large image containing a complex toolchain (such as a browser, desktop environment, and development libraries), and starting multiple internal services and waiting for them to become ready. This entire process can take several minutes, significantly impairing the immediacy and smoothness of user interaction with the agent. Even if a large number of environments could be created in advance, it would still result in substantial resource idleness and cost waste, which is unacceptable for AI agent applications requiring real-time interaction (such as transient AI tasks).
[0096] To address this, this application proposes a sandbox pooling pre-warmed mechanism to expedite the lengthy cold start process. Pre-warming means that before a sandbox instance is allocated to an agent, it has completed all time-consuming initialization processes, including container startup, image pulling, startup and health checks of internal services (such as VNC, browser, and API server), placing it in a ready-to-use state. Thus, users no longer request sandbox creation, but rather retrieve a sandbox from a pre-ready pool, reducing latency from minutes to milliseconds, fundamentally solving the latency problem and providing AI agents with a seamless, ready-to-use experience.
[0097] Based on this, refer to Figure 3 The flowchart shown in Embodiment 3 of this application illustrates the sandbox resource processing method, which can maintain the number of available sandboxes in the pool by actively creating sandbox instances in the background. Figure 3 As shown, the sandbox resource processing method proposed in this embodiment may include:
[0098] Step S31: Pre-create and maintain at least one sandbox instance in an unassigned state in the central state storage.
[0099] In this embodiment, during system initialization or operation, the processing device pre-creates a batch of sandbox instances on the resource orchestration platform. These sandbox instances have completed all time-consuming processes such as image pulling, container startup, internal service initialization, and health checks, and are in a readily available, unallocated state, forming a sandbox pool. This pool is stored in the sandbox pool record in the central state storage, such as in the Redis sandbox:pool list, i.e., the second instance information. When a user or agent requests a sandbox, it is directly allocated from the sandbox pool instead of being created from scratch, thus achieving millisecond-level instant provisioning.
[0100] Optionally, this application can acquire an atomic lock with a timeout (such as SET...NX EX...) in Redis to ensure that, in a distributed deployment environment, only one service instance executes the sandbox pool maintenance logic at any given time, thus avoiding duplicate creation.
[0101] Step S32: Obtain the number of sandbox instances in the unallocated state in the central state storage.
[0102] Step S33: If the quantity is lower than the preset water level, create a new sandbox instance and record the identification information of the sandbox instance in the third instance information.
[0103] Step S34: After the sandbox instance is created and enters the unallocated state, remove its identification information from the third instance information and add it to the second instance information for instances in the unallocated state.
[0104] After successfully acquiring the distributed lock, to determine whether new sandbox instances need to be created and the required number of sandbox instances, three key metrics can be read from the central state store: the number of available sandbox instances in the current sandbox pool (LLEN poolKey, i.e., currentSize), the number of sandbox instances being created (creatingCounter), and the configured target sandbox pool size (poolSize). This determines the number of sandbox instances that need to be added, calculated as needed = poolSize - (currentSize + creatingCount). If needed > 0, it indicates that the number of unallocated sandbox instances in the sandbox pool is currently low, and new sandbox instances need to be added immediately. Conversely, if the number is low, new sandbox instances can be temporarily suspended, and the number in the sandbox pool can continue to be monitored.
[0105] Therefore, this application can periodically detect whether the number of available sandbox instances in the sandbox pool is lower than a preset water level. If it is lower than the preset water level, it indicates that the number of sandbox instances in the current sandbox pool may not meet the actual needs, and the process of creating additional sandbox instances is initiated. The preset water level can be determined based on the target sandbox pool size and the number of sandbox instances currently being created, such as the difference between the target sandbox pool size and the number of sandbox instances currently being created, i.e., poolSize-creatingCount. If (poolSize-creatingCount) > currentSize, a new sandbox instance is created and added to the sandbox pool. This application does not limit the size of the preset water level used to determine whether the current sandbox pool needs to be replenished with new sandbox instances.
[0106] In the case where `needed > 0`, a corresponding number of asynchronous creation tasks can be submitted to the scheduling thread pool (ScheduledExecutorService). Delays can be set between tasks to stagger creation requests and avoid instantaneous impact on the resource orchestration platform's API. Each asynchronous task executes the creation of a single sandbox instance. The creation process can include: atomically incrementing the `creatingCounterKey` in Redis (INCR creatingCounterKey), which is the count information of the sandbox instance being created. This is equivalent to incrementing the corresponding creation record in the central state storage, recording the sandbox's identification information in the third instance information. Afterward, the resource orchestration platform API can be called to create the instance, such as calling the `SandboxService.createPoolableSandbox` method to interact with the container orchestration platform (DomeOS) by calling the underlying `SandboxManager`, creating a sandbox instance and configuring it with unique identification information.
[0107] The identification information can be generated by embedding a placeholder ID into the resource name (Pod / Deployment) of the container orchestration platform. That is, during creation, a globally unique placeholder UUID is generated locally, for example: 123e4567-e89b-12d3-a456-426614174000. To ensure uniqueness within the Kubernetes namespace, this UUID can be embedded into the physical resource name, generating a Pod name like sandbox-pod-123e4567, which serves as the identification information for this sandbox instance. It should be noted that this identification information remains unchanged throughout the entire lifecycle of the sandbox instance. When the AI agent requests the sandbox instance, this identification information is associated with the actual business tenant information of the current request (such as task identifier or task executor identifier) and written as a key-value pair into the Redis mapping table, i.e., the sandbox record. This achieves anonymous placeholders at the physical layer and dynamic real-name binding at the logical layer, enabling millisecond-level rapid delivery and meeting real-time response requirements.
[0108] After creation, you can wait for the services inside the container (such as Xvfb, noVNC, Python API Server, etc.) to start and pass the health check. At this point, createPoolableSandbox returns a CompleteableFuture. If creation is successful when this Future completes, it means the health check has passed and the instance has entered the ready state. You can then serialize the complete information of the created sandbox instance (including network location, access credentials, etc.) into JSON data and store it in the sandbox pool record. At this point, you can add its identification information to the unassigned instance information in the second instance information. If creation fails, such as if the creation wait time exceeds a certain waiting threshold, an error log can be logged. Regardless of whether creation succeeds or fails, in the creation completion block of the task, the creation counter (creatingCounterKey) must be atomically decremented, and the identification information must be removed from the creation collection (i.e., its identification information is deleted from the third instance information) to ensure the accuracy of the creation counter information.
[0109] As can be seen, by using the sandbox pool management service, a preheated sandbox pool is proactively created and maintained on the container orchestration platform. The sandbox instances within it have completed all time-consuming startup processes and are immediately available. When an AI agent makes a request, the system can directly allocate a sandbox instance from the sandbox pool instead of creating one, achieving near-zero latency provisioning. Specifically, on the Kubernetes / DomeOS container cloud platform, each preheated sandbox instance physically corresponds to an independent Pod or a Pod controlled by a single-replica Deployment. This Pod encapsulates the operating system image, basic toolchain, and related API services required for the agent to execute tasks.
[0110] In some embodiments, this application can configure a timeout threshold for the above-mentioned creation waiting process, such as 60 to 120 seconds, which can be determined based on the maximum cumulative time limit of container cold start and internal service initialization. Optionally, this application can dynamically configure the timeout by evaluating parameters such as image pull latency, container startup and internal process auto-start time, and health check retry buffer, so that if the creation waiting time exceeds the timeout threshold and no HTTP200 is returned, it indicates that the creation has failed, and the occupied physical resources are immediately terminated and released to avoid waiting indefinitely for the asynchronous creation task.
[0111] The image pull latency is as follows: a sandbox image integrating tools such as a desktop and browser is typically between 2GB and 5GB in size. Without local caching on the node, pulling this image across the network usually takes approximately 30-90 seconds. Container startup and internal process auto-start time: after the container starts running, the internal supervisord sequentially starts and initializes services such as Xvfb, noVNC, and the Python API Server, which takes 5-15 seconds. Health check retry buffer: the management end polls the sandbox's readiness status via an HTTP interface, typically requiring a 5-10 second buffer for multiple retries. This embodiment evaluates the total latency of these parameters and sets a waiting threshold to cover most network fluctuations and system loads; its size is not limited.
[0112] Therefore, this embodiment can maintain a certain number of ready sandboxes in the sandbox pool, i.e., sandbox instances in an unallocated state, through the above method. This transforms the process of an agent acquiring a sandbox from time-consuming creation to millisecond-level allocation, thereby fundamentally eliminating startup delay.
[0113] Reference Figure 4 This is a flowchart illustrating the sandbox resource processing method proposed in Embodiment 4 of this application. This embodiment describes the sandbox instance allocation process, that is, a possible implementation method for an agent to acquire sandbox instances, such as... Figure 4 As shown, the implementation method may include:
[0114] Step S41: In response to the sandbox retrieval request, retrieve a sandbox instance in an unallocated state from the sandbox pool.
[0115] When an AI agent or user initiates a sandbox retrieval request, the processing unit atomically pops a sandbox instance from the pool records in the central state store and returns it to the caller. For example, executing the Redis RPOP command attempts to atomically pop JSON data of a sandbox instance from the Redis sandbox pool list, carrying identification information at creation time (such as placeholder IDs). At this point, the sandbox instance is in an unallocated state, has completed all initialization work, and is immediately available.
[0116] It should be noted that if there are no sandbox instances in the sandbox pool, that is, no sandbox instances in the central state storage are in an unallocated state, it can be downgraded to the on-demand creation mode. The sandboxService.createAndAssignSandbox is directly called to create a new sandbox instance for the current user. That is, a corresponding sandbox instance is created for the task that initiates the sandbox acquisition request to meet the actual business needs. The creation process can be referred to the description of the corresponding part of the above embodiment, and will not be repeated in this embodiment.
[0117] In some embodiments, when handling high-concurrency requests in an empty sandbox pool state, a classification and degradation strategy based on tenant-level identifiers can be adopted. This is especially relevant when different concurrent requests belong to different projects (different projectIds) or different task executors. Due to the multi-tenant isolation requirements of the sandbox, each project must have its own independent and secure physical runtime environment. That is, a separate sandbox instance is created for each task identifier, and they cannot share the same sandbox. Therefore, when processing requests from multiple different projects, this application will concurrently create multiple independent sandbox instances, i.e., concurrently call the Kubernetes API to create separate physical Pods for different projects. The creation process is not detailed here.
[0118] Preferably, to prevent the physical cluster from crashing due to excessive instantaneous concurrent creation, the system can implement threshold-based rate limiting on the overall concurrent creation speed, limited by the creatingCounter (in-process creation counter) and the underlying thread pool. In other words, based on the number of sandboxes currently being created and the thread pool limit, the upper limit of concurrent creation is determined, and the number of multiple concurrent requests is then controlled accordingly.
[0119] Furthermore, for multiple concurrent requests from the same task executor, idempotent (anti-duplicate) control can be implemented using a distributed lock based on task identifiers. This prevents the repeated creation of multiple useless and expensive sandbox instances for the same task identifier. Specifically, before entering the creation logic, the processing device attempts to acquire a distributed lock based on the task identifier, such as SETsandbox:lock: <projectid>In NX EX 10, for requests that successfully acquire the lock, the creation and allocation of sandbox instances can be performed, as described in the previous embodiment, and the lock is released after successful verification. For requests that fail to acquire the lock, a short-term polling wait can be performed on the lock. When the lock is released, the relevant information of the sandbox instance already created by the successful lock acquirer is directly read from the central state store and returned directly, achieving request merging and allocation idempotency.
[0120] Step S42: Update the status of the acquired sandbox instance to the assigned status, and record the association information between the sandbox instance and the task identifier.
[0121] The processing device binds the popped-up sandbox instance with the task identifier (such as projectId) in the request. For example, in the allocation record of the central state storage (such as the second instance information like the sandbox:assigned hash table), a binding record containing information such as the sandbox instance's identifier, allocation time, and creation time is stored using the task identifier as the key. This application does not restrict the content of the binding information or its recording method. Each binding record can be written into the allocation record of the central state storage, such as into the second instance information, for later reference.
[0122] Simultaneously, the current timestamp is stored in the heartbeat record as the starting point of the active time information, i.e., the initial heartbeat value. This is stored in association with the task identifier to achieve timely cleanup of allocated but inactive sandbox instances according to the method described in the above embodiment. At this time, the sandbox instance changes from an unallocated state to an allocated state, and can be updated in the instance information section of the second instance information corresponding to different states, or the state information of the corresponding sandbox instance in the second instance information can be directly updated.
[0123] Step S43: Obtain the network location information of the sandbox instance and write it into the location record of the central state storage.
[0124] When allocating sandbox instances, their network location information can also be read directly from the central state storage. This information may include, but is not limited to, internal IP, VNC port, API port, etc., such as 10.1.2.3:6080. When the sandbox instance is created, this application uses the task identifier as the key and the network location information as the value to solidify the dynamically changing network location information in the central state storage. That is, the network location information corresponding to each sandbox instance is written into the location record of the central state storage. This can be a dedicated sandbox_locations hash table for external systems such as unified gateway routing services or monitoring systems to query and access. For example, it provides an O(1) fast query plane for external systems without parsing the complex complete JSON data of the sandbox instance.
[0125] As can be seen, the acquisition of network location information for each sandbox instance does not rely on real-time calls to the resource orchestration platform's API. Instead, it is determined during the preheating (sandbox pre-creation) phase. During allocation, the network location information is directly extracted from the deserialized sandbox object. This avoids the impact on the resource orchestration platform's API under high-concurrency allocation scenarios, which could lead to its crash. At the same time, it reduces allocation latency, avoiding the 100ms-500ms increase in latency caused by real-time API queries for each allocation, thus slowing down immediate response.
[0126] In practical applications, configuration information such as identification information and network location information (location record) for the same sandbox instance can be configured by building and maintaining a mapping table (i.e., a mapping relationship, but not limited to the mapping table storage method; this application only uses this as an example for illustration) in the central state storage. This mapping table can be queried and used by external systems to achieve comprehensive observability and accessibility of the sandbox instance's state.
[0127] Furthermore, not just any external system has the right to query and access this mapping table. Only authorized and trusted internal infrastructure components, such as unified gateway routing services, system monitoring dashboards, and log audit aggregators, which reside in the same trusted domain within the central state store, can query this mapping table. To address this, the central state store itself can enable ACLs (Access Control Lists) and password authentication to prevent ordinary third-party users or unauthorized systems from directly querying the mapping table.
[0128] Step S44: If the association information recording fails or the location record writing fails, the identification information of the corresponding sandbox instance will be re-added to the second instance information.
[0129] To ensure transactional consistency in allocation operations and prevent sandbox instances from becoming unallocated and neither in the sandbox pool nor successfully allocated due to intermediate failures after being popped by RPOP, this application introduces a local try-catch interception and Redis transaction rollback handling method. This encapsulates the entire allocation process—popping from the sandbox pool, writing the allocation record, and writing the position record—within an exception handling block (such as a Java try-catch block). If an exception occurs during the writing of the allocation or position record, such as intermittent Redis outages or network jitter, the exception handling block can promptly capture the exception and immediately perform compensation operations. For example, it can re-add the newly popped sandbox instance to the sandbox pool record (second instance information), or re-serialize the complete information of the newly popped sandbox instance and push it back to the central state storage sandbox pool by executing the Redis RPUSH command, and re-add its identification information back to the second instance information, ensuring that the sandbox instance can be reused by the next normal request. At the same time, the business side of the current request can return a failure response to guide the caller (such as an AI agent) to back off and retry, ensuring that it will not get a sandbox instance in a semi-bound and corrupted state.
[0130] Step S45: If it fails to re-add the sandbox instance's identifier information to the second instance information, call the application programming interface of the resource orchestration platform to destroy the sandbox instance.
[0131] As described above, if the compensatory pushback operation also fails (e.g., Redis completely disconnects), the processing device can execute a fallback mechanism after catching the RPUSH exception. This involves bypassing the central state store and directly calling the resource orchestration platform's API to asynchronously destroy the physical Pod of the sandbox instance, preventing physical resource escape. After Redis recovers, the orphan sandbox instance cleanup process described above can be relied upon to achieve eventual consistency convergence between the physical and logical layers. The implementation process is not detailed in this application.
[0132] In summary, the state of the sandbox pool, its allocation status, heartbeat, routing information, and distributed locks used for cleanup are all centralized in a high-performance central state store. This does not depend on the sandbox instance's own registration or heartbeat reporting. Instead, an omniscient management service running on the processing device actively maintains and drives all state changes, forming a complete closed loop from provisioning and monitoring to recycling. This architecture ensures data consistency and management simplicity, solving the problems of high cost and complexity in traditional methods.
[0133] In practical applications, sandbox instances preheated according to the method described above will continuously maintain their internal services in a running state (hot-standby) while waiting for allocation, instead of entering a dormant or suspended state. This avoids the hundreds of milliseconds to several seconds of latency that would occur when an agent requests a transient AI task, due to waking up the container, restoring memory state, and rebinding network routes. This reduces the extremely low latency (e.g., millisecond-level response) of the preheated sandbox, meeting the immediacy requirements of transient AI tasks. Transient AI tasks are characterized by short lifecycles (typically seconds to minutes, such as 5 seconds to 10 minutes, like an AI agent temporarily executing code, fetching a specific webpage, performing single-step format conversion, or calling a single planning tool), frequent starts and stops (exhibiting high concurrency, high dynamism, and burstiness; in production environments, the start-stop concurrency for a single tenant or single system can reach tens to hundreds of times per second), volatile states, and the possibility of abnormal interruptions during execution.
[0134] In practical applications, the lifecycle and start / stop frequency of AI tasks can be dynamically determined by the management end (AI agent application layer) that calls the sandbox based on business logic at the start and end of the task. These are not preset static values. The end of the lifecycle can be dynamically determined through a heartbeat detection mechanism. Combined with the relevant description of the identification method of inactive sandbox instances above, if no instruction is detected when the time difference between the creation time of the sandbox instance and the last heartbeat time is greater than the inactive threshold, it is determined that the lifecycle of the currently requested transient AI task has ended, and the cleanup process can be triggered.
[0135] It is evident that sandbox instances of AI agents are essentially one-time, disposable resources serving transient AI tasks. They should not, and cannot, assume the responsibility of reliably managing their own lifecycle. For example, a sandbox instance executing a code compilation task should be destroyed immediately after the task ends; it has no opportunity to deregister itself. A sandbox instance that crashes due to executing malicious code is even less likely to issue an offline notification. This application adopts a centralized push model, completely entrusting the responsibility of state management to a reliable, long-term management service. It downgrades the sandbox to a pure, stateless execution unit, not relying on this unreliable transient resource itself to report its state. From an overall architectural perspective, this ensures that even if the sandbox instance itself is unreliable, its metadata state in the system remains accurate and manageable, avoiding management chaos and resource leaks.
[0136] It should be understood that this application can also be applied to long-running AI tasks. For this, the corresponding configuration information needs to be adjusted accordingly, such as setting the inactivity threshold to infinity or marking long-running AI tasks as exempt from cleanup whitelisting. This is to avoid using the default logic applicable to managing transient AI tasks to manage long-running AI tasks. If a long-running AI task does not interact with the management API for a long time during local high-load computation, and the heartbeat time is not updated, the cleanupInactiveSandboxes service may mistakenly determine that it is dead and forcibly destroy the container, leading to a false kill risk. Furthermore, long-running AI tasks occupy sandbox pool resources that should be frequently used, causing the replenishment speed of idle sandbox instances to lag behind, thus undermining the high-frequency turnover design intention, resulting in resource exhaustion and decreased turnover rate. Additionally, long-running sandboxes are prone to memory leaks, disk overload, or dependency library pollution, which violates the design intention of this application to maintain an absolutely clean and secure isolated environment after use, leading to environmental degradation and state accumulation. Therefore, for long-term AI tasks, this application configures the routing to a non-pooled, exclusive sandbox lifecycle management link through the control plane, skipping the cleanup and detection of inactive sandbox instances, thereby achieving the classification and parallel management of long-term and transient AI tasks.
[0137] In some embodiments, pre-warmed (pre-created) sandbox instances in the sandbox pool are not maintained indefinitely. Dynamic lifecycle auditing and periodic rolling updates can be used to prevent environmental degradation, i.e., they are in a pre-warming waiting state. However, long-running containers may experience minor system drift, memory fragmentation, or connection hangs. In one possible implementation, such as in high-frequency scenarios, unallocated sandbox instances in the sandbox pool are quickly allocated and consumed by RPOP, while new sandbox instances are continuously replenished in the background. The sandbox instances in the sandbox pool have a very short dwell time, thereby achieving natural replacement.
[0138] In another possible implementation, the sandbox pool can be actively audited and timed out (re-warmed up). Based on the above analysis, the management background daemon thread will periodically (e.g., every 10 minutes) audit the time-to-live (TTL) of the pre-warmed sandbox instances in the Redis queue, which is the creation time mentioned above. Once it is found that the time-to-live of a sandbox instance in the sandbox pool exceeds a preset value (e.g., 2 hours), a cleanup operation can be performed on it, that is, it can be removed from the sandbox pool and safely destroyed. At the same time, checkAndReplenishPool is triggered to recreate and pre-warm a new sandbox instance, thereby ensuring that a certain number of sandbox instances available at any time in the sandbox pool are always kept in the latest, healthiest, and uncontaminated state.
[0139] Based on the above analysis, for inactive sandbox instances, user-released sandbox instances, and sandbox instances undergoing rolling updates and elimination, in scenarios where the container cloud platform API is called to initiate physical destruction requests and complete the corresponding cleanup operations, a physical layer and logical layer state reverse reconciliation awareness algorithm can be used to perceive the physical state of the sandbox instance and synchronize the Redis state. Optionally, the reconciliation thread can periodically pull the identification information of all currently valid sandbox instances from Redis to construct a logical fact set L, i.e., the second instance information, and actively pull the physical full Pod list through the Kubernetes API to construct a physical reality set P, i.e., the first instance information. Difference set reconciliation calculations are then performed on these P sets to determine sandbox instances that logically exist but have physically disappeared, i.e., the difference set Dlost = LP.
[0140] Once a non-empty Dlost set is detected, the management service recognizes that these sandbox instances have experienced abnormal physical demise and can automatically trigger a reverse repair process, forcibly removing the logical record of the corresponding sandbox instance from Redis's allocation records, heartbeat records, and routing mapping tables. Thus, this application, through proactive control plane comparison and reverse repair, ensures that even if Kubernetes automatically reschedules or the container passively deadlocks, the Redis ledger can detect and automatically synchronize its state within seconds. This proactive cleanup method and periodic difference set reconciliation method ensure consistency between the logical state of Redis and the physical state of the container platform.
[0141] In practical applications of distributed architectures, container orchestration platforms may experience temporary unavailability due to API overload, rate limiting, or network jitter. Based on the above analysis, system robustness can be ensured through strategies such as asynchronous fault tolerance capture, rate limiting backoff, and two-phase state protection. Specifically, when a background thread attempts to create a sandbox, if the Kubernetes API throws an exception (e.g., ConnectException, SocketTimeoutException, or HTTP 503 / 429 error), the exception handling block described above can atomically decrement the creation counter in Redis during the asynchronous creation task. This prevents logical deadlocks caused by the creation counter failing to decrement due to exceptions, thus preventing the system from mistakenly believing that sandboxes are constantly being created and ceasing to add new sandboxes. Since sandbox creation fails, the JSON data of the sandbox instance will not be pushed into the Redis queue. In the next timed period, such as the maintainPools task, the preset watermark can be rechecked and creation attempted again, achieving natural retry fault tolerance.
[0142] In summary, during the process of responding to sandbox acquisition requests and executing RPOP to attempt to acquire ready sandbox instances from the pool, even if Kubernetes is temporarily unavailable, as long as there are still existing sandbox instances in the Redis sandbox pool, the AI agent's requests can still achieve near 100% seamless delivery, realizing deep decoupling between physical layer anomalies and business layer delivery. If there are no sandbox instances in the sandbox pool, and on-demand creation fails due to Kubernetes downtime, a rate limiting / degradation error message can be returned to the business side, but the entire Redis state machine will not be disrupted. Furthermore, when cleaning up inactive sandbox instances, if the Kubernetes delete interface call fails, the system will not immediately clear the corresponding state data in Redis, such as the records in `sandbox:assigned`. Instead, it will retain the records in Redis, waiting for the next cleanup cycle to attempt to delete the inactive sandbox instances from Kubernetes again. Only when the Kubernetes interface successfully returns a destruction / cleanup confirmation, or when it is confirmed in subsequent reconciliation that the physical resources have been cleaned up, is the state data in Redis finally cleaned up, thus ensuring transactional consistency of state changes.
[0143] In some embodiments, the AI agent, upon receiving an assigned sandbox instance, can send Bash commands to that sandbox instance via the SandboxController to execute commands within that instance when performing tasks. Optionally, this application can set a maximum allowed time for command execution, such as setting an execution timeout in ExecuteReq (e.g., a default of 30 seconds), which can be determined based on the business-normal time limit for transient AI tasks and security escape prevention strategies. Based on this, if the AI agent generates malicious code or executes a command requiring interactive blocking, this execution timeout mechanism acts as a forced hard shutdown, forcibly terminating the command. This effectively prevents computing resources from being exhausted by malicious or uncontrolled code, avoids a single task dragging down the entire sandbox or even affecting the host system, ensures that no command can run indefinitely, and guarantees the stability and security of the entire sandbox system.
[0144] Reference Figure 5 This is a flowchart illustrating the sandbox resource processing method proposed in Embodiment 5 of this application. Since not every intelligent agent or user has access to the sandbox instances in the sandbox pool, this embodiment describes a possible implementation method for controlling access to sandbox instances, such as... Figure 5 As shown, the implementation method may include:
[0145] Step S51: In response to the access request for the target sandbox instance, determine the task identifier carried in the access request.
[0146] Step S52: Determine whether there is any association information between the task identifier and the target sandbox instance based on the allocation record in the central state storage.
[0147] Step S53: If associated information exists, allow access requests for the target sandbox instance.
[0148] In practical applications, AI agents or front-end users cannot directly access sandbox instances via IP and port; they need to be proxied through a unified routing gateway (control plane). Therefore, when the gateway receives an access request for a specific sandbox instance (referred to as the target sandbox instance for ease of description), it first extracts the caller's task identifier from the access request, such as extracting the projectId from the HTTP header, and uses this task identifier to determine the task or tenant to which the access request belongs.
[0149] Afterwards, you can query the allocation records in the central state storage, such as the sandbox:assigned mapping table, using the task identifier as the key to verify whether the task identifier is bound to the target sandbox instance. If the sandbox instance in the query result is consistent with the target sandbox instance, it means that there is an association between the task identifier and the target sandbox instance. If no record is found or the sandbox instance in the record is inconsistent, it means that there is no association between the task identifier and the target sandbox instance.
[0150] In this embodiment, the gateway (control plane) will only forward the access request to the target sandbox instance after the association information verification passes, so that the corresponding task can be completed in the target sandbox instance. If the verification fails, the gateway intercepts the request and returns a rejection response, such as HTTP 403 Forbidden, thus cutting off unauthorized access at the application layer. This achieves access isolation in a multi-tenant environment, prevents unauthorized access between sandboxes of different task executors, and ensures data security and the isolation of the execution environment. In this way, when facing large-scale systems with hundreds or thousands of dynamic sandbox instances serving different agents simultaneously, the above method can accurately and securely route user access requests to the dynamic IP address of a specific sandbox instance in the cluster. It also achieves reasonable and secure management of the lifecycle of AI tasks with transient characteristics such as frequent start-stop and possible abnormal interruptions, avoiding problems such as resource leakage and system instability.
[0151] In some embodiments, container cloud platforms deploy Kubernetes NetworkPolicies to configure strict Layer 3 / 4 firewall rules, such as default deny rules and whitelist admission rules, on underlying network components (e.g., Calico / Flannel), to achieve secure isolation at the data plane (network layer). Default deny rules can be used to prohibit direct cross-network connections between sandbox Pods and between ordinary user containers and sandbox Pods. Whitelist admission rules only allow gateway nodes marked with `role=gateway` to send traffic to specific ports of sandbox Pods; even if a malicious user obtains the internal IP of other sandbox instances through unauthorized means, their physical network packets will be dropped at the network layer. Thus, access authentication and control plane isolation, combined with underlying network policy isolation (such as KubernetesNetworkPolicies), constitute a dual layer of protection.
[0152] Therefore, based on Dynamic Session Token Verification, while User A's agent is running in Sandbox 1, User B cannot access Sandbox 1, ensuring multi-tenant data isolation. Specifically, during the sandbox allocation phase, the management service, while deserializing the sandbox object, can dynamically generate a unique single-use token (Session Token / Access Token, i.e., dynamic session token) for this session using a secure random algorithm. On the management side, this session token is written to a mapping table in the central state storage, such as Redis's sandbox:assignedHash, along with the tenant binding relationship (as mentioned above). On the execution side (within the sandbox), the management service injects this session token into the memory configuration of the internal service of Sandbox 1 via API, thus returning only the identification information of the successfully allocated sandbox instance and its corresponding session token to User A's calling end.
[0153] Thus, when user B initiates an access request for a target sandbox instance (such as sandbox 1 mentioned above), such as requesting to view the VNC desktop or call the Shell interface, this access request needs to explicitly carry user B's session credentials and the identification information of the target sandbox instance. Afterwards, when the gateway receives this access request, it can verify through Redis whether the session credentials match the registered credentials for sandbox 1. Since user B does not have the exclusive session credentials for sandbox 1 assigned to user A, the verification fails directly.
[0154] Specifically, the session credentials mentioned above can be forcibly bound to the lifecycle of the sandbox instance. Once the sandbox instance is cleaned up, the session credentials will be immediately invalidated in the central state storage and the physical sandbox instance. In this way, even if the session credentials are passively leaked during network transmission, due to the extremely short lifecycle of transient AI tasks, the session credentials will automatically expire within minutes, thus maximizing the defense against replay attacks.
[0155] Reference Figure 6 This is a schematic diagram of the structure of a sandbox resource processing device proposed in an embodiment of this application, as shown below. Figure 6 As shown, the sandbox resource processing device may include:
[0156] The first instance information acquisition module 61 is used to acquire the first instance information of all sandbox instances currently running on the resource orchestration platform.
[0157] The second instance information acquisition module 62 is used to acquire the second instance information of all legal sandbox instances recorded in the central state storage. The legal sandbox instances include sandbox instances in the allocated state and sandbox instances in the unallocated state.
[0158] The instance information comparison module 63 is used to compare the first instance information with the second instance information to obtain unrecorded sandbox instances in the first instance information that are not recorded in the second instance information;
[0159] The first cleanup module 64 is used to perform cleanup operations on the unrecorded sandbox instances.
[0160] Optionally, the second instance information mentioned above also includes the sandbox instance recorded in the third instance information stored in the central state storage. Based on this, the above device may further include:
[0161] The third instance information recording module is used to record the identification information of the sandbox instance in the third instance information when the sandbox instance creation begins.
[0162] The instance information update module is used to remove the identification information from the third instance information and add it to the instance information in the second instance information that is in the unallocated state after the sandbox instance is created and enters the unallocated state.
[0163] Optionally, the above-mentioned device may further include:
[0164] The creation time information acquisition module is used to acquire the creation time information of the unrecorded sandbox instance;
[0165] The creation time determination module is used to determine the creation time of the unrecorded sandbox instance based on the creation time information.
[0166] Based on this, the first cleanup module 64 is used to not perform a cleanup operation on an unrecorded sandbox instance when there is an unrecorded sandbox instance whose creation time is less than a preset time threshold.
[0167] Optionally, the above-mentioned device may further include:
[0168] The active time information update module is used to respond to the business request initiated by the intelligent agent to the sandbox instance and update the active time information of the corresponding sandbox instance in the central state storage according to the task identifier corresponding to the business request.
[0169] The inactive sandbox instance determination module is used to determine, based on the active time information, the sandbox instances that have entered the inactive state from among the sandbox instances that are in the allocated state.
[0170] The second cleanup module is used to perform cleanup operations on the sandbox instances that have entered an inactive state.
[0171] Optionally, the above-mentioned device may further include:
[0172] The quantity acquisition module is used to acquire the number of sandbox instances in the unallocated state in the second instance information;
[0173] The first creation module is used to create a new sandbox instance if the number is lower than a preset water level.
[0174] Optionally, the above-mentioned device may further include:
[0175] The sandbox instance acquisition module is used to retrieve sandbox instances in an unallocated state from the second instance information in response to a sandbox acquisition request.
[0176] The status update module is used to update the status of the acquired sandbox instance to the assigned status and record the association information between the sandbox instance and the task identifier.
[0177] The second creation module is used to create a corresponding sandbox instance for the task that initiates the sandbox acquisition request if there is no sandbox instance in the central state storage that is in an unallocated state.
[0178] Based on this, the above-mentioned device may further include:
[0179] The location record storage module is used to obtain the network location information of the sandbox instance and write it into the location record of the central state storage;
[0180] The identification information update module is used to re-add the identification information of the corresponding sandbox instance to the second instance information if the association information recording fails or the location record writing fails.
[0181] The destruction module is used to destroy the sandbox instance by calling the application programming interface of the resource orchestration platform if it fails to re-add the identifier information of the sandbox instance to the second instance information.
[0182] Optionally, the above-mentioned device may further include:
[0183] The task identifier determination module is used to determine the task identifier carried in the access request in response to the access request for the target sandbox instance.
[0184] The association information determination module is used to determine whether there is association information between the task identifier and the target sandbox instance based on the allocation record in the central state storage;
[0185] An access control module is used to allow the access request when associated information exists.
[0186] Reference Figure 7 This is a schematic diagram of the hardware structure of a sandbox resource processing system proposed in an embodiment of this application, as shown below. Figure 7 As shown, the sandbox resource processing system may include a processing device 71 and a central state storage 72. The processing device 71 can run the above-mentioned management service and is configured to implement the steps of the sandbox resource processing method proposed in any embodiment of this application. The implementation process will not be described in detail in this application.
[0187] The central state storage 72 can be used to store various state information of sandbox instances, including but not limited to the allocation record / second instance information of the aforementioned sandbox pool, which can store the identification information of sandbox instances in an unallocated state; the allocation record of sandbox instances, which can store the association information between allocated sandbox instances and task identifiers; the location record, which can store the network location information of sandbox instances; the heartbeat record, which can store the active time information of sandbox instances; and the creation record, which can store the identification information of sandbox instances that are being created.
[0188] In practical applications, the processing device 71 communicates with the resource orchestration platform to obtain information about the actual running sandbox instances and issues commands such as creation and destruction to it. The processing device 71 also communicates with the central state storage 72 to read and write various state information of the sandbox instances. Through the collaborative work of the processing device 71 and the central state storage 72, this system achieves full lifecycle management of sandbox instances, including orphan sandbox instance cleanup, heartbeat-inactive sandbox instance cleanup, pooling preheating, allocation, rollback, and access control. The implementation process of each function can be referred to the description of the corresponding method embodiments above.
[0189] It should be understood that, Figure 7 The structure of the system shown does not constitute a limitation on the system in the embodiments of this application. In practical applications, the system may include more... Figure 7 The application does not provide detailed examples of all of the following components, including more or fewer components, or combinations of certain components such as gyroscopes, accelerometers, gravity sensors, image sensors, and other sensing units used to obtain sensing parameters, power management modules, antennas, input / output components, and other communication elements.
[0190] This application also provides a computer-readable storage medium that carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the sandbox resource processing methods provided in this application.
[0191] The computer-readable storage medium can be any available medium that an electronic device can store, or a data storage device such as a training device or data center that integrates one or more available media. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
[0192] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the sandbox resource processing methods provided in this application.
[0193] When computer-readable instructions are loaded and executed on an electronic device, all or part of the processes or functions according to the embodiments of this application are generated. The electronic device may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer-readable instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer-readable instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means, which can be determined according to the actual application scenario.
[0194] This application also proposes an electronic device that may include at least one memory, a computer program stored in the memory, and at least one processor capable of running an intelligent agent. The intelligent agent can execute the computer program through the processor to implement the method steps on the intelligent agent side of the sandbox resource processing method described in the method embodiment of this application. The implementation process can be referred to the description of the corresponding method embodiment above, and will not be repeated in this embodiment.
[0195] Finally, it should be noted that the system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the system embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0196] In the above embodiments, the implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented using software plus necessary general-purpose hardware, or it can be implemented using dedicated hardware including dedicated integrated circuits, dedicated CPUs, dedicated memory, dedicated components, etc. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The various embodiments in this specification are described in a progressive or parallel manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to mutually. For the apparatuses and systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to in the method section.< / projectid> < / timestamp> < / projectid> < / deploymentid> < / deploymentid> < / deploymentid> < / projectid> < / projectid>
Claims
1. A sandbox resource processing method, characterized in that, include: Get the first instance information of all sandbox instances currently running on the resource orchestration platform; Obtain the second instance information of all legitimate sandbox instances recorded in the central state storage, wherein the legitimate sandbox instances include sandbox instances in the allocated state and sandbox instances in the unallocated state; The first instance information is compared with the second instance information to obtain the unrecorded sandbox instances in the first instance information that are not recorded in the second instance information; Perform a cleanup operation on the unrecorded sandbox instance.
2. The method according to claim 1, characterized in that, The second instance information also includes the sandbox instance recorded in the third instance information of the central state storage, and further includes: When a sandbox instance is created, its identification information is recorded in the third instance information. After the sandbox instance is created and enters the unallocated state, the identification information is removed from the third instance information and added to the unallocated instance information in the second instance information.
3. The method according to claim 1 or 2, characterized in that, Before performing the cleanup operation, the following is also included: Obtain the creation time information of the unrecorded sandbox instance; Based on the creation time information, determine the creation duration of the unrecorded sandbox instance; If there are unrecorded sandbox instances whose creation time is less than a preset time threshold, no cleanup operation will be performed on the unrecorded sandbox instances.
4. The method according to claim 1, characterized in that, Also includes: In response to a business request initiated by an agent to a sandbox instance, the active time information of the corresponding sandbox instance in the central state storage is updated according to the task identifier corresponding to the business request. Based on the active time information, determine the sandbox instances that have entered the inactive state from among the sandbox instances that were in the assigned state; Perform a cleanup operation on the sandbox instance that has entered an inactive state.
5. The method according to claim 2, characterized in that, Also includes: Obtain the number of sandbox instances in the unassigned state from the second instance information; If the quantity is lower than the preset water level, create a new sandbox instance.
6. The method according to claim 1, characterized in that, Also includes: In response to the sandbox retrieval request, a sandbox instance in an unallocated state is retrieved from the second instance information; The status of the acquired sandbox instance is updated to the assigned status, and the association information between the sandbox instance and the task identifier is recorded.
7. The method according to claim 6, characterized in that, Also includes: Obtain the network location information of the sandbox instance and write it into the location record of the central state storage; If the association information recording fails or the location record writing fails, the identification information of the corresponding sandbox instance will be re-added to the second instance information; If it fails to re-add the sandbox instance's identifier information to the second instance information, the application programming interface of the resource orchestration platform is invoked to destroy the sandbox instance.
8. The method according to claim 6, characterized in that, In response to a sandbox acquisition request, the following are also included: If there is no sandbox instance in the central state storage that is in an unallocated state, a corresponding sandbox instance is created for the task that initiated the sandbox acquisition request.
9. The method according to claim 6, characterized in that, Also includes: In response to an access request for a target sandbox instance, determine the task identifier carried in the access request; Based on the allocation records in the central state storage, determine whether there is any association information between the task identifier and the target sandbox instance; If relevant information exists, the access request is permitted.
10. A sandbox resource processing system, characterized in that, include: Processing unit, and central state storage; The processing device is configured as follows: Get the first instance information of all sandbox instances currently running on the resource orchestration platform; Obtain the second instance information of all legitimate sandbox instances recorded in the central state storage, wherein the legitimate sandbox instances include sandbox instances in the allocated state and sandbox instances in the unallocated state; The first instance information is compared with the second instance information to obtain the unrecorded sandbox instances in the first instance information that are not recorded in the second instance information; Perform a cleanup operation on the unrecorded sandbox instance.