Gray scale test method of cloud mobile phone server and related equipment
Through incremental snapshot technology and dynamic shunt method, differential snapshots of new and old versions of servers are generated, combined with hardware trusted execution environment and network micro-segmentation technology, the problem of waste of resources and high risks in the upgrade of cloud mobile server version is solved, and efficient and secure grayscale testing is achieved.
Patent Information
- Application Number
- CN202510326847.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-04
AI Technical Summary
The version upgrade method of traditional cloud mobile servers has long deployment time, huge resource consumption and high risk. The existing grayscale testing solution has failed to effectively solve the problems of resource waste and distortion of test results.
Incremental snapshot technology is used to generate differential snapshots of new and old server file systems, dynamically divert user requests through software definition network, and performance indicators are monitored in real time, triggering a rollback mechanism to ensure environmental stability, and physical isolation is achieved by combining hardware trusted execution environments and network micro-segmentation technology.
Significantly reduce the amount and time consumption of deployment data, improve testing efficiency and security, quickly identify abnormalities and automatically roll back, ensure the stability of the production environment, and reduce the impact of risks.
Smart Images

Figure CN120256222A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud computing technology, and in particular, to a gray-box testing method for cloud mobile phone servers and related devices. Background Art
[0002] In the field of cloud computing technology, the version upgrade and function verification of cloud mobile phone servers are key links to ensure service stability. Traditional testing methods usually adopt a full-scale deployment method, that is, directly replacing the old version server in the production environment with a new version server. This method requires a complete replication of the production environment, resulting in a long deployment time and huge resource consumption. For example, the full-scale image transmission may take several hours or even longer, and at the same time, it occupies a large amount of storage resources. In addition, the full-scale deployment has extremely high risks. If there are defects in the new version, it may directly affect all users, causing service interruptions or data anomalies, seriously threatening the stability of the production environment.
[0003] In the prior art, although gray-box testing has been introduced to reduce risks, conventional solutions mostly rely on full-scale images or containerized deployments. The full-scale image still needs to transmit complete data, and it fails to solve the problems of resource waste and deployment efficiency. Although the containerized solution achieves a certain degree of environmental isolation, the basic images are shared among containers, which may lead to distorted test results due to configuration differences, and it cannot completely avoid potential interference with the production environment. Therefore, there is an urgent need for a gray-box testing method for cloud mobile phone servers to solve the above-mentioned technical problems. Summary of the Invention
[0004] A series of simplified concepts are introduced in the Summary of the Invention section, which will be further described in detail in the Detailed Description section. The Summary of the Invention section of this application does not mean to attempt to define the key features and essential technical features of the claimed technical solution, nor does it mean to attempt to determine the protection scope of the claimed technical solution.
[0005] In a first aspect, this application provides a gray-box testing method for a cloud mobile phone server, characterized in that the method includes:
[0006] Obtain the file system of the old version server and the file system of the new version server;
[0007] Based on the file system of the old version server and the file system of the new version server, perform a differential analysis to generate an incremental snapshot file;
[0008] Select some nodes in the server cluster to load the incremental snapshot file and build a gray-box testing environment;
[0009] Through software-defined network, distribute user requests to the gray-box testing environment and the stable environment according to a preset traffic splitting ratio;
[0010] Collect the performance metrics of the gray test environment in real time. When it is detected that the performance metrics of the gray test environment exceed the preset threshold, trigger the rollback mechanism, unload the incremental snapshot file, and switch the corresponding node to the old version.
[0011] In some embodiments, perform a differential analysis based on the file systems of the old version server and the new version server to generate an incremental snapshot file, including:
[0012] Adopt a block-level comparison algorithm to compare the file systems of the old version server and the new version server, and find the different parts. Among them, the block-level comparison algorithm adopts the copy-on-write block-level comparison algorithm of Btrfs or ZFS;
[0013] Compress and encrypt the different parts, and add a version identifier to generate an incremental snapshot file. Among them, the incremental snapshot file is set in the incremental snapshot repository of the storage layer, and the snapshot version dependency relationship and node status are recorded through the version metadata database.
[0014] In some embodiments, select some nodes in the server cluster to load the incremental snapshot file and build a gray test environment, including:
[0015] According to the preset node selection rules, select candidate nodes with the same performance and configuration as the production environment from the server cluster;
[0016] Encrypt and isolate the memory data of the candidate nodes through the hardware trusted execution environment module;
[0017] Mount the incremental snapshot file on the candidate nodes, and synchronize the production environment dependencies and configuration parameters to complete the deployment of the gray test environment.
[0018] In some embodiments, distribute user requests to the gray test environment and the stable environment according to the preset traffic splitting ratio through software-defined networking, including:
[0019] Based on the traffic splitting rules defined by the policy engine, dynamically adjust the preset traffic splitting ratio according to the device type, user IP segment, or geographical location;
[0020] Inject a gray identification label into the HTTP header of the user request, and isolate the gray traffic and the production traffic through the network micro-segmentation controller;
[0021] Monitor the traffic load of the gray test environment in real time, and dynamically optimize the traffic splitting strategy according to the feedback of the performance metrics.
[0022] In some embodiments, collect the performance metrics of the gray test environment in real time. When it is detected that the performance metrics of the gray test environment exceed the preset threshold, trigger the rollback mechanism, unload the incremental snapshot file, and switch the corresponding node to the old version, including:
[0023] Obtain the performance metrics of nodes in the gray test environment through a second-level metric collector, where the performance metrics include at least one of CPU usage rate, memory usage rate, request latency, network bandwidth, and response time;
[0024] Push the performance metrics to the anomaly detection engine, and determine whether to trigger a rollback signal in combination with multi-dimensional thresholds;
[0025] When it is detected that the performance metrics continuously exceed the preset thresholds, generate an anomaly event log and mark the abnormal nodes;
[0026] Freeze the write operation of the gray test environment, and migrate the unfinished user sessions to a stable environment, where the stable environment is an old version server cluster that has not loaded the incremental snapshot file, to process the unshunted user requests;
[0027] Execute the incremental snapshot rollback command to restore the file system of the gray test environment to the old version state;
[0028] Update the node version information through the version metadata database, and clean up the test data in the temporary storage volume.
[0029] In some embodiments, it further includes:
[0030] Generate an incremental snapshot version chain based on the dependency tree recorded in the version metadata database;
[0031] After the rollback operation is completed, trace back to the previous stable version according to the incremental snapshot version chain, and reconstruct the snapshot dependency relationship of the gray test environment;
[0032] Perform a security check on the rolled-back nodes through the hardware trusted execution environment module to verify whether the data is secure and complete.
[0033] In some embodiments, generating an incremental snapshot version chain includes:
[0034] Starting from the baseline version, record the generation time, differential block hash value, and compression parameters of each incremental snapshot;
[0035] Perform anti-tampering protection on the incremental snapshot version chain through a digital signature algorithm, and store the signature key in an independent security module.
[0036] In a second aspect, the present application proposes a gray test device for a cloud mobile phone server, including:
[0037] A file acquisition unit, configured to acquire the file systems of the old version server and the new version server;
[0038] A snapshot generation unit that performs differential analysis based on the file systems of the old - version server and the new - version server to generate an incremental snapshot file;
[0039] An environment construction unit that selects some nodes in the server cluster to load the incremental snapshot file to construct a gray - scale test environment;
[0040] A request allocation unit that distributes user requests to the gray - scale test environment and the stable environment according to a preset diversion ratio through software - defined networking;
[0041] A monitoring and rollback unit that collects the performance metrics of the gray - scale test environment in real - time. When it detects that the performance metrics of the gray - scale test environment exceed the preset threshold, it triggers a rollback mechanism to unload the incremental snapshot file and switch the corresponding nodes to the old version.
[0042] In a third aspect, an electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program stored in the memory, it implements the steps of the gray - scale test method for the cloud mobile phone server according to any one of the first aspects.
[0043] In a fourth aspect, the present application proposes a computer - readable storage medium with a computer program stored thereon. When the computer program is executed by a processor, it implements the gray - scale test method for the cloud mobile phone server according to any one of the first aspects.
[0044] In summary, through the gray - scale test method based on incremental snapshots, the present application improves the efficiency and security of the version verification of the cloud mobile phone server. By only generating the differential snapshots of the old and new version file systems, the deployment data volume and time consumption are significantly reduced. Dynamically diverting user requests to the gray - scale test environment, combined with real - time performance monitoring and preset threshold judgment, can quickly identify anomalies and trigger an automatic rollback mechanism to ensure the stability of the production environment. In addition, this method realizes physical isolation through the hardware trusted execution environment and network micro - segmentation technology, avoiding interference between versions, and at the same time supports flexible adjustment of the test scope, taking into account both test efficiency and risk control, providing an efficient, secure and scalable solution for the gray - scale test of the cloud mobile phone server. Description of the Drawings
[0045] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to limit this specification. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0046] Figure 1 It is a schematic flowchart of the gray - scale test method for the cloud mobile phone server provided by the embodiment of the present application;
[0047] Figure 2 Structural schematic diagram of the gray box testing device for the cloud mobile phone server provided by the embodiment of the present application;
[0048] Figure 3 Structural diagram of the electronic device for gray box testing of the cloud mobile phone server provided by the embodiment of the present application. Specific implementation manners
[0049] Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims and the above-mentioned drawings of the present application are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order different from that shown or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.
[0050] Please refer to Figure 1 , which is a schematic flow chart of a gray box testing method for a cloud mobile phone server provided by the embodiment of the present application, and specifically may include:
[0051] S110. Obtain the file system of the old version server and the file system of the new version server;
[0052] Exemplarily, in the gray box testing process of the cloud mobile phone server, obtaining the file systems of the old version and the new version servers is a basic link for building a test environment. This operation aims to completely capture all the data states of the two version systems, including the operating system, application programs, configuration files, user data, etc. By comprehensively obtaining the two file systems, subsequent difference analysis can accurately identify the modified content between the versions, provide raw data support for generating incremental snapshots, and thus achieve efficient version comparison and environment deployment.
[0053] The core principle of this step is to ensure the comprehensiveness and accuracy of subsequent differential analysis by completely obtaining the file system. The old version of the file system serves as a reference benchmark, while the new version of the file system reflects the changes to be tested. The differential data between the two will directly determine the generation logic of the incremental snapshot, thereby affecting the construction efficiency of the gray-box testing environment and the reliability of the test results. This process provides a data basis for subsequent traffic diversion, performance monitoring, and rollback mechanisms in gray-box testing, and is a key prerequisite for achieving low-risk and high-efficiency version verification.
[0054] S120. Perform differential analysis based on the file systems of the old-version server and the new-version server to generate an incremental snapshot file.
[0055] Exemplarily, performing differential analysis based on the file systems of the old version and the new version of the server aims to accurately identify the changes between the two, providing a minimized data update package for subsequent gray-box testing. This step extracts only the differential parts between the versions by comparing block-level data of the file system (such as the copy-on-write mechanism of Btrfs or ZFS), avoiding resource waste caused by full-data transmission. The core of differential analysis lies in ensuring the atomicity and consistency of data comparison, thereby generating an independent snapshot file containing only incremental changes, laying an efficient and reliable data foundation for gray-box environment deployment.
[0056] The process of generating the incremental snapshot file relies on file system snapshot technology and compression algorithms. By freezing the file system state in read-only mode, the accuracy of differential analysis is ensured. The generated incremental snapshot is encrypted and version-identified, then stored in a dedicated repository and linked with the version metadata database to record the dependency relationships and node states of the snapshot. This mechanism not only reduces storage and transmission overheads but also ensures data consistency between the gray-box testing environment and the production environment through version traceability, providing reliable data support for subsequent dynamic traffic allocation and exception rollback.
[0057] S130. Select some nodes in the server cluster to load the incremental snapshot file and construct a gray-box testing environment.
[0058] Exemplarily, the core of selecting some nodes in the server cluster to construct a gray-box testing environment lies in ensuring the hardware configuration consistency and controllability between the testing environment and the production environment. By preset node selection rules (such as hardware configuration matching, low-load status, or specific area grouping), candidate nodes with the same hardware configuration as the production environment are selected, thus avoiding test result deviations caused by hardware differences. At the same time, based on the grouping rules defined by the policy engine (such as business type or geographical location), the testing scope can be flexibly divided to ensure that gray-box testing can cover the target scenarios while minimizing the impact on the production environment.
[0059] The construction of the gray-scale test environment depends on the loading of incremental snapshot files and the synchronization of environment configurations. Candidate nodes quickly obtain the differential data of the new version file system by mounting incremental snapshot files, and combine the dependencies and configuration parameters of the synchronized production environment to ensure that the software running conditions of the test environment are highly consistent with those of the production environment. In addition, the hardware trusted execution environment (TEE) module encrypts and isolates the node memory data, and the network micro-segmentation controller is used to implement independent network segment division, ensuring the security isolation between the gray-scale environment and the production environment at the physical level, preventing interference between versions and data leakage, and providing a reliable operation basis for subsequent traffic diversion and performance monitoring.
[0060] S140. Allocate user requests to the gray-scale test environment and the stable environment according to a preset diversion ratio through software-defined networking;
[0061] Exemplarily, realizing the dynamic diversion of user requests through software-defined networking (SDN) is the core execution link of gray-scale testing. The principle lies in that based on the diversion rules predefined by the policy engine (such as by device type, user IP segment, or geographical location), the SDN controller calculates and allocates user requests to the gray-scale test environment and the stable environment in real time. For example, by directing 30% of the traffic to the gray-scale nodes that load incremental snapshots, and the remaining 70% of the traffic is still processed by the stable environment, a controllable test scenario is constructed in the production environment. Through the centralized control ability of SDN, this process realizes the flexible adjustment and real-time effectiveness of the traffic allocation ratio, ensuring that the verification scope of the new version can be dynamically expanded or contracted.
[0062] The essence of traffic diversion is to achieve version isolation and data tracking through network layer technologies. The SDN gateway injects a gray-scale identification label (such as X-Version:Gray in the HTTP header) into user requests, enabling accurate identification and monitoring of gray-scale traffic during transmission. At the same time, combined with the network micro-segmentation controller to divide independent VLANs and encrypt the transmission channel (such as IPsec), it ensures the physical isolation between gray-scale traffic and the stable environment, preventing side-channel attacks between versions. This hierarchical control mechanism not only guarantees the stability of the production environment but also provides data support for subsequent anomaly detection and rollback decisions by collecting the performance metrics (such as latency, CPU utilization) of gray-scale nodes in real time, forming a closed-loop control system of "diversion - monitoring - feedback".
[0063] S150. Collect the performance metrics of the gray-scale test environment in real time. When it is detected that the performance metrics of the gray-scale test environment exceed the preset threshold, trigger the rollback mechanism, unmount the incremental snapshot file, and switch the corresponding node to the old version.
[0064] Exemplarily, the key to achieving controllable risks is to collect the performance metrics of the grayscale test environment in real time. Core metrics such as CPU utilization, memory usage, and request latency of grayscale nodes are obtained in real time through a second-level metric collector, and this data is continuously pushed to the anomaly detection engine for multi-dimensional analysis. A preset threshold (such as a latency greater than 200 ms and lasting for 10 seconds) is used as a trigger condition. When any metric continuously exceeds the standard, the system determines that there is an anomaly risk in the new version and automatically triggers a rollback mechanism, thereby achieving rapid protection of the production environment.
[0065] The core of the rollback mechanism is to restore system stability through version backtracking. Once the trigger condition is met, the automatic rollback controller immediately freezes the write operation of the grayscale node, calls the orchestration scheduler to execute an incremental snapshot rollback command (such as Btrfs-rollback--incremental), and restores the file system to the old version state. At the same time, unfinished user sessions are migrated to a stable environment, and temporary storage volumes are cleared to ensure that test data does not affect the production environment. The version metadata database synchronously updates the node status, records the rollback event, and provides a data basis for subsequent problem troubleshooting and version optimization.
[0066] In summary, this application significantly improves the efficiency and security of grayscale testing of cloud mobile phone servers through generating incremental snapshot files, building a grayscale test environment, dynamically diverting user requests, and an automatic rollback mechanism. Specifically, the incremental snapshot technology only transfers the different parts of the file system, reducing the data transfer volume by more than 90% compared to full-scale deployment, shortening the test environment construction time by more than 50%, and significantly reducing resource consumption. The physical isolation design of the grayscale environment and the stable environment, combined with SDN dynamic traffic control, can accurately control the test scope and monitor performance metrics in real time, ensuring that problems with the new version only affect a specific user group. When an anomaly is detected, the automatic rollback mechanism can complete the environment restoration within seconds, avoid the spread of failures, and at the same time ensure the continuity of the user experience through session migration. This method realizes the full-process optimization from test deployment to risk control, providing reliable support for the rapid iteration and stable operation of cloud servers.
[0067] In some instances, based on the file system of the old version server and the file system of the new version server, perform a difference analysis to generate an incremental snapshot file, including:
[0068] Adopt a block-level comparison algorithm to compare the file system of the old version server and the file system of the new version server to find the different parts, where the block-level comparison algorithm adopts the copy-on-write block-level comparison algorithm of Btrfs or ZFS;
[0069] Compress and encrypt the different parts, and add a version identifier to generate an incremental snapshot file. The incremental snapshot file is set in the incremental snapshot repository of the storage layer, and the snapshot version dependency relationship and node status are recorded through the version metadata database.
[0070] Exemplarily, this application uses the built-in copy-on-write (COW) block-level comparison algorithm of the Btrfs or ZFS file system to implement file system difference analysis. This algorithm freezes the data state of the production environment by mounting the file system in read-only mode to ensure that the data is not modified during the comparison process. At the block level granularity (such as 4KB / block), the algorithm scans the old and new version file systems block by block, identifies the changed disk sectors, and only extracts the data blocks of the different parts. This mechanism avoids the resource waste of full data comparison, and at the same time ensures the accuracy and consistency of difference analysis through atomic operations, ensuring that the generated incremental snapshot only contains the actual changed content between versions.
[0071] For the identified different data blocks, use the zstd compression algorithm for efficient compression (the compression level is preset to 15:1), significantly reducing the volume and transmission overhead of the snapshot file. The compressed data is encrypted through a digital signature algorithm (such as RSA-2048), and the key is securely stored in combination with the hardware trusted execution environment (TEE) module to prevent the snapshot data from being tampered with or leaked during transmission or storage. The encrypted incremental snapshot file is added with a unique version identifier (such as a hash value, a timestamp, and a baseline version number) to form an immutable snapshot package, ensuring the reliability of version tracing and dependency management.
[0072] The generated incremental snapshot file is stored in the incremental snapshot repository of the storage layer and is classified and managed using the naming specification of "baseline version_target version_timestamp". At the same time, the version metadata database records the generation time, the hash value of the different blocks, and the compression parameters of each snapshot by maintaining a snapshot dependency tree (such as v1.2→v1.3→v1.4), supporting the backtracking and verification of the version chain. The version metadata database is synchronized with the orchestration scheduler in real time to dynamically update the snapshot loading status and health metrics of the nodes, providing data support for gray environment deployment, traffic dynamic diversion, and abnormal rollback. This mechanism protects the integrity of the version chain through a tamper-proof signature (such as HMAC-SHA256), ensuring that the system can accurately backtrack to the specified state during rollback or version switching, realizing the controllability and security of the whole process.
[0073] In some instances, select some nodes in the server cluster to load the incremental snapshot file to build a gray test environment, including:
[0074] According to the preset node selection rules, select candidate nodes from the server cluster with the same performance and configuration as the production environment;
[0075] Encrypt and isolate the memory data of candidate nodes through the hardware trusted execution environment module;
[0076] Mount the incremental snapshot file on the candidate node, and synchronize the production environment dependencies and configuration parameters to complete the deployment of the gray test environment.
[0077] Exemplarily, when selecting some nodes from the server cluster to build the gray test environment, first filter the candidate nodes according to the preset node selection rules. The rules include: hardware configuration consistency, requiring that the hardware parameters such as CPU model, memory capacity, and storage type of the candidate nodes are exactly the same as those of the production environment to ensure the performance comparability between the test environment and the production environment; version dependency matching, verifying the old version base dependencies (such as system libraries, configuration files) deployed on the candidate nodes through the version metadata database to ensure the compatibility of the incremental snapshot loading and running; network isolation feasibility, allocating an independent network segment (such as a specific VLAN or IPsec encrypted channel) for the candidate nodes by the network micro-segmentation controller to achieve physical isolation between the gray traffic and the production traffic. The node selection process is completed by the orchestration scheduler in collaboration with the policy engine, and preferentially selects low-load nodes or servers in a specified area to minimize the impact on the production environment.
[0078] After selecting the candidate nodes, encrypt and isolate their memory data through the hardware trusted execution environment module (such as Intel SGX). Based on the trusted execution environment technology, this module divides the runtime memory of the gray nodes into an encrypted protected area and a non-protected area. The encrypted protected area stores the incremental snapshot file, sensitive configuration parameters, and user session data, and realizes static and dynamic data protection through hardware-level encryption algorithms (such as AES-256) to prevent side-channel attacks or unauthorized access. At the same time, the network micro-segmentation controller configures an independent security group policy for the gray nodes to restrict the network communication between cross-version nodes and ensure double isolation between the gray environment and the production environment at the logical and physical levels.
[0079] When mounting the incremental snapshot file on the candidate node, load the differential data blocks in read-only mode, and use the file system write-back redirection technology (such as the snapshot mounting mechanism of Btrfs / ZFS) to overlay the changed content on the old version base image to form a complete new version running environment. During the deployment process, the configuration management center synchronizes the dependencies (such as database connection strings, service port numbers) and environment variables from the production environment to the gray nodes to ensure the functional consistency between the test environment and the production environment. After completing the mounting and synchronization, the version metadata database records the snapshot version number, dependency relationship, and health status of the nodes to form the metadata baseline of the gray test environment. This process is uniformly coordinated by the orchestration scheduler and realized through automated scripts or API calls for one-key deployment, and finally constructs a gray test cluster that is highly consistent with the production environment and safely isolated.
[0080] In some instances, user requests are distributed to the gray - scale test environment and the stable environment according to a preset traffic - splitting ratio through software - defined networking, including:
[0081] Based on the traffic - splitting rules defined by the policy engine, dynamically adjust the preset traffic - splitting ratio according to device type, user IP segment, or geographical location;
[0082] Inject a gray - scale identification label into the HTTP header of the user request, and isolate the gray - scale traffic and production traffic through a network micro - segmentation controller;
[0083] Real - time monitor the traffic load of the gray - scale test environment, and dynamically optimize the traffic - splitting strategy according to the feedback of performance metrics.
[0084] Exemplarily, define the traffic - splitting rules for user requests through the policy engine. Based on multi - dimensional conditions such as device type (e.g., mobile / PC), user IP segment (e.g., a specific geographical area), or geographical location (e.g., city / country), dynamically adjust the preset traffic - splitting ratio (e.g., 30% gray - scale traffic). The policy engine is linked with the version metadata database to obtain the version status and health metrics of the gray - scale nodes in real time, and generate a dynamic traffic - splitting strategy in combination with a load - balancing algorithm (such as weighted round - robin or consistent hashing). For example, give priority to allocating high - end device users to the gray - scale environment to verify new functions, or dynamically reduce the traffic - splitting ratio according to the CPU utilization rate of the nodes to avoid overload risks. The traffic - splitting strategy is sent to the SDN gateway cluster through the SDN controller to ensure strict consistency between traffic allocation and policy definition.
[0085] After receiving the user request, the SDN gateway cluster injects a gray - scale identification label (such as X - Version:Gray) into the HTTP request header according to the traffic - splitting strategy, for identification and routing to the gray - scale node pool. The network micro - segmentation controller synchronously divides an independent network channel for the gray - scale traffic, for example, through VLAN isolation or an IPsec encrypted tunnel, to achieve physical isolation from the production traffic. In specific implementation, the controller configures an independent security group policy for the gray - scale nodes, restricts cross - version communication with the stable nodes, and ensures the bandwidth priority of the gray - scale traffic through traffic - shaping technology (such as QoS policy). The collaborative mechanism of label injection and network isolation not only avoids data interference between versions but also ensures the traceability and security of the gray - scale traffic.
[0086] The metric collector collects performance metrics (including request latency, error rate, resource utilization, etc.) of the gray environment at the second-level granularity and pushes them to the anomaly detection engine for multi-dimensional threshold judgment (such as latency > 200ms for 10 seconds). The monitoring data is fed back to the policy engine in real time, triggering the dynamic optimization of the traffic splitting policy: for example, when it is detected that the CPU utilization rate of the gray node exceeds 80%, the traffic splitting ratio is automatically reduced to 10%; or when the error rate in the gray environment surges, the traffic splitting rule is temporarily switched to hash distribution by user ID to disperse the load. The optimized policy takes effect immediately through the SDN controller, forming a closed-loop control of "monitoring → analysis → adjustment". This mechanism records the policy change history through the version meta-database and collaborates with the orchestration scheduler to ensure the flexibility of traffic scheduling and the stability of the system.
[0087] In some instances, the performance metrics of the gray test environment are collected in real time. When it is detected that the performance metrics of the gray test environment exceed the preset threshold, a rollback mechanism is triggered to unload the incremental snapshot file and switch the corresponding node to the old version, including:
[0088] Obtain the performance metrics of the nodes in the gray test environment through a second-level metric collector, where the performance metrics include at least one of CPU usage, memory usage, request latency, network bandwidth, and response time;
[0089] Push the performance metrics to the anomaly detection engine and combine multi-dimensional thresholds to determine whether to trigger a rollback signal;
[0090] In the case of detecting that the performance metrics continuously exceed the preset threshold, generate an anomaly event log and mark the abnormal nodes;
[0091] Freeze the write operations in the gray test environment and migrate the unfinished user sessions to the stable environment, where the stable environment is an old version server cluster that has not loaded the incremental snapshot file to handle the unsplit user requests;
[0092] Execute the incremental snapshot rollback command to restore the file system of the gray test environment to the old version state;
[0093] Update the node version information through the version meta-database and clean up the test data in the temporary storage volume.
[0094] Exemplarily, when collecting the performance metrics of the grayscale test environment in real time, key parameters such as the CPU usage rate, memory occupancy rate, request latency, network bandwidth, and response time of the node are obtained through a second-level metric collector at a sampling interval of 500 ms. The collector is implemented based on a lightweight proxy or kernel module to ensure low invasiveness to system resources. After receiving the metric data, the anomaly detection engine uses a multi-dimensional threshold judgment logic (such as independent threshold triggers: latency > 200 ms for 10 seconds, CPU utilization > 90% for 30 seconds), and combines a sliding window algorithm to verify the continuous over-standard situation of the metrics. When it detects that the preset threshold is breached, the engine generates an anomaly event log (including timestamp, node ID, over-standard metric type and value), and pushes it to the automatic rollback controller through the event bus to trigger the rollback decision process.
[0095] After triggering the rollback mechanism, the system first freezes the write operations of the grayscale nodes, and prohibits new data from being written through a distributed lock or file system lock (such as the write lock of Btrfs / ZFS) to prevent data inconsistency during the rollback process. Uncompleted user sessions are migrated to the stable environment in real time by the SDN gateway cluster. The session persistence technology (such as the multi-path transmission of the QUIC protocol) is used during the migration process to ensure that the connection is not interrupted and the user is unaware. The stable environment is an old version server cluster without loading the incremental snapshot. It synchronizes the latest production configuration through the version metadata database to ensure seamless takeover of the grayscale traffic. After the migration is completed, the orchestration scheduler calls the underlying file system command (such as btrfs rollback --incremental) to restore the file system of the grayscale node to the old version state and completely remove the changes of the incremental snapshot.
[0096] After the rollback operation is executed, the version metadata database updates the version information of the abnormal node (such as rolling back from v1.3 to v1.2), and records the metadata of the rollback event (including rollback time, trigger metrics, operators, etc.). The uncommitted test data in the temporary storage volume (such as / tmp / gray_cache) is cleaned up through an automated script to avoid residual data interfering with subsequent tests or occupying storage resources. At the same time, the hardware trusted execution environment module performs an integrity check on the rolled-back node to verify the consistency of the file system hash value with the old version baseline to ensure data security. The network micro-segmentation controller synchronously removes the independent network segment configuration of the grayscale node and re-integrates it into the production environment network policy to complete the full closed-loop management from anomaly detection to environment recovery.
[0097] In some instances, it also includes:
[0098] Based on the dependency tree recorded in the version metadata database, an incremental snapshot version chain is generated, including:
[0099] Starting from the baseline version, record the generation time, differential block hash value, and compression parameters of each incremental snapshot;
[0100] The incremental snapshot version chain is protected against tampering through a digital signature algorithm, and the signature key is stored in an independent security module;
[0101] After the rollback operation is completed, trace back to the previous stable version according to the incremental snapshot version chain, and reconstruct the snapshot dependency relationship of the gray test environment;
[0102] The rolled-back nodes are securely verified through the hardware trusted execution environment module to verify whether the data is secure and complete.
[0103] Exemplarily, when generating the incremental snapshot version chain based on the dependency tree recorded in the version metadata database, starting from the baseline version (such as v1.0), the generation time, differential block hash value, compression parameters (such as zstd compression level 15:1), and digital signature of each incremental snapshot are recorded in a chained manner. The dependency tree maintains the parent-child relationship between versions through a directed acyclic graph (DAG) structure. For example, versions v1.1 → v1.2 → v1.3 ensure the linear logic of version tracing. Each snapshot package is digitally signed through an asymmetric encryption algorithm (such as RSA-2048), and the signature key is stored in the hardware security module (HSM) to prevent the version chain from being tampered with. The version metadata database synchronously stores the metadata and dependency relationship of the snapshot package to form a complete version history record, providing a verifiable traceability basis for the rollback operation.
[0104] After the rollback operation is completed, the system traces back to the previous stable version according to the incremental snapshot version chain (such as rolling back from v1.3 to v1.2). The orchestrator extracts the snapshot dependency relationship and node configuration information of the target version from the version metadata database, and calls the file system command (such as btrfs rollback --incremental) to restore the file system state of the gray nodes. At the same time, the abnormal version node (v1.3) is removed from the dependency tree, and the logical link from v1.2 to subsequent versions is reconstructed (if there are subsequent legal snapshots). The configuration management center synchronously updates the environment variables and dependencies of the gray nodes to ensure their consistency with the production environment. This process ensures the integrity of the operation through an atomic transaction. If any sub-step fails, the transaction rollback mechanism is triggered to avoid inconsistent system states.
[0105] After rollback, the node needs to be securely verified through a hardware trusted execution environment module (such as Intel SGX). The verification process includes: data integrity verification, calculating the hash value of the file system after rollback and comparing it with the baseline version hash recorded in the version metadata database to ensure no data tampering; memory encryption status check, verifying the seal of the memory encryption protection area of the gray node to prevent sensitive data leakage; key legality audit, verifying through the HSM whether the signature key and session key used during the rollback process have not been revoked or leaked. The verification results are recorded in the security audit log. If anomalies are found (such as hash mismatch or key failure), an alarm is triggered and the node is isolated. The network micro-segmentation controller synchronously updates the network policy of the node, restricting its communication range until the verification passes, ultimately achieving the security and trustworthiness of the rollback environment.
[0106] The technical solution of this application will be further described in detail through specific embodiments below.
[0107] Embodiment 1 is an analysis of the technical implementation process and effects of 5G cloud game version updates. In this embodiment, taking the 5G cloud game server version update as the scenario, the specific application of the gray-box testing method based on incremental snapshots is demonstrated.
[0108] Compare the file systems of the old version v3.2 and the new version v3.3 of the game engine. Use the COW block-level difference analysis algorithm built into the ZFS file system to scan each block to identify the changed data. By mounting the production environment file system in read-only mode, ensure the stability of the data state during the difference analysis process. Only extract the actually changed disk blocks (4KB granularity) such as model files and configuration files to generate an incremental snapshot file. Use the zstd compression algorithm (compression ratio 15:1) to efficiently compress the difference data, reducing the snapshot volume from 120GB to 8GB, and through digital signature (RSA-2048) and TLS1.3 encrypted transmission channel, ensure data integrity and security. The finally generated incremental snapshot file is stored in the incremental snapshot repository of the storage layer, and the version metadata database records its dependency tree (v3.2→v3.3) and metadata (such as difference block hash value, compression parameters).
[0109] According to the preset node selection rules, select 20% of the candidate nodes in the East China Region 2 server cluster with the same hardware configuration as the production environment (such as the same CPU model and memory capacity), and allocate an independent VLAN (such as VLAN1002) for them through the network micro-segmentation controller to achieve physical isolation. After encrypting and isolating the memory data of the candidate nodes using the hardware Trusted Execution Environment (TEE) module, pull the encrypted snapshot file from the incremental snapshot repository, mount it to the candidate nodes through the Btrfs / ZFS snapshot loading function, synchronize the production environment dependencies (such as database connection configuration, middleware version) and environment variables, and build a gray test environment that is exactly the same as the production environment. The remaining 80% of the nodes form a stable node pool and continue to run version v3.2 to ensure the stability of the production environment.
[0110] Allocate user requests to the gray node pool and the stable node pool according to the device type. The SDN gateway cluster performs dynamic traffic splitting. For user requests for high-end GPU devices (such as RTX4090), inject 100% gray labels (such as HTTP header X-Version:Gray) and direct them to the gray node pool to test the new version's image quality enhancement function; requests for mid- and low-end devices remain in the original version. At the same time, the SDN gateway monitors the traffic load of the gray nodes in real time, obtains data such as CPU utilization and request latency through the second-level metric collector, and pushes it to the anomaly detection engine for multi-dimensional analysis to dynamically optimize the traffic splitting strategy to balance the load.
[0111] During the gray test process, the monitoring system detects that the delay peak of the gray nodes continuously exceeds 200ms (preset threshold) for 10 seconds due to a defect in the new version of the ray tracing module, triggering the automatic rollback mechanism. The automatic rollback controller immediately freezes the write operations of the gray nodes, calls the Btrfs incremental rollback command (btrfs rollback --incremental) to restore the file system to version v3.2, and migrates the session to the stable node pool through the SDN gateway. At the same time, clean up the temporary storage volume of the gray nodes (such as / tmp / gray_cache), update the node status to v3.2 in the version metadata database, and record the rollback event. The entire rollback process takes 15 seconds and successfully restores the stability of the production environment.
[0112] Verification through this embodiment shows that the gray-box testing method based on incremental snapshots significantly improves the version update efficiency and security. The deployment time is reduced from 2 hours for full-scale deployment to 6 minutes, a 95% reduction in deployment time. In terms of resource consumption, the incremental snapshot volume is only 6.7% of the full-scale image, saving 93.3% of storage and transmission resources. New version issues only affect users of high-end GPU devices (accounting for <5%), avoiding affecting all users. Version rollback is completed within 15 seconds, with a 40-fold improvement in recovery speed compared to traditional solutions. This embodiment fully demonstrates the technical advantages of the present invention in reducing testing risks, shortening the iteration cycle, and optimizing resource utilization, and is applicable to 5G cloud game scenarios with extremely high requirements for real-time performance and stability.
[0113] Embodiment 2 is an analysis of the implementation process and effects of the hot upgrade technology for AI model services. Taking the hot upgrade of AI model services as the scenario, this embodiment demonstrates the specific application of the gray-box testing method proposed in this application.
[0114] During the hot upgrade of AI model services, the file system differences between the old version BERT-base model and the new version BERT-large model are analyzed through the Btrfs block-level comparison algorithm. First, the file system of the production environment is mounted in read-only mode to ensure an accurate comparison in a data-frozen state. The algorithm scans the file system at a block granularity of 4KB to identify the changed parts that only involve model parameters (such as weight matrices, bias terms) and configuration files (such as hyperparameter settings). The generated differential data blocks are efficiently compressed through the zstd compression algorithm, with a compression ratio as high as 3.7:1, reducing the snapshot volume from 3.2GB in full scale to 860MB. The compressed incremental snapshot is stored in the incremental snapshot repository after being digitally signed with RSA-2048, and its dependency relationship (BERT-base → BERT-large) and metadata are recorded through the version metadata database.
[0115] In the inference service cluster, according to the hardware configuration consistency rules (CPU model, GPU model, memory capacity matching the production environment) and network isolation requirements, some nodes are selected to build a gray-box testing environment. The selected nodes encrypt and isolate the memory data through the hardware trusted execution environment module (Intel SGX) and allocate an independent VLAN (such as VLAN1003) to achieve physical isolation from the production traffic. During deployment, the orchestration scheduler pulls the BERT-large snapshot file from the incremental snapshot repository, mounts it on the nodes in read-only mode, and synchronizes the dependencies (such as database connection addresses, API keys) and running parameters of the production environment. The remaining nodes are reserved as a stable node pool to continue running the BERT-base version to serve the user requests that are not diverted.
[0116] Define the traffic splitting strategy through the Software-Defined Network (SDN) controller: Dynamically allocate traffic according to the user ID hash value. The specific rules are as follows: 10% of user requests are directed to the gray node pool to verify the inference performance and accuracy of the BERT-large model; 90% of user requests are retained in the stable node pool to maintain the BERT-base version service. The SDN gateway cluster injects the X-Model-Version:Gray label into the HTTP header of user requests and isolates the gray traffic through an IPsec encrypted tunnel. The real-time monitoring system collects the inference latency, accuracy, and resource utilization metrics of gray nodes and dynamically optimizes the traffic splitting strategy (for example, when it is detected that the node load exceeds 70%, temporarily reduce the traffic splitting ratio to 5%).
[0117] During the gray box testing process, the second-level metric collector detected that the inference accuracy of some gray nodes decreased by 8% (caused by the overfitting problem of the BERT-large model). The anomaly detection engine determined that the preset threshold was met (accuracy decrease > 5%), triggering the automatic rollback mechanism:
[0118] Freeze the write operation of gray nodes through the Btrfs file system lock. Uncompleted inference requests are seamlessly migrated to the stable node pool through the QUIC protocol to ensure seamless user switching; perform an incremental snapshot rollback, execute the btrfs rollback --incremental command to restore the file system of gray nodes to the BERT-base version state, and unmount the BERT-large incremental snapshot;
[0119] Delete the test logs and cache data in the temporary storage volume, update the node version to BERT-base in the version metadata database, and verify the file system hash value through the hardware TEE module to ensure data integrity.
[0120] The incremental snapshot technology in the embodiments of this application reduces the full-scale deployment time from 30 minutes to zero downtime, and the resource consumption is reduced by 74% (from 3.2GB to 860MB); only 10% of user requests participate in the gray box testing, avoiding affecting all users, and reducing the problem impact range by 10 times; it only takes 10 seconds from anomaly detection to complete rollback, ensuring the high availability of the AI model service and the continuity of the user experience; through hardware encryption isolation and network micro-segmentation technology, the zero data leakage risk between the gray box environment and the production environment is achieved. This Embodiment 2 verifies the efficiency, security, and flexibility of the present invention in complex scenarios, providing a reusable technical paradigm for the version iteration of large-scale AI services.
[0121] Embodiment 3 is the implementation process and effect analysis of the batch update technology for edge computing nodes; this embodiment focuses on the batch version update scenario of edge computing nodes, comprehensively demonstrating the technical advantages and application value of the gray box testing method proposed in this application.
[0122] During the upgrade of the Edge Operating System (OS) image from v2.1 to v2.2, the copy-on-write block-level comparison algorithm of the ZFS file system is adopted to analyze the differences between the file systems of the two versions. By mounting the file system of the edge nodes in the production environment in read-only mode, an accurate comparison under the data frozen state is ensured. The algorithm scans the file system at a block level granularity of 4KB to identify the changed parts involving the operating system kernel modules, device drivers, and configuration files (such as network policies). The generated differential data blocks are efficiently compressed by the zstd compression algorithm (compression level 15), with a compression ratio as high as 13.8:1. The snapshot volume is reduced from the full 18GB to 1.3GB. The compressed incremental snapshot is stored in the incremental snapshot repository after being digitally signed with RSA-2048, and its dependencies (v2.1→v2.2) and metadata (such as generation timestamp, differential block hash value) are recorded through the version metadata database.
[0123] The deployment process is divided into two stages to verify cross-region compatibility. In the first stage, in the edge computing cluster in the North China region, according to the hardware configuration consistency rules (such as CPU architecture, storage type matching the production environment) and network isolation requirements, 5% of the nodes are selected to build a gray test environment. The selected nodes encrypt and isolate the memory data through the hardware trusted execution environment module (such as Intel SGX) and allocate an independent VLAN (such as VLAN1004) to achieve traffic isolation. The orchestration scheduler pulls the v2.2 snapshot file from the incremental snapshot repository, mounts it on the nodes in read-only mode, and synchronizes the dependencies of the edge services (such as IoT device drivers, local database configurations).
[0124] In the second stage, 10% of the nodes are selected in the South China region to expand the gray test scope, focusing on verifying the version compatibility under different geographical environments (such as humidity, network latency differences). During the deployment process, the network micro-segmentation controller configures an independent IPsec encryption channel for the South China nodes to ensure the security of cross-region traffic.
[0125] By defining a geographical location-driven traffic diversion strategy through the Software-Defined Network (SDN) controller, for requests from Beijing users, 100% are routed to the North China gray nodes to test the performance of the v2.2 version in the northern network environment; for requests from Guangzhou users, 100% are routed to the South China gray nodes to verify the stability of the version in the high-humidity environment in the south.
[0126] The SDN gateway cluster injects the X-Edge-Version:Gray label into the HTTP header of the user request and achieves accurate routing through geographical DNS resolution. The real-time monitoring system collects the CPU utilization rate, local storage I / O, and network jitter metrics of the gray nodes and dynamically optimizes the traffic distribution (such as when it is detected that the load of the North China nodes exceeds 80%, some requests are temporarily diverted to the South China nodes).
[0127] During the gray-box testing process, the second-level metric collector detected that the CPU utilization rate of some nodes in the North China region continuously exceeded 90% (caused by the incompatibility between the new device driver of v2.2 version and the hardware). The anomaly detection engine determined that the preset threshold was met (CPU > 90% for 5 consecutive minutes), triggering the full-region automatic rollback mechanism, that is:
[0128] Freeze the write operations of the abnormal nodes through the ZFS file system lock. The unfinished edge computing tasks (such as real-time video analysis) are migrated to the stable node pool through a low-latency transmission protocol (such as Web RTC) to ensure business continuity; execute the zfsrollback command to restore the file systems of the gray-box nodes in North China and South China to the v2.1 version state, and uninstall the v2.2 incremental snapshot; delete the cache data and log files in the temporary storage volume, update the node version to v2.1 in the version metadata database, and verify the file system hash value through the hardware TEE module to ensure data integrity. The network micro-segmentation controller synchronously releases the independent network policy of the gray-box nodes and incorporates them back into the production environment.
[0129] The incremental snapshot technology in the embodiments of this application shortens the full-batch update time from 3 days to 2 hours, and reduces the resource consumption by 92.8% (from 18GB to 1.3GB); the phased deployment strategy effectively identifies the compatibility issues between the North China and South China environments and reduces the global risk; it only takes <10 seconds for a single node from anomaly detection to the completion of full-region rollback, and the overall efficiency is 40 times higher than the traditional solution; dynamic traffic control realizes load balancing and avoids service interruption caused by overloading of edge nodes; through hardware encryption isolation and geographical traffic segmentation, it ensures zero leakage risk of sensitive data (such as local IoT device information). Embodiment 3 of this application verifies the high efficiency, flexibility, and robustness of the present invention in the distributed edge computing scenario, and provides an extensible technical solution for the iteration of large-scale edge nodes.
[0130] Please refer to Figure 2 , which is a schematic structural diagram of a gray-box testing device for a cloud mobile phone server provided by the embodiments of this application, including:
[0131] A file acquisition unit 21, configured to acquire the file system of the old-version server and the file system of the new-version server;
[0132] A snapshot generation unit 22, which performs differential analysis based on the file system of the old-version server and the file system of the new-version server to generate an incremental snapshot file;
[0133] An environment construction unit 23, configured to select some nodes in the server cluster to load the incremental snapshot file and construct a gray-box testing environment;
[0134] A request allocation unit 24 is configured to allocate user requests to a gray-scale test environment and a stable environment according to a preset traffic splitting ratio through a software-defined network;
[0135] A monitoring and rollback unit 25 is configured to collect performance metrics of the gray-scale test environment in real time. When it detects that the performance metrics of the gray-scale test environment exceed a preset threshold, it triggers a rollback mechanism to uninstall the incremental snapshot file and switch the corresponding node to the old version.
[0136] Please refer to Figure 3 , an embodiment of the present application further provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored on the memory 310 and executable on the processor. When the processor 320 executes the computer program 311, it implements the steps of any method of the gray-scale test device of the cloud mobile phone server.
[0137] Since the electronic device introduced in this embodiment is the device used to implement a gray-scale test device of a cloud mobile phone server in an embodiment of the present application, based on the method introduced in the embodiment of the present application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in the embodiment of the present application will not be described in detail here. As long as the device used by those skilled in the art to implement the method in the embodiment of the present application belongs to the scope protected by the present application.
[0138] In the specific implementation process, when the computer program 311 is executed by the processor, it can implement any implementation manner in the corresponding embodiment of the first aspect.
[0139] It should be noted that in the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0140] Those skilled in the art should understand that the embodiments of the present application can provide a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media containing computer-readable program code.
[0141] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in one block or multiple blocks.
[0142] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in one block or multiple blocks.
[0143] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in one block or multiple blocks.
[0144] Embodiments of the present application also provide a computer program product, which includes computer software instructions. When the computer software instructions run on a processing device, the processing device is caused to execute Figure 1 the process of a gray-box testing method for a cloud mobile phone server in a corresponding embodiment.
[0145] A computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, they wholly or partly generate a process or function in accordance with the embodiments of the present application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium may be any available medium that can be stored by the computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium, an optical medium, or a semiconductor medium, etc.
[0146] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0147] In several embodiments provided in the present application, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other may be through some interfaces, and the indirect coupling or communication connection of the devices or units may be in an electrical, mechanical, or other form.
[0148] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or distributed to multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0149] In addition, the functional units in the various embodiments of the present application may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware and / or software functional units.
[0150] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device to execute all or part of the steps of the methods of the various embodiments of this application.
[0151] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of this application.
[0152] Although the preferred embodiments of this specification have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of this specification.
[0153] Obviously, those skilled in the art can make various changes and deformations to this specification without departing from the spirit and scope of this specification. In this way, if these modifications and deformations of this specification fall within the scope of the claims of this specification and their equivalent technologies, this specification also intends to include these modifications and deformations.
Claims
1. A gray-box testing method for a cloud mobile phone server, characterized in that, The method includes: Obtaining the file system of the old version server and the file system of the new version server; Performing a difference analysis based on the file system of the old version server and the file system of the new version server to generate an incremental snapshot file; Selecting some nodes in the server cluster to load the incremental snapshot file to construct a gray test environment; Allocating user requests to the gray test environment and the stable environment according to a preset shunting ratio through software-defined networking; Real-time collecting the performance metrics of the gray test environment. When it is detected that the performance metrics of the gray test environment exceed a preset threshold, triggering a rollback mechanism to unload the incremental snapshot file and switch the corresponding nodes to the old version.
2. The method according to claim 1, wherein The performing a difference analysis based on the file system of the old version server and the file system of the new version server to generate an incremental snapshot file includes: Adopting a block-level comparison algorithm to compare the file system of the old version server and the file system of the new version server to find the different parts, where the block-level comparison algorithm adopts the copy-on-write block-level comparison algorithm of Btrfs or ZFS; Compressing and encrypting the different parts and adding a version identifier to generate the incremental snapshot file, where the incremental snapshot file is set in the incremental snapshot repository of the storage layer, and the snapshot version dependency relationship and node status are recorded through the version metadata database.
3. The method according to claim 1, wherein The selecting some nodes in the server cluster to load the incremental snapshot file to construct a gray test environment includes: Selecting candidate nodes with the same performance and configuration as the production environment from the server cluster according to a preset node selection rule; Encrypting and isolating the memory data of the candidate nodes through the hardware trusted execution environment module; Mounting the incremental snapshot file on the candidate nodes and synchronizing the production environment dependencies and configuration parameters to complete the deployment of the gray test environment.
4. The method according to claim 1, characterized in that, The allocating user requests to the gray test environment and the stable environment according to a preset shunting ratio through software-defined networking includes: Dynamically adjusting the preset shunting ratio based on the shunting rules defined by the policy engine according to device type, user IP segment, or geographical location; Injecting a gray identification label into the HTTP header of the user request and isolating the gray traffic and production traffic through the network micro-segmentation controller; Real-time monitoring the traffic load of the gray test environment and dynamically optimizing the shunting strategy according to the feedback of the performance metrics.
5. The method according to claim 1, wherein The real-time collecting the performance metrics of the gray test environment. When it is detected that the performance metrics of the gray test environment exceed a preset threshold, triggering a rollback mechanism to unload the incremental snapshot file and switch the corresponding nodes to the old version includes: Obtaining the performance metrics of the nodes in the gray test environment through a second-level metric collector, where the performance metrics include at least one of CPU usage rate, memory usage rate, request latency, network bandwidth, and response time; Pushing the performance metrics to the anomaly detection engine and combining multi-dimensional thresholds to determine whether to trigger a rollback signal; When it is detected that the performance metrics continuously exceed the preset threshold, generating an anomaly event log and marking the anomaly nodes; Freeze the write operation of the grayscale test environment, and migrate the unfinished user sessions to the stable environment, where the stable environment is an old version server cluster that has not loaded the incremental snapshot file to process the unshunted user requests; Execute the incremental snapshot rollback command to restore the file system of the grayscale test environment to the old version state; Update the node version information through the version metadata database and clean up the test data in the temporary storage volume.
6. The method according to claim 5, wherein It also includes: Generate an incremental snapshot version chain based on the dependency tree recorded in the version metadata database; After the rollback operation is completed, trace back to the previous stable version according to the incremental snapshot version chain and reconstruct the snapshot dependency relationship of the grayscale test environment; Perform a security check on the rolled-back nodes through the hardware trusted execution environment module to verify whether the data is secure and complete.
7. The method according to claim 6, characterized in that, The generating of the incremental snapshot version chain includes: Starting from the baseline version, record the generation time, differential block hash value, and compression parameters of each incremental snapshot; Protect the incremental snapshot version chain against tampering through a digital signature algorithm and store the signature key in an independent security module.
8. A gray box testing device for a cloud mobile phone server, characterized in that, It includes: A file acquisition unit for acquiring the file systems of the old version server and the new version server; A snapshot generation unit that performs differential analysis based on the file systems of the old version server and the new version server to generate an incremental snapshot file; An environment construction unit for selecting some nodes in the server cluster to load the incremental snapshot file to construct a grayscale test environment; A request allocation unit for allocating user requests to the grayscale test environment and the stable environment according to a preset shunting ratio through a software-defined network; A monitoring and rollback unit for real-time collecting the performance metrics of the grayscale test environment, and triggering a rollback mechanism when it detects that the performance metrics of the grayscale test environment exceed the preset threshold, unloading the incremental snapshot file and switching the corresponding nodes to the old version.
9. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is used to implement the steps of the grayscale test method of the cloud mobile phone server as described in any one of claims 1 to 7 when executing the computer program stored in the memory.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The computer program, when executed by the processor, implements the grayscale test method of the cloud mobile phone server as described in any one of claims 1 to 7.
Citation Information
Cited By
Server content updating and rollback method based on static technology
CN120892234A