A software remote testing system and method based on a distributed architecture
By building an exception detection mechanism and topological clipping based on state trajectory, the optimal scheduling path is dynamically generated, which solves the task interruption problem of the distributed remote testing system when node failures is solved, and seamless migration and state reconstruction of tasks in a distributed environment is realized, improving the stability and consistency of the system.
Patent Information
- Application Number
- CN202510704115.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-05-29
AI Technical Summary
When existing distributed remote testing systems face sudden failures of test nodes or network exceptions, they cannot realize early perception and dynamic response to node status, resulting in task interruption, loss of results, wasted resources and inconsistent task semantics, poor test execution stability, insufficient adaptability and reliability.
By collecting the multi-dimensional running state of remote test nodes, building time series state vectors and state trajectories, generating exception scores, eliminating faulty nodes, building a new topology diagram and updating the scheduling structure, extracting task status snapshots and performing cross-platform encapsulation, performing state reconstruction and consistency scoring, and dynamically determining task execution strategies.
Accurate modeling and risk prediction of remote test nodes are realized, optimal scheduling paths are generated dynamically, and the stability and continuity of tasks in distributed environments are ensured, and the scheduling stability and task execution consistency of the system in node fluctuation scenarios are improved.
Smart Images

Figure CN120234253B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software testing, and more specifically, to a software remote testing system and method based on a distributed architecture. Background Art
[0002] In the development and operation and maintenance process of modern large-scale software systems, the complexity and coverage of software testing are continuously expanding. The traditional testing mode in a single-machine or local network environment has difficulty meeting the testing requirements of distributed, heterogeneous, and highly real-time systems. With the rapid development of cloud computing, edge computing, and microservice architectures, the deployment environment of software systems tends to be decentralized, the collaboration logic between components is complex, and the test objects span multiple physical nodes and network regions. The distributed architecture distributes test tasks to multiple distributed nodes and executes them in parallel in a remote testing environment, so as to achieve high-concurrency, high-coverage, and high-efficiency testing of large-scale software systems.
[0003] Deficiencies of the prior art: When a distributed remote testing system faces sudden failures of testing nodes or network anomalies, it generally relies on a static scheduling structure and a retry-based task recovery mechanism, and cannot achieve early perception and dynamic response to node states; once a task is interrupted during operation, it usually needs to restart the task, resulting in the loss of results of the completed part, waste of testing resources, and task semantic inconsistency problems; in addition, the current testing process generally lacks the ability to structurally manage task states, cannot achieve hot migration of tasks in the running state, and cannot effectively verify the logical consistency of tasks after migration, resulting in poor task execution stability and uncontrollable recovery process, seriously restricting the adaptability and reliability of remote testing systems in high-dynamic and high-heterogeneous environments. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, the present invention provides a software remote testing system and method based on a distributed architecture to solve the problem of low task migration execution efficiency in the software remote testing process in the above-mentioned background art.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A software remote testing method based on a distributed architecture includes the following steps:
[0007] Collect the multi-dimensional running states of remote testing nodes, construct a time series state vector and extract a state trajectory, generate an anomaly score for the current time point of the node according to the state trajectory, and judge the node state;
[0008] Remove the faulty nodes and topological edges in the abnormal state, and construct a new topology graph. Model the scheduling capabilities of the remaining nodes, construct a cost function in combination with the path delay, generate a minimum-cost path mapping, and synchronously update the global scheduling structure and routing information;
[0009] Extract the structured state of the tasks being executed in the failed nodes and encapsulate them as cross-platform snapshots; allocate target nodes according to the scheduling graph and perform state reconstruction;
[0010] Compare the original state snapshot of the node with the restored state snapshot, generate an asymmetric consistency scoring function for node state recovery determination, and determine the node execution policy based on the determination result.
[0011] In a preferred embodiment, collect the running state information of the remote test nodes, construct a time series state vector and extract the state trajectory. The specific process is as follows:
[0012] Collect the running state information of each remote test node. The running state information includes the CPU occupancy rate, memory usage rate, disk I / O saturation in the resource metrics, the average delay, packet loss rate, and bandwidth occupancy rate in the network metrics, and the average response time and error rate in the task metrics. Then fuse the obtained data to construct the original state vector;
[0013] Normalize the state vectors of different dimensions in the original state vector;
[0014] Use a sliding window to extract the state trajectory of the node within the set time. The state trajectory is a sequence composed of consecutive standardized state vectors, which is used to describe the dynamic change characteristics of the node state.
[0015] In a preferred embodiment, generate an anomaly score for the node at the current time point according to the state trajectory and judge the node state. The specific process is as follows:
[0016] Input the state trajectory into a one-dimensional convolutional network or a gated recurrent neural network to compress the temporal information in the state trajectory into a state feature vector;
[0017] According to the state feature vector, define an anomaly scoring function for the node at the current time point and determine the anomaly score value of the node;
[0018] Compare the dynamic anomaly threshold with the anomaly score value to judge the node state;
[0019] When the dynamic anomaly threshold is greater than the anomaly score value, trigger an anomaly response and mark the node state as an abnormal state.
[0020] In a preferred embodiment, faulty nodes and topological edges in abnormal states are removed, and a new topological graph is constructed. The scheduling capabilities of the remaining nodes are modeled. The specific steps are as follows:
[0021] Summarize the nodes marked as abnormal states as the abnormal node identification set;
[0022] Remove all nodes in the abnormal node identification set from the graph, and clear the incoming and outgoing edges corresponding to the nodes. The graph after removing the faulty nodes and paths is used as the new topological graph;
[0023] Conduct a capacity analysis on the remaining available nodes in the new topological graph, construct the current resource adaptation map, and determine the scheduling capacity of each node as: , where represents the current available resources; represents the total amount of resources of this node; is a numerical stability constant to avoid division by zero; represents the relative resource margin of the current node.
[0024] In a preferred embodiment, a cost function is constructed in combination with the path delay to generate the minimum-cost path mapping, and the global scheduling structure and routing information are synchronously updated. The specific process is as follows:
[0025] Define the set of test tasks to be scheduled as , where m is the total number of test tasks;
[0026] Construct a task scheduling mapping graph , where represents a mapping relationship from the test task to the target execution node;
[0027] Set the path selection based on the optimization criterion. The total cost of each feasible path from the scheduling center to the node is: , where is the path delay (including link transmission time); is the node scheduling capacity index; , is the weight parameter, and the values are all greater than 0;
[0028] By traversing all paths, select the target node with the minimum total cost for each task: represents the k-th test task in the set of test tasks to be scheduled ;
[0029] After all tasks are mapped, update the global scheduling graph structure, including updating the task routing table and topological cache graph inside the scheduler.
[0030] In a preferred embodiment, the structured states of the tasks being executed in the failed node are extracted and encapsulated as a cross-platform snapshot; the target node is allocated according to the scheduling graph and the state is reconstructed. The specific process is as follows:
[0031] When the scheduler receives the notification of node exception, it immediately freezes all the active tasks on it and records them as the active set , where m is the total number of active tasks;
[0032] Extract the internal state of the task status of each active task and generate a snapshot tuple;
[0033] The content of the snapshot tuple includes: the position of the script execution pointer, the set of internal variables, the environment binding tuple, the description of the temporary file handle and the file buffer structure, and the summary of the network session status;
[0034] Convert the snapshot tuple and encapsulate it into a cross-platform expression form;
[0035] The encapsulation process includes performing a dependency mapping transformation on the environmental dependencies, renaming the incompatible component paths, serializing the set of internal variables into a platform-neutral format; converting the network status summary into a replayable connection status description; establishing an offset mapping for the position of the script pointer;
[0036] After completing the snapshot encapsulation, the scheduler maps each task to be restored to the target execution node and obtains the final task status recovery snapshot after the migration is completed.
[0037] In a preferred embodiment, the original state snapshot of the node is compared with the recovery state snapshot to generate an asymmetric consistency scoring function for node state recovery determination. The specific steps are as follows:
[0038] Construct a state space mapping function to unify the task status recovery snapshot and the snapshot tuple into the same semantic space;
[0039] Construct an asymmetric consistency residual score for the hot migration scenario of the target node. The expression of the asymmetric consistency residual score is: , where d is the dimension of the state vector; 、 are the feature components of the original state and the migration state at the k-th item; is the call stack similarity at the k-th item; is the residual of the call stack similarity at the k-th item.
[0040] In a preferred embodiment, and based on the determination result, determine the node execution policy. The specific steps are as follows:
[0041] Define a risk tolerance threshold and compare it with the asymmetric consistency residual score;
[0042] If the risk tolerance threshold is less than or equal to the asymmetric consistency residual score, it is determined that the current state is valid, the recovery process is successful, and the task continues to be executed on the target test node;
[0043] If the risk tolerance threshold is greater than the asymmetric consistency residual score, the node state recovery is not fully credible, there is a potential error risk, and the target test node task rollback mechanism is started.
[0044] A software remote test system based on a distributed architecture, used to implement the above-mentioned software remote test method based on a distributed architecture, including:
[0045] A node data acquisition module, used to collect the multi-dimensional operating status of remote test nodes, construct a time series state vector and extract the state trajectory, generate an anomaly score for the current time point of the node according to the state trajectory, and judge the node state;
[0046] A node path update module, used to eliminate faulty nodes and topological edges in the abnormal state, construct a new topological graph, model the scheduling capabilities of the remaining nodes, construct a cost function in combination with the path delay, generate a minimum cost path mapping, and synchronously update the global scheduling structure and routing information;
[0047] A scheduling graph reconstruction module, used to extract the structured state of the tasks being executed in the failed nodes and encapsulate them into cross-platform snapshots; allocate target nodes according to the scheduling graph and perform state reconstruction;
[0048] A node state recovery analysis module, used to compare the original state snapshot of the node with the recovered state snapshot, generate an asymmetric consistency scoring function to determine the node state recovery, and determine the node execution strategy based on the determination result.
[0049] The technical effects and advantages of the present invention:
[0050] The present invention realizes the accurate modeling and risk prediction of the operating status of remote test nodes by constructing a node anomaly detection mechanism based on state trajectory analysis, combines topology pruning and resource capacity evaluation, dynamically generates the optimal scheduling path and updates the task mapping structure, effectively avoiding the scheduling risks brought by unstable nodes; during the task execution process, extracts the structured state snapshot of the interrupted task and performs platform-independent encapsulation to realize the seamless migration and state reconstruction of the task on the target node; further constructs an asymmetric residual scoring model by comparing the semantic consistency of the original state and the recovered state, quantifies the credibility of the state recovery, and dynamically determines whether the task continues to execute or enters the rollback mechanism based on this, thereby realizing the state-aware intelligent scheduling and hot migration control of test tasks in a distributed environment, and significantly improving the stability, scheduling continuity and logical consistency of task execution of the system in the node fluctuation scenario. Description of the Drawings
[0051] Figure 1 This is a flowchart of a software remote testing method based on a distributed architecture according to the present invention.
[0052] Figure 2 This is a schematic structural diagram of a software remote testing system based on a distributed architecture according to the present invention. Specific embodiments
[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0054] Embodiment 1: As Figure 1 shown, a software remote testing method based on a distributed architecture includes the following steps:
[0055] Collect the multi-dimensional operating status of remote test nodes, construct a time series state vector and extract the state trajectory, generate an anomaly score for the current time point of the node according to the state trajectory, and judge the node status;
[0056] Eliminate the faulty nodes and topological edges in the abnormal state, and construct a new topological graph. Model the scheduling ability of the remaining nodes, combine the path delay to construct a cost function, generate a minimum cost path mapping, and synchronously update the global scheduling structure and routing information;
[0057] Extract the structured state of the tasks being executed in the failed nodes and encapsulate them as cross-platform snapshots; allocate target nodes according to the scheduling graph and perform state reconstruction;
[0058] Compare the original state snapshot of the node with the restored state snapshot, generate an asymmetric consistency scoring function for node state recovery determination, and determine the node execution strategy based on the determination result.
[0059] During the distributed software remote testing process, the test tasks are distributed to remote test nodes for execution. Since these nodes are in heterogeneous networks and inconsistent resource environments, it is easy to have situations such as resource overload, link degradation, or execution anomalies, resulting in test interruption or delay;
[0060] Step 1, perform state perception and anomaly detection of remote test nodes, that is, by collecting multi-dimensional state data, constructing a time series model to describe the behavior evolution of the nodes, calculating the anomaly score value, and identifying potential faulty nodes in advance, so as to output a trigger signal for topological reconstruction and task migration. The specific steps are as follows:
[0061] Periodically collect the running status of each node before task scheduling, identify potential abnormal trends by constructing a state evolution model, and make predictions before the node status deteriorates;
[0062] First, collect the running status information of each remote test node at time t. The running status information includes resource metrics (such as CPU occupancy rate, memory usage rate, disk I / O saturation), network metrics (such as average latency, packet loss rate, bandwidth occupancy rate), and task metrics (such as average response time, error rate, etc.). The obtained data is fused to construct an original state vector;
[0063] Normalize the state vectors of different dimensions in the original state vector, perform scale unification processing on states with different dimensions, and obtain a unified state vector after standardization;
[0064] Introduce a sliding window mechanism to evolve the state trajectory of the node in a recent period of time. This trajectory is a sequence H(t) composed of continuous standardized state vectors, which is used to describe the dynamic change characteristics of the node status;
[0065] Specifically, use a sliding window to construct the state trajectory for the standardized state vector. Let the window length be , then the state trajectory is: , at this time, the state trajectory represents the multi-dimensional behavior evolution of the node in the recent time instants. Using this trajectory, trend characteristics (such as continuous deterioration or fluctuating instability) can be judged. R represents the set of real numbers, that is, the set composed of all real numbers;
[0066] By analyzing this trajectory, determine whether the node is in an unstable state such as gradually depleting resources, frequent network jitter, or continuous decline in response ability. The state trajectory is then input into an embedding encoding model, such as a lightweight one-dimensional convolutional network or a gated recurrent neural network (such as GRU), so as to compress this multi-dimensional time series information into a state feature vector z(t) for the next step of risk assessment;
[0067] According to the state feature vector z(t), define the abnormal scoring function of the node at the current time point: , where w is a learnable feature weight vector, T represents the transpose operation, The function is used to enhance the pulling effect of the abnormal trend on the final score; b is a bias term to adjust the center point of the score;
[0068] Calculate an abnormal score value according to this state feature vector , used to quantify whether the current node is in a potentially high-risk state. This scoring value does not depend on the absolute value of any single metric, but is learned from data in multiple dimensions and at multiple time points. Therefore, it has better anti-jitter performance and generalization ability. The higher the anomaly score, the greater the degree to which the node deviates from the normal state at the current time point;
[0069] Generate a dynamic anomaly threshold based on the historical fluctuation of the score , which is calculated from the score sequence in a recent period of time, such as taking its 95th percentile value and adding a fluctuation tolerance factor;
[0070] Maintain a set of historical score sequences , and based on its change trend and the relationship between the current value and the historical fluctuation range, determine whether to trigger an anomaly response. Let the judgment function be: , where is the dynamic anomaly threshold, which is composed of the upper quantile of the historical score plus a tolerance factor, and is used to dynamically adapt to the state characteristics of different nodes;
[0071] If the trigger value is 1, immediately mark the node state as an abnormal state and output the following content: node identifier and anomaly mark, current state trajectory snapshot H(t), anomaly score value .
[0072] Step 2, distributed topology reconstruction and scheduling graph update. After completing the node state perception and anomaly detection in Step 1, one or more test nodes with abnormal running states will be marked. To ensure that tasks are not distributed to unstable nodes and maintain the high availability and load balance of the system, it is necessary to immediately start the topology reconstruction mechanism to adjust the original node connection structure and task mapping path;
[0073] Not be overly sensitive during periods when the node states generally fluctuate greatly, and be able to respond more promptly when the states are relatively stable. Once the current anomaly score value exceeds this dynamic anomaly threshold, it is considered that the node is in an abnormal state. The system immediately generates an anomaly trigger signal, removes the node from the schedulable pool, and marks the anomaly;
[0074] After receiving the anomaly node detection results from Step 1, the system first needs to remove these no-longer-safe nodes from the current scheduling topology graph. The scheduling graph is the basic structure of all task distribution logics. If the abnormal nodes are not cleared in time, it will cause test tasks to still be wrongly distributed to faulty nodes, resulting in incorrect responses or data loss;
[0075] The abnormal node identification set output from step 1 is recorded as F, which represents the set of nodes currently judged to be unschedulable. The original scheduling topology can be represented as a graph structure G = (V, E), where V is the set of all nodes, E is the set of connecting edges, and edges represent routable paths for tasks.
[0076] Remove all nodes in the abnormal node identification set F from the graph and clear their corresponding topological edges (incoming and outgoing edges) to generate a new pruned topological graph. ,in, New topology The node set is composed of the original set V after deleting the abnormal node set F, that is, retaining all healthy and schedulable nodes; New topology The edge set represents the result after deleting all edges associated with any abnormal node in the original edge set E; Represents an edge in the original graph; Indicates that at least one of the two endpoints connected by the edge belongs to the abnormal node set F;
[0077] Eliminate failed nodes and related paths to ensure that subsequent tasks are scheduled on stable nodes, avoiding task loss caused by scheduling paths passing through unavailable nodes.
[0078] It should be noted that this involves not only the physical removal of nodes but also the updating of their logical connectivity. This is particularly important when pruning edges in the graph. For example, if the communication path between two healthy nodes relies on a relay node, and that relay node is abnormal, the connection between the two healthy nodes should also be considered broken. This strategy ensures that subsequent path planning is based entirely on trusted connections.
[0079] After completing structural elimination, task path reconstruction cannot be performed immediately. This is because, although the health of the remaining nodes in the system meets the basic scheduling conditions, their carrying capacity is uneven. Some nodes may have sufficient idle resources, while others are approaching resource bottlenecks. If tasks are dispatched without considering this, the local system load will be increased, reducing overall execution efficiency, and even triggering new congestion or node overload.
[0080] Therefore, we analyze the capabilities of the remaining available nodes in the new topology, build the current resource adaptation map, and define each node The scheduling capability is: , where Indicates currently available resources (such as the number of remaining CPU cores, free memory, etc.); Indicates the total amount of resources of the node; is a numerically stable constant to avoid division by zero; Indicates the relative resource margin of the current node.
[0081] The resource adaptation map is used for subsequent task distribution and path weight assignment, and is the input basis for optimizing the scheduling path.
[0082] Perform resource capacity modeling for each remaining node, extract information such as its computing resource availability, remaining task processing space, and concurrent support strength at the current moment, and quantify it into a comparable scheduling capacity scoring metric, the relative resource margin The smaller it is, the more crowded the node is and the less suitable it is to continue to undertake tasks; on the contrary, a high relative resource margin indicates that the node has abundant resources and strong response capabilities.
[0083] Perform scheduling path reconstruction and minimum-cost mapping generation. After node capacity evaluation, enter the most critical reconstruction stage. At this time, the scheduler needs to select an optimal path for each task in the current set of tasks to be scheduled in the trimmed topology graph, pointing to a suitable test node. The specific steps are as follows:
[0084] Define the set of test tasks to be scheduled as , where m is the total number of test tasks, and the goal is to reconstruct the task assignment path on the trimmed new topology graph;
[0085] Construct a task scheduling mapping graph , where represents a mapping relationship from test tasks to target execution nodes, that is, a function that assigns each task to which node for execution (for example, represents the mapping function from tasks to nodes), and path selection is based on the following optimization criteria;
[0086] Let the total cost of each feasible path from the scheduling center to node be: , where is the path delay (including link transmission time); is the node scheduling capacity metric; , is the weight parameter, and its value is greater than 0, which is used to balance path transmission efficiency and resource adaptability;
[0087] By traversing all paths, select the target node with the minimum total cost for each task: represents the set of test tasks to be scheduled The kth test task in
[0088] Perform task graph synchronization update and topology cache replacement, that is, when all task mappings are completed, update the global scheduling graph structure. This update includes two parts:
[0089] Update the task routing table inside the scheduler , indicating the path along which the task should be sent to the new target node;
[0090] Update the topological cache graph of the system, replace the original graph with the updated new topological graph and use it as the scheduling graph for node testing, and record the adjustment timestamp to support topological rollback after the abnormal node recovers;
[0091] In addition, a path adjustment notice can be generated, including: the task identifier of the reconstructed path, the node information before and after scheduling, and the details of the path cost change.
[0092] It should be noted that the task routing table refers to a task-level mapping structure maintained inside the scheduler, which is used to record the target execution node corresponding to each test task currently and the scheduling path information it needs to pass through. It is one of the core data structures for the scheduler to perform task distribution and status tracking. The task routing table records the task-node mapping relationship, stores the scheduling path, and tracks the task execution status, etc.; the topological cache graph refers to a copy of the graph structure maintained in the system scheduler, which is used to save the state snapshot of the effective scheduling graph in the current test system and serves as the working graph or intermediate graph version for scheduling execution, so as to support the system to retain the topological historical state and assist in restoring scheduling consistency after operations such as node state change, path reconstruction, and task migration.
[0093] In summary, this step integrates the three factors of task characteristics, path constraints, and node status, and realizes the optimal matching of scheduling performance on the premise of ensuring availability. Therefore, when some nodes in the test environment fail, the system can quickly complete topological update, task remapping, and path optimization, and keep the overall scheduling uninterrupted.
[0094] Step 3, perform test task status snapshot and hot migration execution. After completing the topological reconstruction and task scheduling graph update in Step 2, the scheduler has completed the path reconstruction and target node mapping of the tasks on the failed nodes. However, for those test tasks that are already running on the faulty nodes, if they are directly aborted and then re-executed, it will cause unnecessary test repetition, data loss, and context confusion problems. Therefore, a task status snapshot extraction and hot migration mechanism is designed to ensure that the task can seamlessly continue to execute on the new target node after the task running is interrupted. The specific steps are as follows:
[0095] When the scheduler receives the notice of node abnormality, immediately freeze all the active task sets on it, denoted as the active set: , where m is the total number of active tasks. Each active task is being executed, and its internal state needs to be extracted as a structured snapshot;
[0096] The content of the task status snapshot (snapshot tuple) includes but is not limited to:
[0097] The position of the script execution pointer : Record the current test execution position;
[0098] Set of internal variables : Include the values of context variables during execution;
[0099] Environment binding tuple : Represent the execution environment of task dependencies;
[0100] Description of temporary file handle and file buffer structure (if any); (if any);
[0101] Summary of network session status : Describe whether the communication session is being maintained.
[0102] Extracting these running states will cause the task to be forced to terminate and reset, losing its original execution continuity. Therefore, the task status needs to be quickly extracted before the node exits (or through an asynchronous log compensation mechanism), and the above status information is organized into a snapshot tuple: .
[0103] After extracting the snapshot tuple, it cannot be directly transmitted to the target node. The reason is that there may be differences in the underlying running platforms of different test nodes (such as operating system versions, test toolchains, library dependency paths, etc.). If the format and semantics of the snapshot content are not uniformly converted, it may cause the failure of state loading after migration or task logic deviation;
[0104] Convert the snapshot tuple into a cross-platform neutral expression form, and define a snapshot encapsulation function: , where is the encapsulation function, is the snapshot tuple;
[0105] The encapsulation process of the encapsulation function includes: performing dependency mapping conversion on environment dependencies, renaming incompatible component paths, serializing the set of internal variables into a platform-neutral format; converting the network status summary into a replayable connection status description to support session reconstruction; establishing an offset mapping for the script pointer position to make it consistent with the target node's parsing.
[0106] After completing the snapshot encapsulation, the scheduler maps each task to be restored to its target execution node according to the scheduling graph constructed in step 2, and the target execution receives the encapsulated snapshot , starts the task recovery process in the container isolation environment, thereby eliminating the task termination problem caused by node failure, and finally the task status restores the snapshot is regarded as equivalent to the logical state of the original task at the time of interruption and can continue to execute on the target node.
[0107] During the cross-node reconstruction of the task status, recovery distortion or context loss may occur due to reasons such as platform environment, serialization deviation, and status mapping conflicts. Therefore, before the task starts to execute again, a status consistency check is performed.
[0108] Step 4, perform a consistency verification and recovery confirmation mechanism, and the specific steps are as follows:
[0109] Receive the task status recovery snapshot after the migration in Step 3 and at the same time retain the snapshot tuple of the original migration The structural descriptions in both are structured state vectors, containing multiple heterogeneous fields such as variables, execution pointers, environment dependencies, session information, etc.;
[0110] Construct a state space mapping function to unify the two sets of state structures into the same semantic space S for subsequent consistency comparison: where, represents the semantic vector of the original state, represents the semantic reconstruction vector of the target state;
[0111] This state space mapping function should satisfy two basic properties: each logical field in the state is semantically comparable; it can adapt to the running data structures in different system platforms (such as variable names, library versions, system description information, etc.).
[0112] Construct an asymmetric consistency residual score for the hot migration scenario of the target node, allowing for slight differences in certain dimensions of the state, but maintaining strict consistency for key variables or path structures. Define the asymmetric consistency residual score as follows: where d is the dimension of the state vector (i.e., the number of state components); 、 are the feature components of the original state and the migrated state at the k-th item; is the state difference measure function for the k-th item (such as structural edit distance, string matching difference, call stack similarity, etc.); is the residual for the k-th item.
[0113] Asymmetric consistency residual score The lower the value of, the closer the state recovery is to the original state.
[0114] After completing the multi-dimensional semantic comparison and consistency residual calculation of the task status, the system needs to make a comprehensive determination of the current recovery state of the hot migration task, so as to decide whether to allow the task to continue executing on the target node, or whether to roll back the task status to the original snapshot or an alternative migration path;
[0115] Define the risk tolerance threshold and compare it with the asymmetric consistency residual score to determine the sensitivity of the current task to consistency, so as to judge whether the current recovery state is trustworthy enough. For example, interface test tasks have extremely high requirements for session recovery, while stateless script tasks can tolerate a certain degree of execution deviation;
[0116] If the risk tolerance threshold is less than or equal to the asymmetric consistency residual score, the system determines that the current state is valid, the recovery process is successful, and the task can continue to be executed on the target test node; at the same time, record the successful migration path and state mapping parameters into the task scheduler log for subsequent migration strategy optimization;
[0117] If the risk tolerance threshold is greater than the asymmetric consistency residual score, the system will consider that the state recovery is not completely trustworthy and there is a potential error risk. At this time, the task rollback mechanism of the target test node will be started, which generally includes the following processing paths: reschedule the task to a pre-configured backup node, reload the snapshot and execute the recovery again; or in the scenario where there is no backup resource, mark the task as migration failed, change it to the failed processing state, and wait for manual intervention or the system fault tolerance strategy to intervene.
[0118] It should be noted that the thresholds involved in the embodiments are set according to specific scenarios and requirements.
[0119] The present invention realizes accurate modeling and risk prediction of the running state of remote test nodes by constructing a node anomaly detection mechanism based on state trajectory analysis, combines topology pruning and resource capacity evaluation, dynamically generates the optimal scheduling path and updates the task mapping structure, effectively avoiding the scheduling risks brought by unstable nodes; during the task execution process, extract the structured state snapshot of the interrupted task and perform platform-independent encapsulation to realize the seamless migration and state reconstruction of the task on the target node; further, by comparing the semantic consistency between the original state and the restored state, construct an asymmetric residual score model to quantify the credibility of state recovery, and accordingly dynamically determine whether the task continues to execute or enters the rollback mechanism, thus realizing intelligent scheduling and hot migration control based on state awareness of test tasks in a distributed environment, significantly improving the stability, scheduling continuity and logical consistency of task execution of the system in the node fluctuation scenario.
[0120] Embodiment 2: A software remote test system based on a distributed architecture, as Figure 2 shown, specifically including:
[0121] A node data acquisition module, used to collect the multi-dimensional running state of remote test nodes, construct a time series state vector and extract the state trajectory, generate an anomaly score for the current time point of the node according to the state trajectory, and judge the node state;
[0122] The node path update module is used to eliminate faulty nodes and topological edges in abnormal states, construct a new topological graph, model the scheduling capabilities of the remaining nodes, construct a cost function in combination with path delay, generate a minimum-cost path mapping, and synchronously update the global scheduling structure and routing information;
[0123] The scheduling graph reconstruction module is used to extract the structured state of the tasks being executed in the failed nodes and encapsulate them as cross-platform snapshots; allocate target nodes according to the scheduling graph and perform state reconstruction;
[0124] The node state recovery analysis module is used to compare the original state snapshot of the node with the recovered state snapshot, generate an asymmetric consistency scoring function for node state recovery determination, and determine the node execution strategy based on the determination result.
[0125] The above formulas are all calculated by taking the numerical values without dimensions. Specific dimension elimination can be achieved by various means such as standardization, which will not be elaborated here. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0126] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more collections of available media. The available media can be magnetic media (such as floppy disks, ATA hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state ATA hard disk.
[0127] It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0128] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0129] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0130] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0131] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0132] As mentioned above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A software remote testing method based on a distributed architecture, characterized in that, It includes the following steps: Collect the multi-dimensional operating status of remote test nodes, construct a time-series state vector and extract the state trajectory, generate an anomaly score for the node at the current time point according to the state trajectory, and judge the node status; Eliminate the faulty nodes and topological edges in the abnormal state, construct a new topology graph, model the scheduling capabilities of the remaining nodes, combine the path delay to construct a cost function, generate a minimum-cost path mapping, and synchronously update the global scheduling structure and routing information; Extract the structured state of the tasks being executed in the failed nodes and encapsulate them as cross-platform snapshots; allocate target nodes according to the scheduling graph and perform state reconstruction; Compare the original state snapshot of the node with the restored state snapshot, generate an asymmetric consistency scoring function for node state recovery determination, and determine the node execution strategy based on the determination result; Compare the original state snapshot of the node with the restored state snapshot, generate an asymmetric consistency scoring function for node state recovery determination, and the specific steps are as follows: Construct a state space mapping function to unify the task state recovery snapshot and the snapshot tuple into the same semantic space; Construct an asymmetric consistency residual score for the hot migration scenario of the target node. The expression for the asymmetric consistency residual score is: , where d is the dimension of the state vector; , are the feature components of the original state and the migrated state at the k-th item; is the call stack similarity at the k-th item; is the residual of the call stack similarity at the k-th item; And determine the node execution strategy based on the determination result, and the specific steps are as follows: Define a risk tolerance threshold for comparison with the asymmetric consistency residual score; If the risk tolerance threshold is less than or equal to the asymmetric consistency residual score, it is determined that the current state is valid, the recovery process is successful, and the task continues to be executed on the target test node; If the risk tolerance threshold is greater than the asymmetric consistency residual score, the node state recovery is not completely credible, there is a potential error risk, and the task rollback mechanism of the target test node is started.
2. The software remote testing method based on a distributed architecture according to claim 1, wherein: Collect the operating status information of remote test nodes, construct a time-series state vector and extract the state trajectory, and the specific process is as follows: Collect the operating status information of each remote test node. The operating status information includes the CPU occupancy rate, memory usage rate, disk I / O saturation rate in the resource metrics, the average delay, packet loss rate, and bandwidth occupancy rate in the network metrics, and the average response time and error rate in the task metrics, and fuse the obtained data to construct an original state vector; Normalize the state vectors of different dimensions in the original state vector; Use a sliding window to extract the state trajectory of the node within the set time. The state trajectory is a sequence composed of continuous normalized state vectors, which is used to describe the dynamic change characteristics of the node state.
3. A software remote testing method based on a distributed architecture according to claim 2, characterized in that: Generate an anomaly score for the node at the current time point according to the state trajectory, and judge the node status, and the specific process is as follows: Input the state trajectory into a one-dimensional convolutional network or a gated recurrent neural network to compress the temporal information in the state trajectory into a state feature vector; According to the state feature vector, define an anomaly scoring function for the node at the current time point and determine the anomaly score value of the node; Compare the dynamic anomaly threshold with the anomaly score value to judge the node status; When the dynamic anomaly threshold is greater than the anomaly score value, trigger an anomaly response and mark the node status as an abnormal state.
4. A software remote testing method based on a distributed architecture according to claim 3, characterized in that: Eliminate the faulty nodes and topological edges in the abnormal state, construct a new topology graph, and model the scheduling capabilities of the remaining nodes. The specific steps are as follows: Summarize the nodes marked as abnormal states as an abnormal node identification set; Remove all nodes in the abnormal node identification set from the graph, clear the incoming and outgoing edges corresponding to the nodes, and use the removed failed nodes and path graph as the new topology graph; Perform a capability analysis on the remaining available nodes in the new topology graph, construct the current resource adaptation map, and determine each node 's scheduling capability as: , where represents the currently available resources; represents the total node resources; is a numerical stability constant to avoid division by zero; represents the relative resource margin of the current node, is the node set of the new topology graph.
5. A software remote testing method based on a distributed architecture according to claim 4, characterized in that: Construct a cost function in combination with path delay, generate a minimum cost path mapping, and synchronously update the global scheduling structure and routing information. The specific process is as follows: Define the set of test tasks to be scheduled as , where m is the total number of test tasks; Construct a task scheduling mapping graph , where represents a mapping relationship from a test task to a target execution node; The path selection is set based on an optimization criterion, and the total cost of each feasible path from the dispatching center to the node is: , where is the path delay (including the link transmission time); is the node scheduling capacity index; , is the weight parameter, and the values are all greater than 0; By traversing all paths, select the target node with the minimum total cost for each task: , represents the k-th test task in the set of test tasks to be scheduled ; After all tasks are mapped, update the global scheduling graph structure, including updating the task routing table and topology cache graph inside the scheduler.
6. A software remote testing method based on a distributed architecture according to claim 5, characterized in that: Extract the structured state of the tasks being executed in the failed nodes and encapsulate them into cross-platform snapshots; allocate target nodes according to the scheduling graph and perform state reconstruction. The specific process is as follows: When the scheduler receives a notification of node exception, it immediately freezes all active tasks on it and records them as the active set , where h is the total number of active tasks; Extract the internal state of the task status of each active task and generate a snapshot tuple; The content of the snapshot tuple includes: the position of the script execution pointer, the set of internal variables, the environment binding tuple, the description of the temporary file handle and file buffer structure, and the summary of the network session status; Convert and encapsulate the snapshot tuple into a cross-platform representation form; The encapsulation process includes performing a dependency mapping conversion on the environment dependencies, renaming incompatible component paths, serializing the set of internal variables into a platform-neutral format; converting the network status summary into a replayed connection status description; establishing an offset mapping for the script pointer position; After completing the snapshot encapsulation, the scheduler maps each task to be restored to the target execution node and obtains the final task status recovery snapshot after the migration is completed.
7. A software remote testing system based on a distributed architecture, which is used to implement a software remote testing method based on a distributed architecture according to any one of claims 1-6, characterized in that, Including: A node data collection module, which is used to collect the multi-dimensional operating status of remote test nodes, construct a time series status vector and extract the status trajectory, generate an abnormal score for the current time point of the node according to the status trajectory, and judge the node status; A node path update module, which is used to remove the failed nodes and topology edges in the abnormal state, construct a new topology graph, model the scheduling capabilities of the remaining nodes, construct a cost function in combination with path delay, generate a minimum cost path mapping, and synchronously update the global scheduling structure and routing information; A scheduling graph reconstruction module, which is used to extract the structured state of the tasks being executed in the failed nodes and encapsulate them into cross-platform snapshots; allocate target nodes according to the scheduling graph and perform state reconstruction; A node status recovery analysis module, which is used to compare the original state snapshot of the node with the restored state snapshot, generate an asymmetric consistency scoring function for node status recovery determination, and determine the node execution strategy based on the determination result.
Citation Information
Patent Citations
Distributed automatic test vector generation method, device and system capable of dynamically expanding capacity
CN117851107A
System and method for using an automated process to identify bugs in software source code
US20050223357A1