Software remote testing system and method based on distributed architecture

By building an exception detection mechanism and topological clipping based on state trajectory, the optimal scheduling path is dynamically generated, and the stability and task execution consistency of distributed remote testing systems in node fluctuations scenarios are achieved, task interruption and resource waste caused by node failures in the existing technology is solved, and the system's adaptability and reliability are improved.

CN120234253AActive Publication Date: 2025-07-01SHENZHEN SOFT ALLIANCE TECH SERVICE CO LTD

Patent Information

Application Number
CN202510704115.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-01
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

When a distributed remote testing system faces a sudden failure of the test node or network exception, it cannot realize early perception and dynamic response to the node state, resulting in task interruption, loss of results, wasted resources and inconsistent task semantics, poor test execution stability, insufficient adaptability and reliability.

Method used

By collecting the multi-dimensional running state of remote test nodes, building time series state vectors and state trajectories, generating exception scores, eliminating faulty nodes, building a new topology diagram and updating the scheduling structure, extracting task status snapshots and performing cross-platform encapsulation, combining consistency scoring functions to perform task recovery judgments, and achieving seamless migration and state reconstruction of tasks.

Benefits of technology

It realizes accurate modeling and risk prediction of remote test nodes, dynamically generates the optimal scheduling path, improves the stability and scheduling continuity of distributed test systems in node fluctuations scenarios, and ensures the logical consistency of task execution and efficient operation of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234253A_ABST
    Figure CN120234253A_ABST
Patent Text Reader

Abstract

The invention discloses a software remote testing system and method based on a distributed architecture, particularly relates to the technical field of software testing, and is used for solving the problem of low task migration execution efficiency in a software remote testing process. According to the method, an anomaly detection mechanism is constructed based on a node state track, potential faults of remote test nodes are accurately identified, modeling is performed in combination with topology cutting and resource capability, an optimal task scheduling path is dynamically generated, and a scheduling structure is updated; structured snapshots are extracted from running tasks in the failure nodes, cross-platform packaging is carried out, state reconstruction is completed at target nodes, and uninterrupted task migration is achieved; by comparing the original state with the recovery state, an asymmetric residual scoring function is constructed, the consistency degree is quantified, and whether the task continues to be executed or rolls back is dynamically judged, so that the scheduling stability and the execution continuity of the distributed test under the fluctuation of the test node are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of software testing, and more specifically, to a software remote testing system and method based on a distributed architecture. Background Art

[0002] In the development and operation and maintenance process of modern large-scale software systems, the complexity and coverage of software testing are constantly expanding. The traditional testing mode in a single-machine or local network environment has been difficult to meet the testing requirements of distributed, heterogeneous, and highly real-time systems. With the rapid development of cloud computing, edge computing, and microservice architectures, the deployment environment of software systems tends to be decentralized, the collaboration logic between components is complex, the test objects span multiple physical nodes and network regions. The distributed architecture distributes test tasks to multiple distributed nodes and executes them in parallel in a remote testing environment, so as to achieve high-concurrency, high-coverage, and high-efficiency testing of large-scale software systems.

[0003] Deficiencies of the prior art: When the distributed remote testing system faces sudden failures of test nodes or network anomalies, it generally relies on a static scheduling structure and a retry-based task recovery mechanism, and cannot achieve early perception and dynamic response to node states; once the task is interrupted during operation, it usually needs to restart the task, resulting in the loss of results of the completed part, waste of test resources, and task semantic inconsistency problems; in addition, the current testing process generally lacks the ability to structurally manage task states, cannot achieve hot migration of tasks in the running state, and cannot effectively verify the logical consistency of tasks after migration, resulting in poor task execution stability and uncontrollable recovery process, seriously restricting the adaptability and reliability of the remote testing system in a highly dynamic and highly heterogeneous environment. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, the present invention provides a software remote testing system and method based on a distributed architecture to solve the problem of low task migration execution efficiency in the software remote testing process in the above-mentioned background art.

[0005] To achieve the above object, the present invention provides the following technical solutions: A software remote testing method based on a distributed architecture, comprising the following steps: Collect multi-dimensional operating states of remote testing nodes, construct a time series state vector and extract a state trajectory, generate an anomaly score for the current time point of the node according to the state trajectory, and judge the node state; Eliminate faulty nodes and topological edges in the abnormal state, and construct a new topological graph, model the scheduling ability of the remaining nodes, construct a cost function in combination with path delay, generate a minimum cost path mapping, and synchronously update the global scheduling structure and routing information; Extract the structured status of the tasks being executed in the failed nodes and encapsulate it into a cross-platform snapshot; allocate target nodes according to the scheduling graph and perform status reconstruction; Compare the original node status snapshot with the restored status snapshot, generate an asymmetric consistency scoring function to determine node status recovery, and determine the node execution policy based on the determination result.

[0006] In a preferred embodiment, collect the running status information of remote test nodes, construct a time series status vector and extract the status trajectory. The specific process is as follows: Collect the running status information of each remote test node. The running status information includes the CPU occupancy rate, memory usage rate, disk I / O saturation rate in resource metrics, the average latency, packet loss rate, and bandwidth occupancy rate in network metrics, as well as the average response time and error rate in task metrics. And fuse the obtained data to construct the original status vector; Normalize the status vectors of different dimensions in the original status vector; Use a sliding window to extract the status trajectory of the node within the set time. The status trajectory is a sequence composed of continuous standardized status vectors, which is used to describe the dynamic change characteristics of the node status.

[0007] In a preferred embodiment, generate the anomaly score of the node at the current time point according to the status trajectory and judge the node status. The specific process is as follows: Input the status trajectory into a one-dimensional convolutional network or a gated recurrent neural network to compress the temporal information in the status trajectory into a status feature vector; According to the status feature vector, define the anomaly scoring function of the node at the current time point and determine the anomaly score value of the node; Compare the dynamic anomaly threshold with the anomaly score value to judge the node status; When the dynamic anomaly threshold is greater than the anomaly score value, trigger an anomaly response and mark the node status as an abnormal state.

[0008] In a preferred embodiment, remove the faulty nodes and topological edges in the abnormal state, and construct a new topological graph. Model the scheduling capabilities of the remaining nodes. The specific steps are as follows: Summarize the nodes marked as abnormal states as the abnormal node identification set; Remove all the nodes in the abnormal node identification set from the graph, clear the incoming and outgoing edges corresponding to the nodes, and use the graph after removing the failed nodes and paths as the new topological graph; Perform a capacity analysis on the remaining available nodes in the new topological graph, construct the current resource adaptation map, and determine the scheduling capacity of each node as: , where Indicates the currently available resources; Indicates the total amount of resources of this node; Is a numerical stability constant to avoid division by zero; Indicates the relative resource margin of the current node.

[0009] In a preferred embodiment, a cost function is constructed in combination with path delay to generate a minimum-cost path mapping, and the global scheduling structure and routing information are synchronously updated. The specific process is as follows: Define the set of test tasks to be scheduled as , where m is the total number of test tasks; Construct a task scheduling mapping graph , where Represents a mapping relationship from a test task to a target execution node; Set the path selection based on the optimization criterion, and the total cost of each feasible path from the scheduling center to the node is: , where Is the path delay (including link transmission time); Is the node scheduling ability index; , Are weight parameters, and their values are all greater than 0; By traversing all paths, select the target node with the minimum total cost for each task: Represents the set of test tasks to be scheduled The kth test task in; After all tasks are mapped, update the global scheduling graph structure, including updating the task routing table and topology cache graph inside the scheduler.

[0010] In a preferred embodiment, extract the structured state of the tasks being executed in the failed nodes and encapsulate them into a cross-platform snapshot; allocate target nodes according to the scheduling graph and perform state reconstruction. The specific process is as follows: When the scheduler receives a notification of node exception, immediately freeze all active tasks on it and record them as the active set , where m is the total number of active tasks; Extract the internal state of the task status of each active task and generate a snapshot tuple; The content of the snapshot tuple includes: the position of the script execution pointer, the set of internal variables, the environment binding tuple, the description of the temporary file handle and the file buffer structure, and the summary of the network session status; Convert the snapshot tuple and encapsulate it into a cross-platform expression form; The encapsulation process includes performing dependency mapping conversion on environmental dependencies, renaming incompatible component paths, serializing the internal variable set into a platform-neutral format; converting the network state summary into a replayable connection state description; establishing an offset mapping for the script pointer position; After completing the snapshot encapsulation, the scheduler maps each task to be restored to the target execution node and obtains the final task state recovery snapshot after migration is completed.

[0011] In a preferred embodiment, the original state snapshot of the node is compared with the recovery state snapshot to generate an asymmetric consistency scoring function for node state recovery determination. The specific steps are as follows: Construct a state space mapping function to unify the task state recovery snapshot and the snapshot tuple into the same semantic space; Construct an asymmetric consistency residual score for the live migration scenario of the target node. The expression of the asymmetric consistency residual score is: where d is the dimension of the state vector; 、 are the feature components of the original state and the migration state at the k-th item; is the call stack similarity at the k-th item; is the residual of the call stack similarity at the k-th item.

[0012] In a preferred embodiment, based on the determination result, determine the node execution policy. The specific steps are as follows: Define a risk tolerance threshold and compare it with the asymmetric consistency residual score; If the risk tolerance threshold is less than or equal to the asymmetric consistency residual score, determine that the current state is valid, the recovery process is successful, and the task continues to execute on the target test node; If the risk tolerance threshold is greater than the asymmetric consistency residual score, the node state recovery is not completely credible, there is a potential error risk, and start the task rollback mechanism of the target test node.

[0013] A software remote test system based on a distributed architecture for implementing the above-mentioned software remote test method based on a distributed architecture, including: A node data acquisition module for collecting the multi-dimensional operating state of the remote test node, constructing a time series state vector and extracting the state trajectory, generating an anomaly score for the current time point of the node according to the state trajectory, and judging the node state; A node path update module for removing the faulty nodes and topological edges in the abnormal state, constructing a new topology graph, modeling the scheduling ability of the remaining nodes, constructing a cost function in combination with the path delay, generating a minimum cost path mapping, and synchronously updating the global scheduling structure and routing information; The scheduling graph reconstruction module is used to extract the structured status of the tasks being executed in the failed nodes and encapsulate them into cross-platform snapshots; allocate target nodes according to the scheduling graph and perform status reconstruction; The node status recovery analysis module is used to compare the original node status snapshot with the restored status snapshot, generate an asymmetric consistency scoring function to determine the node status recovery, and determine the node execution strategy based on the determination result.

[0014] The technical effects and advantages of the present invention: The present invention realizes the accurate modeling and risk prediction of the running status of remote test nodes by constructing a node anomaly detection mechanism based on state trajectory analysis, combines topology pruning and resource capacity evaluation, dynamically generates the optimal scheduling path and updates the task mapping structure, and effectively avoids the scheduling risks brought by unstable nodes; during the task execution process, extracts the structured status snapshot of the interrupted task and performs platform-independent encapsulation to realize the seamless migration and status reconstruction of the task on the target node; further constructs an asymmetric residual scoring model by comparing the semantic consistency of the original status and the restored status, quantifies the credibility of the status recovery, and dynamically determines whether the task continues to execute or enters the rollback mechanism based on this, thereby realizing the state-aware intelligent scheduling and hot migration control of test tasks in a distributed environment, and significantly improving the stability, scheduling continuity and logical consistency of task execution of the system in the node fluctuation scenario. Description of the Drawings

[0015] Figure 1 It is a flowchart of a software remote test method based on a distributed architecture according to the present invention.

[0016] Figure 2 It is a schematic structural diagram of a software remote test system based on a distributed architecture according to the present invention. Detailed Embodiments

[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0018] Embodiment 1: As Figure 1 shown, a software remote test method based on a distributed architecture includes the following steps: Collect the multi-dimensional running status of remote test nodes, construct a time series status vector and extract the state trajectory, generate an anomaly score for the current time point of the node according to the state trajectory, and judge the node status; Remove the faulty nodes and topological edges in the abnormal state, and construct a new topological graph. Model the scheduling capabilities of the remaining nodes, construct a cost function in combination with the path delay, generate the minimum-cost path mapping, and synchronously update the global scheduling structure and routing information; Extract the structured state of the tasks being executed in the failed nodes and encapsulate them as cross-platform snapshots; allocate target nodes according to the scheduling graph and perform state reconstruction; Compare the original state snapshot of the node with the restored state snapshot, generate an asymmetric consistency scoring function for node state recovery determination, and determine the node execution strategy based on the determination result.

[0019] During the remote testing process of distributed software, the test tasks are distributed to remote test nodes in different locations. Since these nodes are in heterogeneous networks and inconsistent resource environments, it is easy to have situations such as resource overload, link degradation, or abnormal execution, resulting in test interruption or delay; Step 1, perform state awareness and anomaly detection of remote test nodes, that is, by collecting multi-dimensional state data, constructing a time series model to depict the behavior evolution of nodes, calculating the anomaly score value, and identifying potential faulty nodes in advance, so as to output a trigger signal for topology reconstruction and task migration. The specific steps are as follows: Periodically collect the running states of each node before task scheduling, and identify potential abnormal trends by constructing a state evolution model to make a prediction before the node state deteriorates; First, collect the running state information of each remote test node at time t. The running state information includes resource metrics (such as CPU occupancy rate, memory usage rate, disk I / O saturation), network metrics (such as average delay, packet loss rate, bandwidth occupancy rate), and task metrics (such as average response time, error rate, etc.). The obtained data is fused to construct an original state vector; Normalize the state vectors of different dimensions in the original state vector, perform scale unification processing on states with different dimensions, and obtain a unified state vector after standardization; Introduce a sliding window mechanism to evolve the state trajectory of the node in the recent period of time. This trajectory is a sequence H(t) composed of continuous standardized state vectors, which is used to describe the dynamic change characteristics of the node state; Specifically, use a sliding window to construct the state trajectory for the standardized state vector. Let the window length be , then the state trajectory is: , at this time, the state trajectory represents the multi-dimensional behavior evolution of the node in the recent time instants. Using this trajectory, trend characteristics (such as continuous deterioration or fluctuating instability) can be judged. R represents the set of real numbers, that is, the set composed of all real numbers; By analyzing the trajectory, it is determined whether the node is in an unstable state such as gradually depleting resources, frequent network jitters, or continuous decline in response capabilities. The state trajectory is then input into an embedding encoding model, such as a lightweight one-dimensional convolutional network or a gated recurrent neural network (such as GRU), to compress the multi-dimensional time-series information into a state feature vector z(t) for the next step of risk assessment; According to the state feature vector z(t), an anomaly scoring function for the node at the current time point is defined: , where w is a learnable feature weight vector, T represents the transpose operation, The function is used to enhance the pulling effect of the anomaly trend on the final score; b is a bias term to adjust the center point of the score; An anomaly score value is calculated based on the state feature vector , which is used to quantify whether the current node is in a potentially high-risk state. This score value does not depend on the absolute value of any single metric, but is learned from data in multiple dimensions and at multiple time points, so it has better anti-jitter and generalization capabilities. The higher the anomaly score, the greater the degree to which the node deviates from the normal state at the current time point; A dynamic anomaly threshold is generated based on the historical fluctuations of the score , which is calculated based on the score sequence in a recent period of time, such as taking its 95% quantile value and adding a fluctuation tolerance factor; Maintain a set of historical score sequences , and determine whether to trigger an anomaly response based on its change trend and the relationship between the current value and the historical fluctuation range. Let the decision function be: , where, is the dynamic anomaly threshold, which consists of the upper quantile of the historical score plus a tolerance factor, and is used to dynamically adapt to the state characteristics of different nodes; If the trigger value is 1, immediately mark the node state as an abnormal state and output the following content: node identifier and anomaly mark, the current state trajectory snapshot H(t), the anomaly score value .

[0020] Step 2, Distributed topology reconstruction and scheduling graph update. After completing the node state perception and anomaly detection in Step 1, one or more test nodes with abnormal operating states will be marked. To ensure that tasks are not distributed to unstable nodes and maintain the high availability and load balance of the system, it is necessary to immediately start the topology reconstruction mechanism to adjust the original node connection structure and task mapping path; It is not overly sensitive during periods when node status generally fluctuates greatly, and can respond more promptly when the status is relatively stable. Once the current anomaly score exceeds the dynamic anomaly threshold, the node is considered to be in an abnormal state, and the system immediately generates an anomaly trigger signal, removes the node from the schedulable pool, and marks it as abnormal. After receiving the abnormal node detection results from step 1, the system first needs to remove these no longer safe nodes from the current scheduling topology. The scheduling graph is the basic structure of all task distribution logic. If the abnormal nodes are not removed in time, the test tasks may still be incorrectly distributed to the faulty nodes, resulting in erroneous responses or data loss. The abnormal node identification set output from step 1 is recorded as F, which represents the node set currently determined to be unschedulable. The original scheduling topology can be represented as a graph structure G=(V, E), where V is the set of all nodes, E is the set of connecting edges, and the edges represent the routable paths of tasks. All nodes in the abnormal node identification set F are removed from the graph, and their corresponding topological edges (incoming and outgoing edges) are cleared to generate a pruned new topological graph. ,in, New topology The node set is composed of the original set V after deleting the abnormal node set F, that is, all healthy and schedulable nodes are retained; New topology The edge set represents the result after deleting all edges associated with any abnormal node in the original edge set E; Represents an edge in the original graph; Indicates that at least one of the two endpoints connected by the edge belongs to the abnormal node set F; Eliminate failed nodes and related paths to ensure that subsequent task scheduling is performed on stable nodes, and avoid task loss caused by scheduling paths passing through unavailable nodes; It should be pointed out that this not only involves the physical removal of nodes, but also the updating of their logical connectivity. This is particularly important in the pruning of edges in the graph. For example, if the communication path between two healthy nodes depends on a relay node, and the relay node is abnormal, the connection between the two healthy nodes should also be considered broken. This strategy ensures that subsequent path planning is completely based on trusted connections.

[0021] After the structure is eliminated, the task path cannot be rebuilt immediately. The reason is that although the health of the remaining nodes in the system meets the basic scheduling conditions, their carrying capacity is uneven. Some nodes may have sufficient idle resources, while others are close to resource bottlenecks. If tasks are dispatched without considering this, the local load of the system will be increased, the overall execution efficiency will be reduced, and even new congestion or node overload will be triggered. Therefore, perform a capacity analysis on the remaining available nodes in the new topology map, construct the current resource adaptation map, and define each node 's scheduling capacity as: , where represents the current available resources (such as the remaining number of CPU cores, free memory, etc.); represents the total amount of resources of this node; is a numerical stability constant to avoid division by zero; represents the relative resource margin of the current node.

[0022] The resource adaptation map is used for subsequent task distribution and path weight assignment, and is the input basis for scheduling path optimization.

[0023] Perform resource capacity modeling on each remaining node, extract information such as its computing resource availability, remaining task processing space, and concurrent support intensity at the current moment, and quantify it into a comparable scheduling capacity scoring index, the relative resource margin The smaller it is, the more crowded the node is and the less suitable it is to continue to undertake tasks; on the contrary, a high relative resource margin indicates that the node has abundant resources and strong response capabilities.

[0024] Perform scheduling path reconstruction and minimum-cost mapping generation. After node capacity evaluation, enter the most critical reconstruction stage. At this time, the scheduler needs to select an optimal path for each task in the current set of tasks to be scheduled in the trimmed topology map and point to a suitable test node. The specific steps are as follows: Define the set of test tasks to be scheduled as , where m is the total number of test tasks, and the goal is to reconstruct the task assignment path on the trimmed new topology map; Construct a task scheduling mapping graph , where represents a mapping relationship from a test task to a target execution node, that is, a function that assigns each task to a node for execution (for example, represents the mapping function from task to node), and path selection is based on the following optimization criteria; Let the total cost of each feasible path from the scheduling center to node be: , where is the path delay (including link transmission time); is the node scheduling capacity index; , is a weight parameter, and its value is greater than 0, which is used to balance path transmission efficiency and resource adaptability; By traversing all paths, select the target node with the minimum total cost for each task: represents the set of test tasks to be scheduled The k-th test task in

[0025] Perform task graph synchronization update and topology cache replacement. That is, when all tasks are mapped, update the global scheduling graph structure. This update includes two parts: Update the task routing table inside the scheduler , indicating the path along which the task should be sent to the new target node; Update the topology cache graph of the system. Replace the original graph with the updated new topology graph and use it as the scheduling graph for node testing, and record the adjustment timestamp to support topology rollback after the abnormal node recovers; In addition, a path adjustment notice can be generated, including: the task identifier of the reconstructed path, node information before and after scheduling, and details of path cost changes.

[0026] It should be noted that the task routing table is a task-level mapping structure maintained inside the scheduler, used to record the current target execution node corresponding to each test task and the scheduling path information it needs to pass through. It is one of the core data structures for the scheduler to perform task distribution and status tracking. The task routing table records task-node mapping relationships, stores scheduling paths, task execution status tracking, etc.; the topology cache graph is a copy of the graph structure maintained in the system scheduler, used to save the status snapshot of the effective scheduling graph in the current test system, and serve as the working graph or intermediate graph version for scheduling execution, to support the system to retain the topology historical state and assist in restoring scheduling consistency after operations such as node state changes, path reconstruction, and task migration.

[0027] In summary, this step integrates three factors: task characteristics, path constraints, and node status, and achieves the optimal matching of scheduling performance under the premise of ensuring availability. Therefore, when some nodes in the test environment fail, the system can quickly complete topology update, task remapping, and path optimization, and keep the overall scheduling uninterrupted.

[0028] Step 3, perform test task status snapshot and hot migration execution. After completing the topology reconstruction and task scheduling graph update in Step 2, the scheduler has completed the path reconstruction and target node mapping of the tasks on the failed nodes. However, for those test tasks that are already running on the faulty nodes, if they are directly aborted and then restarted, it will cause unnecessary test repetition, data loss, and context confusion problems. Therefore, design a task status snapshot extraction and hot migration mechanism to ensure that the task can seamlessly continue to execute on the new target node after the task operation is interrupted. The specific steps are as follows: When the scheduler receives the notice of node abnormality, immediately freeze all active task sets on it, denoted as the active set: , where m is the total number of active tasks. Each active task is being executed, and its internal state needs to be extracted as a structured snapshot; The content of the task status snapshot (snapshot tuple) includes, but is not limited to: Script execution pointer position : Records the current test execution position; Internal variable set : Contains the values of context variables during execution; Environment binding tuple : Represents the execution environment on which the task depends; Temporary file handle and file buffer structure description (if any); Network session status summary : Describes whether the communication session is being maintained.

[0029] Extracting these running states will cause the task to be forced to terminate and reset, losing its original execution continuity. Therefore, it is necessary to quickly extract the task status before the node exits (or through an asynchronous log compensation mechanism) and organize the above status information into a snapshot tuple: .

[0030] After extracting the snapshot tuple, it cannot be directly transmitted to the target node because the underlying running platforms of different test nodes (such as operating system versions, test toolchains, library dependency paths, etc.) may vary. If the format and semantics of the snapshot content are not uniformly converted, it may cause the post-migration state loading to fail or the task logic to deviate; Convert the snapshot tuple into a cross-platform neutral representation form and define a snapshot encapsulation function: , where is the encapsulation function, is the snapshot tuple; The encapsulation process of the encapsulation function includes: performing dependency mapping conversion on environment dependencies, renaming incompatible component paths, serializing the internal variable set into a platform-neutral format; converting the network status summary into a replayable connection status description to support session reconstruction; establishing an offset mapping for the script pointer position to make it consistent with the target node's parsing.

[0031] After completing the snapshot encapsulation, the scheduler maps each task to be restored to its target execution node according to the scheduling graph constructed in step 2, and the target execution receives the encapsulated snapshot , starts the task recovery process in the container isolation environment, thereby eliminating the task termination problem caused by node failure, and finally the task status restores the snapshot is regarded as equivalent to the logical state of the original task at the time of interruption and can continue to execute on the target node.

[0032] During the cross-node reconstruction of the task status, recovery distortion or context loss may occur due to reasons such as platform environment, serialization deviation, and state mapping conflicts. Therefore, a state consistency check is performed before the task resumes execution.

[0033] Step 4, perform a consistency verification and recovery confirmation mechanism, and the specific steps are as follows: Receive the task status recovery snapshot after the migration in Step 3 and simultaneously retain the snapshot tuple of the original migration The structural descriptions in both are structured state vectors, containing multiple heterogeneous fields such as variables, execution pointers, environment dependencies, session information, etc.; Construct a state space mapping function to unify the two sets of state structures into the same semantic space S for subsequent consistency comparison: where, represents the semantic vector of the original state, represents the semantic reconstruction vector of the target state; This state space mapping function should satisfy two basic properties: each logical field in the state is semantically comparable; it can adapt to the running data structures in different system platforms (such as variable names, library versions, system description information, etc.).

[0034] Construct an asymmetric consistency residual score for the hot migration scenario of the target node, allowing for slight differences in some dimensions of the state, but maintaining strict consistency for key variables or path structures. Define the asymmetric consistency residual score as follows: where d is the dimension of the state vector (i.e., the number of state components); 、 are the feature components of the original state and the migrated state at the k-th item; is the state difference measure function for the k-th item (such as structural edit distance, string matching difference, call stack similarity, etc.); is the residual for the k-th item.

[0035] Asymmetric consistency residual score The lower the value of, the closer the state recovery is to the original state.

[0036] After completing the multi-dimensional semantic comparison and consistency residual calculation of the task status, the system needs to make a comprehensive determination of the current recovery state of the hot migration task to decide whether to allow the task to continue executing on the target node, or whether to roll back the task status to the original snapshot or an alternative migration path; Define the risk tolerance threshold and compare it with the asymmetric consistency residual score to determine the sensitivity of the current task to consistency, so as to judge whether the current recovery state is trustworthy enough. For example, interface test tasks have extremely high requirements for session recovery, while stateless script tasks can tolerate a certain execution deviation; If the risk tolerance threshold is less than or equal to the asymmetric consistency residual score, the system determines that the current state is valid, the recovery process is successful, and the task can continue to be executed on the target test node; at the same time, record the successful migration path and state mapping parameters into the task scheduler log for subsequent migration strategy optimization; If the risk tolerance threshold is greater than the asymmetric consistency residual score, the system will consider that the state recovery is not completely trustworthy and there is a potential error risk. At this time, the task rollback mechanism of the target test node will be started, which generally includes the following processing paths: reschedule the task to a pre-configured backup node, reload the snapshot and execute the recovery again; or in the scenario without backup resources, mark the task as a migration failure, convert it to a failure processing state, and wait for manual intervention or the system fault tolerance strategy to intervene.

[0037] It should be noted that the thresholds involved in the embodiments are set according to specific scenarios and requirements.

[0038] The present invention realizes the accurate modeling and risk prediction of the running state of remote test nodes by constructing a node anomaly detection mechanism based on state trajectory analysis, combines topology pruning and resource capacity evaluation, dynamically generates the optimal scheduling path and updates the task mapping structure, effectively avoiding the scheduling risks brought by unstable nodes; during the task execution process, extracts the structured state snapshot of the interrupted task and performs platform-independent encapsulation to realize the seamless migration and state reconstruction of the task on the target node; further constructs an asymmetric residual score model by comparing the semantic consistency of the original state and the restored state, quantifies the credibility of the state recovery, and dynamically determines whether the task continues to execute or enters the rollback mechanism accordingly, thus realizing the intelligent scheduling and hot migration control of test tasks based on state awareness in a distributed environment, and significantly improving the stability, scheduling continuity and logical consistency of task execution of the system in the node fluctuation scenario.

[0039] Embodiment 2: A software remote test system based on a distributed architecture, as Figure 2 shown, specifically includes: A node data acquisition module, which is used to collect the multi-dimensional running state of remote test nodes, construct a time series state vector and extract the state trajectory, generate an anomaly score for the current time point of the node according to the state trajectory, and judge the node state; The node path update module is used to eliminate faulty nodes and topological edges in abnormal states, construct a new topology graph, model the scheduling capabilities of the remaining nodes, construct a cost function in combination with path delay, generate a minimum-cost path mapping, and synchronously update the global scheduling structure and routing information; The scheduling graph reconstruction module is used to extract the structured state of the tasks being executed in the failed nodes and encapsulate them as cross-platform snapshots; allocate target nodes according to the scheduling graph and perform state reconstruction; The node state recovery analysis module is used to compare the original state snapshot of the node with the restored state snapshot, generate an asymmetric consistency scoring function for node state recovery determination, and determine the node execution strategy based on the determination result.

[0040] The above formulas are all dimensionless and take their numerical values for calculation. Specific dimensionless methods can use various means such as standardization, which will not be elaborated here. The formulas are obtained by collecting a large amount of data for software simulation to obtain a formula closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0041] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, ATA hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state ATA hard disk.

[0042] It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0043] Those of ordinary skill in the art will realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0044] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0045] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0046] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0047] As described above, this is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A software remote testing method based on a distributed architecture, characterized in that It includes the following steps: Collect the multi-dimensional operating status of remote test nodes, construct a time series status vector and extract the status trajectory, generate the anomaly score of the node at the current time point according to the status trajectory, and judge the node status; Eliminate the faulty nodes and topological edges in the abnormal state, construct a new topology graph, model the scheduling ability of the remaining nodes, construct a cost function in combination with the path delay, generate the minimum cost path mapping, and synchronously update the global scheduling structure and routing information; Extract the structured status of the tasks being executed in the failed nodes and encapsulate them into cross-platform snapshots; allocate target nodes according to the scheduling graph and perform status reconstruction; Compare the original state snapshot of the node with the restored state snapshot, generate an asymmetric consistency scoring function to determine the node status recovery, and determine the node execution strategy based on the determination result.

2. The software remote testing method based on a distributed architecture according to claim 1, wherein: Collect the operating status information of remote test nodes, construct a time series status vector and extract the status trajectory. The specific process is as follows: Collect the operating status information of each remote test node. The operating status information includes the CPU occupancy rate, memory usage rate, disk I / O saturation rate in the resource metrics, the average delay, packet loss rate, and bandwidth occupancy rate in the network metrics, and the average response time and error rate in the task metrics, and fuse the obtained data to construct the original state vector; Normalize the state vectors of different dimensions in the original state vector; Use a sliding window to extract the status trajectory of the node within the set time. The status trajectory is a sequence composed of continuous standardized state vectors, which is used to describe the dynamic change characteristics of the node status.

3. A software remote testing method based on a distributed architecture according to claim 2, characterized in that: Generate the anomaly score of the node at the current time point according to the status trajectory and judge the node status. The specific process is as follows: Input the status trajectory into a one-dimensional convolutional network or a gated recurrent neural network to compress the temporal information in the status trajectory into a state feature vector; According to the state feature vector, define the anomaly scoring function of the node at the current time point and determine the anomaly score value of the node; Compare the dynamic anomaly threshold with the anomaly score value to judge the node status; When the dynamic anomaly threshold is greater than the anomaly score value, trigger an anomaly response and mark the node status as an abnormal state.

4. A software remote testing method based on a distributed architecture according to claim 3, characterized in that: Eliminate the faulty nodes and topological edges in the abnormal state and construct a new topology graph. Model the scheduling ability of the remaining nodes. The specific steps are as follows: Summarize the nodes marked as abnormal states as the abnormal node identification set; Eliminate all the nodes in the abnormal node identification set from the graph, clear the incoming and outgoing edges corresponding to the nodes, and use the eliminated failed nodes and path graph as the new topology graph; Perform a capacity analysis on the remaining available nodes in the new topology map, construct the current resource adaptation map, and determine each node 's scheduling capacity as: , where represents the currently available resources; represents the total node resources; is a numerical stability constant to avoid division by zero; represents the relative resource margin of the current node.

5. A software remote testing method based on a distributed architecture according to claim 4, characterized in that: Construct a cost function in combination with the path delay, generate the minimum cost path mapping, and synchronously update the global scheduling structure and routing information. The specific process is as follows: Define the set of test tasks to be scheduled as , where m is the total number of test tasks; Construct a task scheduling mapping graph , where represents a mapping relationship from a test task to a target execution node; The path selection is set based on the optimization criterion, and the total cost of each feasible path from the dispatching center to the node is: , where is the path delay (including the link transmission time); is the node scheduling ability index; , is the weight parameter, and the values are all greater than 0; By traversing all paths, select the target node with the minimum total cost for each task: Denote the k-th test task in the set of test tasks to be scheduled; When all tasks are mapped, update the global scheduling graph structure, including updating the task routing table and topology cache graph inside the scheduler.

6. A software remote testing method based on a distributed architecture according to claim 5, characterized in that: Extract the structured status of the tasks being executed in the failed nodes and encapsulate them into cross-platform snapshots; allocate target nodes according to the scheduling graph and perform status reconstruction. The specific process is as follows: When the scheduler receives a notification of node exception, it immediately freezes all active tasks on it and records them as the active set , where m is the total number of active tasks; Extract the internal state of the task status for each active task and generate a snapshot tuple; The content of the snapshot tuple includes: the position of the script execution pointer, the set of internal variables, the environment binding tuple, the description of the temporary file handle and the file buffer structure, and the summary of the network session status; Convert and encapsulate the snapshot tuple into a cross-platform expression form; The encapsulation process includes performing a dependency mapping transformation on the environmental dependencies, renaming the paths of incompatible components, serializing the set of internal variables into a platform-neutral format; transforming the network status summary into a replay connection status description; establishing an offset mapping for the script pointer position; After completing the snapshot encapsulation, the scheduler maps each task to be restored to the target execution node and obtains the final task status recovery snapshot after the migration is completed.

7. A software remote testing method based on a distributed architecture according to claim 6, characterized in that: Compare the original state snapshot of the node with the restored state snapshot, and generate an asymmetric consistency scoring function to determine the node state recovery, and the specific steps are as follows: Construct a state space mapping function to unify the task status recovery snapshot and the snapshot tuple into the same semantic space; Construct an asymmetric consistency residual score for the hot migration scenario of the target node. The expression for the asymmetric consistency residual score is: , where d is the dimension of the state vector; , are the feature components of the original state and the migrated state at the k-th item; is the call stack similarity at the k-th item; is the residual of the call stack similarity at the k-th item.

8. A software remote testing method based on a distributed architecture according to claim 7, characterized in that: And determine the node execution policy based on the determination result, and the specific steps are as follows: Define a risk tolerance threshold for comparison with the asymmetric consistency residual score; If the risk tolerance threshold is less than or equal to the asymmetric consistency residual score, it is determined that the current state is valid, the recovery process is successful, and the task continues to execute on the target test node; If the risk tolerance threshold is greater than the asymmetric consistency residual score, the node state recovery is not completely credible, there is a potential error risk, and the task rollback mechanism of the target test node is started.

9. A software remote testing system based on a distributed architecture, which is used to implement a software remote testing method based on a distributed architecture according to any one of claims 1-8, and is characterized in that, Including: A node data collection module, which is used to collect the multi-dimensional running status of the remote test node, construct a time series state vector and extract the state trajectory, generate an anomaly score for the current time point of the node according to the state trajectory, and judge the node state; A node path update module, which is used to eliminate the faulty nodes and topological edges in the abnormal state, construct a new topological graph, model the scheduling ability of the remaining nodes, combine the path delay to construct a cost function, generate a minimum cost path mapping, and synchronously update the global scheduling structure and routing information; A scheduling graph reconstruction module, which is used to extract the structured state of the tasks being executed in the failed nodes and encapsulate them into cross-platform snapshots; allocate target nodes according to the scheduling graph and perform state reconstruction; A node state recovery analysis module, which is used to compare the original state snapshot of the node with the restored state snapshot, generate an asymmetric consistency scoring function to determine the node state recovery, and determine the node execution policy based on the determination result.

Citation Information

Patent Citations

  • Distributed automatic test vector generation method, device and system capable of dynamically expanding capacity

    CN117851107A

  • System and method for using an automated process to identify bugs in software source code

    US20050223357A1

  • Synthesis of concurrent schedulers for multicore architectures

    US20110302584A1

Cited By

  • Cross-platform integration-oriented enterprise data collaborative scheduling method and system

    CN120973490A

  • An enterprise data collaborative scheduling method and system for cross-platform integration

    CN120973490B

  • Performance optimization method for reverse debugging deterministic replay

    CN121349842A

  • High-end memory chip asynchronous test system based on distributed architecture

    CN121789747A