Network diagnosis processing method, device, system, equipment, medium and product
By employing cross-layer correlation analysis and a cloud-edge collaborative architecture, the problems of low fault identification accuracy and frequent invalid switching in existing network operation and maintenance solutions have been solved, thereby improving the stability and reliability of device communication.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI SIGE DIGITAL TECHNOLOGY CO LTD
- Filing Date
- 2026-05-29
- Publication Date
- 2026-07-21
AI Technical Summary
Existing network operation and maintenance solutions suffer from low fault identification accuracy and frequent invalid switching, resulting in insufficient stability and reliability of equipment communication. In particular, communication networks are prone to frequent oscillations in complex environments.
By employing cross-layer correlation analysis technology, network diagnostic results are generated by acquiring multi-layer network status parameters of the target device, including fault root cause identification and communication link switching strategies. Based on the strategy execution verification results, link switching operations are performed, and adaptive optimization is carried out in conjunction with the cloud-edge collaborative architecture.
It improved the accuracy of fault identification, avoided invalid switching, ensured the continuous and stable operation of the communication network, and significantly improved the stability and reliability of equipment communication.
Smart Images

Figure CN122437758A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of network technology, and more specifically, relates to a network diagnostic processing method, apparatus, system, equipment, medium and product. Background Technology
[0002] Currently, distributed IoT edge devices such as photovoltaic storage systems generally use hybrid communication links such as Wi-Fi and 4G to achieve data reporting and command interaction. The stability of the communication network directly determines whether the device services can operate continuously.
[0003] Existing network operation and maintenance solutions often rely on single anomaly detection methods, resulting in low fault identification accuracy and a tendency to trigger invalid handovers. Furthermore, traditional link switching is based solely on fixed priorities, leading to unexpected outcomes after switching, poor network stability, and frequent oscillations, ultimately resulting in low reliability of device communication in complex environments. Summary of the Invention
[0004] The purpose of this application is to provide a network diagnostic processing method, apparatus, system, device, medium and product, which aims to solve the technical problems of low fault identification accuracy and frequent invalid switching in existing network operation and maintenance solutions, resulting in insufficient communication stability and reliability of equipment.
[0005] To achieve the above objectives, according to the first aspect of this application, a network diagnostic processing method is provided, applied to a target device, the method comprising: Obtain the multi-layer network status parameters of the current communication link of the target device; Cross-layer correlation analysis is performed on the multi-layer network status parameters to generate network diagnostic results, wherein the network diagnostic results include fault root cause identification and communication link switching strategy; Determine the policy execution verification result corresponding to the communication link switching policy; Based on the verification results of the strategy, a communication link switching operation is performed.
[0006] The beneficial effects of the embodiments in this application compared with the prior art are: The network diagnostic and processing method provided in this application addresses the technical problems of low fault identification accuracy, frequent invalid handovers, and insufficient communication stability and equipment reliability in existing network operation and maintenance solutions. First, by acquiring multi-layer network status parameters of the target device's current communication link, it avoids the limitation of traditional solutions that rely on a single parameter to determine network anomalies. Second, through cross-layer correlation analysis, network diagnosis can analyze the obtained root cause identifiers of faults and generate communication link switching strategies, effectively distinguishing between instantaneous network jitter, unrelated concurrent anomalies, and real composite faults, thus improving fault identification accuracy. Finally, by executing communication link switching operations based on the verification results, meaningless blind switching can be avoided, effectively ensuring the continuous and stable operation of the communication network and significantly improving equipment communication stability and reliability.
[0007] According to a second aspect of this application, a network diagnostic processing method is provided, applied to a server, the method comprising: Obtain the multi-layer network status parameters of the current communication link of the target device; Cross-layer correlation analysis is performed on the multi-layer network status parameters to generate network diagnostic results. The network diagnostic results include fault root cause identification and communication link switching strategy. The communication link switching strategy is determined based on the fault root cause identification, preset link priority ranking, and fault identification threshold. Determine the policy execution verification result corresponding to the communication link switching policy; Based on the verification results of the strategy, adjust at least one of the model parameters used in the link priority sorting, fault identification threshold, and cross-layer correlation analysis.
[0008] According to a third aspect of this application, a network diagnostic processing system is provided, the system comprising: The target device is used to obtain the multi-layer network status parameters of the current communication link of the target device; The server connects to the target device and is used to perform cross-layer correlation analysis on the multi-layer network status parameters and generate network diagnostic results. The network diagnostic results include fault root cause identification and communication link switching strategy. The communication link switching strategy is determined based on the fault root cause identification, preset link priority ranking, and fault identification threshold. The target device is further configured to determine the policy execution verification result corresponding to the communication link switching policy, and to perform a communication link switching operation based on the policy execution verification result; The server is also used to perform verification results according to the strategy and adjust at least one of the model parameters used in the link priority sorting, fault identification threshold and cross-layer correlation analysis.
[0009] According to a fourth aspect of this application, a network diagnostic processing apparatus is provided, applied to a target device, the apparatus comprising: The acquisition unit is used to acquire the multi-layer network status parameters of the current communication link of the target device; The analysis unit is used to perform cross-layer correlation analysis on the multi-layer network status parameters and generate network diagnostic results, wherein the network diagnostic results include fault root cause identification and communication link switching strategy. A determining unit is used to determine the policy execution verification result corresponding to the communication link switching policy; The execution unit is used to perform communication link switching operations based on the verification results of the strategy.
[0010] According to a fifth aspect of this application, a network diagnostic processing apparatus is provided, applied to a server, the apparatus comprising: The acquisition module acquires the multi-layer network status parameters of the current communication link of the target device. The analysis module is used to perform cross-layer correlation analysis on the multi-layer network status parameters and generate network diagnostic results. The network diagnostic results include fault root cause identification and communication link switching strategy. The communication link switching strategy is determined based on the fault root cause identification, preset link priority ranking and fault identification threshold. The determining module is used to determine the policy execution verification result corresponding to the communication link switching policy; An optimization module is used to adjust at least one of the model parameters used in the link priority sorting, fault identification threshold, and cross-layer correlation analysis based on the verification results of the strategy.
[0011] According to a sixth aspect of this application, an electronic device is provided, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the electronic device causes the electronic device to perform the method as described in any one of the claims.
[0012] According to a seventh aspect of this application, a computer-readable storage medium is provided that stores a computer program, which, when executed by a processor, implements the method as described in any one of the claims.
[0013] According to the eighth aspect of this application, a computer program product is provided that, when the computer program product is run on an electronic device, causes the electronic device to perform the method described in any one of the first aspects above.
[0014] It is understandable that the beneficial effects of the second to eighth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a schematic block diagram of a network diagnostic processing system provided in an embodiment of this application; Figure 2 This is a schematic flowchart of a network diagnostic processing method provided in an embodiment of this application; Figure 3 This is a schematic flowchart of a network diagnostic processing method provided in an embodiment of this application; Figure 4 This is a schematic flowchart of a network diagnostic processing method provided in an embodiment of this application; Figure 5 This is a schematic flowchart of a network diagnostic processing method provided in an embodiment of this application; Figure 6 This is a schematic flowchart of a network diagnostic processing method provided in an embodiment of this application; Figure 7 This is a flowchart illustrating another network diagnostic processing method provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of a network diagnostic processing device provided in an embodiment of this application; Figure 9 This is a schematic diagram of another network diagnostic processing device provided in an embodiment of this application; Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0017] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0018] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0019] It should also be understood that, in the description of this application, unless otherwise stated, the " / " used in the specification and appended claims indicates that the related objects are in an "or" relationship. For example, A / B can mean A or B. The "and / or" in this application is merely a description of the relationship between the related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, and A and B exist, or B exists alone. A and B can be singular or plural. Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0020] Furthermore, to facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, but are only used for distinguishing descriptions, and the terms "first" and "second" do not necessarily imply that they are different, nor should they be construed as indicating or implying relative importance.
[0021] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0022] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0023] First, some terms used in the embodiments of this application will be explained to facilitate understanding by those skilled in the art.
[0024] Multi-layer network status parameters: refers to the set of operational status parameters covering all layers of network communication. From top to bottom, it includes various status parameters of the service layer, application layer, connection establishment layer, transport layer, and physical layer. It can completely characterize the full-dimensional operation of the target device's communication link from service interaction and protocol transmission to underlying signal transmission.
[0025] Cross-layer correlation analysis: refers to a comprehensive analysis method that breaks through the limitations of independent detection at a single network level and performs time-series alignment, anomaly screening, causal correlation, and model matching on multi-level network state parameters. It is used to eliminate accidental concurrent anomalies and accurately locate the root cause of network failures.
[0026] Fault root cause identifier: refers to a standardized identifier used to uniquely characterize the attributes of network faults, including fault occurrence level information, fault type information, and fault severity information, and is the decision basis for generating accurate link switching strategies.
[0027] Fault Root Cause Association Model: This refers to an intelligent matching model pre-stored on the server. It internally stores multiple sets of abnormal parameter combinations at different levels and mapping relationships with fault root cause identifiers. Each mapping relationship contains at least two abnormal parameters at different network levels, which is used to achieve intelligent and accurate fault matching.
[0028] Link priority ranking: refers to the pre-configured order of preference for each communication link, which is set by comprehensively considering factors such as link stability, transmission rate, communication cost, and anti-interference capability, and is used to select the optimal switching link from multiple candidate links.
[0029] Policy execution verification results: refers to the effect evaluation results obtained by collecting multi-layer network parameters and cross-layer analysis within a preset verification window period after the target device executes the link switching policy. It is used to characterize the effectiveness of the switching policy, the stability of link operation, and the link jitter status.
[0030] The above is a brief introduction to the terms used in the embodiments of this application, and will not be repeated below.
[0031] With the large-scale deployment of distributed edge devices, optical storage terminals, and IoT terminals, these devices rely on wireless and wired networks to report data and exchange commands. The stability of network communication directly determines the secure and reliable operation of terminal services. Currently, terminal network operation and maintenance generally adopts a local single threshold detection and fixed link switching solution, which has many technical defects and is difficult to adapt to complex and ever-changing field network environments.
[0032] First, existing network diagnostic solutions suffer from low fault identification accuracy and high false positive rate. Traditional solutions rely solely on a single network layer and a single parameter threshold to identify anomalies, failing to distinguish between instantaneous fluctuations in network parameters, unrelated concurrent anomalies, and real systemic faults. This easily leads to issues such as false switching when there is no fault and missed identification of real faults, frequently causing brief service interruptions and network link oscillations.
[0033] Secondly, existing link switching strategies are poorly targeted and prone to haphazard switching. Traditional link switching often adopts a fixed-priority forced switching mode, without matching and adapting links based on the root cause and level of the fault. This can lead to situations where a fault cannot be repaired by link switching but is forcibly switched, or where a valid link is incorrectly replaced, significantly reducing the network's self-healing efficiency.
[0034] Secondly, existing network systems lack adaptive optimization capabilities. Traditional network diagnostics and switching rules use fixed preset parameters and cannot be dynamically updated based on the on-site network environment, equipment operating status, and historical switching results. After long-term operation, they cannot adapt to environmental changes, leading to a continuous decline in the effectiveness of network fault handling.
[0035] Finally, the existing network operation and maintenance architecture has an unreasonable allocation of computing power. Purely local terminal diagnostic computing power is limited and cannot complete complex cross-layer correlation analysis and model matching; pure cloud-centralized architecture completely loses its self-healing ability when the cloud connection is lost, and cannot realize the complementary advantages of cloud and edge, resulting in extremely poor network reliability in extreme scenarios.
[0036] To address the aforementioned technical problems, this application provides an example of a network diagnostic processing system. Please refer to [example provided]. Figure 1 As shown, Figure 1 This application provides a schematic structural block diagram of a network diagnostic processing system, which includes: Target device 101 is used to obtain the multi-layer network status parameters of the current communication link of the target device; Server 102 is communicatively connected to target device 101 and is used to perform cross-layer correlation analysis on multi-layer network status parameters and generate network diagnostic results. The network diagnostic results include fault root cause identification and communication link switching strategy. The communication link switching strategy is determined based on fault root cause identification, preset link priority ranking and fault identification threshold. The target device 101 is also used to determine the policy execution verification result corresponding to the communication link switching policy, and to perform the communication link switching operation according to the policy execution verification result; Server 102 is also used to adjust at least one of the model parameters used in link priority sorting, fault identification threshold and cross-layer correlation analysis based on the policy execution verification results.
[0037] This application provides a network diagnostic processing system that adopts a cloud-edge collaborative architecture. It combines cloud-based intelligent analysis and decision-making with terminal local execution verification and adaptive optimization capabilities to achieve full-process self-healing, including accurate diagnosis of network faults, accurate policy distribution, closed-loop verification of execution effects, and iterative optimization of system parameters. This effectively improves the network communication stability and self-healing capabilities of distributed terminal devices.
[0038] The network diagnostic processing system provided in this application includes a target device and a server. The target device and the server establish a two-way data interaction channel through a communication link to realize closed-loop collaboration of data uploading, policy distribution, result feedback and parameter synchronization.
[0039] The target device is an edge terminal device with network communication, data acquisition, and policy execution capabilities. Specifically, it can be an embedded intelligent device such as a photovoltaic energy management terminal, a photovoltaic inverter, or an IoT operation and maintenance terminal. It mainly undertakes edge-side functions such as network parameter acquisition, link switching execution, and policy effect verification.
[0040] Servers (such as cloud servers) are cloud computing platforms with big data computing, intelligent model matching, and global parameter optimization capabilities. They mainly undertake cloud decision-making functions such as cross-layer fault analysis, diagnostic result generation, switching strategy formulation, and global parameter iterative optimization.
[0041] In some embodiments, the target device is used to collect multi-layer network status parameters corresponding to its current communication link in real time. The target device has a built-in multi-dimensional network parameter acquisition module that can cover all or part of the parameter acquisition of the physical layer, transport layer, connection establishment layer, application layer, and service layer. It can comprehensively capture all-dimensional operational data of the communication link from the underlying signal transmission to the upper-layer service interaction, providing complete and authentic raw data support for cloud-based fault diagnosis and ensuring the comprehensiveness of data analysis.
[0042] In some embodiments, the cloud server establishes a regular communication connection with the target device, receives multi-layer network status parameters uploaded by the target device, and performs cross-layer correlation analysis on the multi-layer network status parameters. Relying on sufficient computing resources, the cloud server completes a series of operations, including time-series alignment of multi-layer parameters, standardization and normalization, layer-by-layer anomaly verification, cross-layer causal correlation, and fault model matching, to generate a fault root cause identifier that includes level, type, and severity.
[0043] Based on this, the cloud server combines fault root cause identification, preset link priority ranking, and fault identification threshold to intelligently generate network diagnostic results adapted to the current fault scenario. These diagnostic results include accurate fault root cause identification and implementable communication link switching strategies, avoiding the problem of communication errors caused by blindly switching communication links in traditional solutions.
[0044] In some embodiments, the target device receives a communication link switching policy from a cloud server and completes the policy implementation and effect verification. For example, the target device continuously collects multi-layer network status parameters of the newly switched link within a preset verification window period, generates a policy execution verification result that includes policy effectiveness, link stability, and link jitter status through local secondary cross-layer analysis, determines the effect of this switching based on the verification result, and finally completes a stable and reliable communication link switching operation to ensure the continuous operation of the target device's terminal services.
[0045] In some embodiments, the cloud server possesses global adaptive iterative optimization capabilities. The cloud server continuously receives policy execution verification results from each target device and dynamically adjusts system parameters based on real-world implementation results. It can selectively optimize at least one of the following: link priority ranking updates, fault identification threshold corrections, and cross-layer correlation analysis model parameter iterations. For effective and stable links, the preferred weight is increased; for invalid fault scenarios, the judgment threshold is corrected; and for novel and unknown faults, the model rules are updated, enabling the continuous self-evolution of the entire network diagnostic system.
[0046] The network diagnostic processing system provided in this embodiment constructs a cloud-edge collaborative architecture of terminal data acquisition and execution, cloud-based analysis and decision-making, and closed-loop iterative optimization, fully leveraging the dual advantages of real-time local execution on the terminal and global intelligent computation in the cloud. It improves the accuracy of network fault root cause identification by replacing traditional single-threshold detection with multi-layer cross-layer correlation analysis; it completely solves the technical drawbacks of blindly switching communication links leading to communication anomalies through fault root cause matching link strategy generation; it allows the system to continuously adapt to complex and ever-changing on-site network environments through adaptive parameter iteration after execution verification; and it reduces the computing power pressure on edge terminals while ensuring the basic self-healing capabilities of terminals in cloud-based link loss scenarios through cloud-edge collaboration, comprehensively improving the stability, self-healing, and intelligence level of distributed terminal device network communication.
[0047] This application provides an example of a network diagnostic processing method. Please refer to [link / reference]. Figure 2 As shown, Figure 2 A schematic flowchart of a network diagnostic processing method provided in this application is shown. This is an example and not a limitation; the method can be applied to or operated in electronic devices, such as photovoltaic and energy storage devices, specifically photovoltaic inverters, energy storage batteries, energy management systems, etc. The method includes: S201, Obtain the multi-layer network status parameters of the current communication link of the target device; S202, perform cross-layer correlation analysis on the status parameters of the multi-layer network to generate network diagnostic results, which include fault root cause identification and communication link switching strategy. S203, Determine the policy execution verification result corresponding to the communication link switching policy; S204, Based on the policy execution verification result, perform a communication link switching operation.
[0048] This embodiment provides a network diagnostic processing method for target devices, particularly suitable for edge devices (i.e., target devices) with independent computing capabilities, such as photovoltaic inverters, energy storage battery management systems, and energy management terminals in distributed photovoltaic-storage systems. This method completes network diagnostics and communication link switching locally on the device, without relying on a cloud server, and can ensure the continuity and reliability of device communication even in extreme scenarios such as a disconnection between the target device and the cloud server.
[0049] In some embodiments, the target device acquires multi-layer network status parameters of the current communication link in real time through its built-in network interface chip and protocol stack. These parameters are collected in a categorized manner according to the hierarchical structure of network communication, covering the physical layer, transport layer, connection establishment layer, application layer, and service layer. Parameters at different layers reflect the operational status of different aspects of network communication. For example, physical layer parameters mainly reflect the basic quality of signal transmission, transport layer parameters reflect the reliability and stability of data transmission, connection establishment layer parameters reflect the efficiency of the network connection establishment process, application layer parameters reflect the protocol interaction status between the device and the server, and service layer parameters directly reflect the network's ability to support core services.
[0050] In this embodiment, the target device performs local cross-layer correlation analysis on the collected multi-layer network status parameters. Unlike traditional single-threshold judgments, which can only identify network anomalies but cannot pinpoint the root cause, cross-layer correlation analysis establishes causal relationships between parameters at different levels, eliminating non-causal concurrent anomalies, thereby accurately identifying the root cause of network faults and generating network diagnostic results that include fault root cause identifiers and communication link switching strategies. The fault root cause identifier includes information on the fault's level, type, and severity. The communication link switching strategy is generated based on the fault root cause, ensuring that only links capable of resolving the current fault are included in the switching scope, avoiding ineffective operations and service interruptions caused by blind switching.
[0051] In this embodiment, the target device determines the policy execution verification result corresponding to the communication link switching strategy. After executing the link switching operation, the target device enters a preset verification window period, during which it continuously collects multi-layer network status parameters of the new communication link. By statistically analyzing the parameters over multiple consecutive sampling periods, the operating status of the new link (the target switching link) is comprehensively judged. For example, the policy execution verification result includes a policy validity indicator, a new link stability indicator, and a jitter indicator, which respectively reflect whether the switching strategy has achieved the expected effect, whether the new link can operate stably, and whether there are frequent fluctuations.
[0052] In this embodiment, the target device performs corresponding communication link switching operations based on the policy execution verification results. When the verification results show that the switching policy is effective and the new link is stable, the target device locks the current new link as the primary communication link; when the verification results show that the switching policy is ineffective or the new link is unstable, the target device automatically triggers a fallback mechanism to restore to the stable link before the switch, or switches to the next alternative link, thereby avoiding the impact of frequent link oscillations on services. Through a local closed-loop diagnosis and switching process, the response time for network faults is shortened, and the reliability and self-healing capability of optical storage equipment communication are improved.
[0053] This embodiment shortens network fault response time by building a complete network diagnostic and adaptive optimization closed loop locally on the device, reducing the average fault repair time from hours to seconds. It solves industry problems associated with traditional cloud-based centralized diagnostic solutions, such as network outages during cloud link failures, high misjudgment rates with single threshold judgments, and service interruptions due to blind switching. It ensures the continuity of communication and the reliability of control functions for optical storage devices even in extreme network environments. Furthermore, it reduces the computational pressure on cloud servers and network bandwidth consumption, lowering the operation and maintenance costs of the optical storage system.
[0054] In some embodiments, the multi-layer network state parameters include one or more of the following: physical layer parameters, transport layer parameters, connection establishment layer parameters, application layer parameters, and service layer parameters. Physical layer parameters include at least one of signal strength, wireless LAN signal quality, and mobile communication signal strength; Transport layer parameters include at least one of network latency, packet loss rate, and jitter; The connection establishment layer parameters include at least one of the following: domain name resolution time, Transmission Control Protocol connection establishment time, and Transport Security Layer protocol handshake time. Application layer parameters include at least one of the following: private protocol connection status and heartbeat response time; The business layer parameters include at least one of the following: data reporting success rate and control command confirmation success rate.
[0055] Immediately, a layered parameter acquisition method should be adopted to comprehensively cover the network operating status from the communication layer to the service layer, facilitating accurate subsequent location of network faults at different levels. In practical applications, any one or more layers of parameters can be selected to complete data acquisition and analysis according to the equipment's operational needs.
[0056] Among them, physical layer parameters are used to characterize the basic quality of the underlying signal transmission, mainly covering signal strength, wireless LAN signal quality, mobile communication signal strength, etc., which intuitively reflect the basic signal conditions of wireless and wired communication links.
[0057] Transport layer parameters are used to characterize the stability and timeliness of data transmission, specifically including network latency, packet loss rate, and jitter, which are used to evaluate the transmission speed, integrity, and fluctuation of data packets during transmission.
[0058] Connection establishment layer parameters are used to characterize the operational efficiency of the network link initialization phase, including domain name resolution time, transmission control protocol connection establishment time, and transport layer handshake time, corresponding to the time consumption indicators of network address resolution, link establishment, and encrypted interaction.
[0059] Application layer parameters are used to characterize the business protocol interaction status between the device and the server. They mainly collect private protocol connection status and heartbeat response time to monitor whether the protocol path between the two parties is normal and whether the interaction is timely.
[0060] Service layer parameters are used to characterize the service operation performance of the equipment, with data reporting success rate and control command confirmation success rate as the main components, to measure the network's support capability for service data transmission and command interaction.
[0061] In some embodiments, such as Figure 3 As shown, cross-layer correlation analysis is performed on the state parameters of a multi-layer network to generate network diagnostic results, including: S301, perform cross-layer correlation analysis on the state parameters of the multi-layer network to obtain the root cause identification of the fault; S302 determines the communication link switching strategy based on the fault root cause identifier, the preset link priority sorting and the fault identification threshold.
[0062] The root cause identifier includes information on the level at which the fault occurred, the type of fault, and the severity of the fault.
[0063] In some embodiments, the process of the target device performing cross-layer correlation analysis on multi-layer network state parameters to generate network diagnostic results is divided into two stages: accurate identification of fault root causes and intelligent generation of switching strategies. The fault root causes are used to identify the essence of the problem by penetrating the phenomenon from the complex fluctuations of multi-layer parameters, while the intelligent generation of switching strategies outputs targeted solutions based on the root cause information.
[0064] The target device performs cross-layer correlation analysis on multi-layer network state parameters to obtain standardized fault root cause identifiers. Unlike traditional single-threshold judgment, which can only detect network anomalies in a general way and cannot distinguish fault types and root causes, cross-layer correlation analysis establishes a causal reasoning model between parameters of different network layers, filters out occasional non-causal concurrent anomalies, and accurately locates the root cause of the fault.
[0065] The root cause identification of a fault includes three core dimensions: fault level information, fault type information, and fault severity information. The level information clarifies the specific location of the fault in the network protocol stack, defining the basic scope for solution selection; the type information further refines the specific manifestation of the fault, distinguishing different types of faults at the same level; and the severity information quantifies the impact of the fault on the business, determining the urgency of subsequent processing and the priority of resource allocation.
[0066] For example, when the physical layer signal strength is detected to be consistently below -85dBm and the transmission layer packet loss rate is simultaneously above 15%, a fault root cause identifier containing three dimensions of information: "physical layer, signal obstruction, and severe" can be generated; when the TCP connection establishment time at the connection establishment layer is detected to be above 500ms and the application layer heartbeat response time is detected to be above 3s, a fault root cause identifier containing three dimensions of information: "connection establishment layer, base station congestion, and moderate" can be generated; when the data reporting success rate at the service layer is detected to be below 80% and all other layer parameters are normal, a fault root cause identifier containing three dimensions of information: "service layer, cloud server overload, and mild" can be generated.
[0067] The target device generates a targeted communication link switching strategy based on the root cause identification of the fault, combined with a preset link priority ranking and fault identification thresholds. The root cause identification serves as the decision-making basis, determining whether link switching is necessary and which links can effectively resolve the current fault. The preset link priority ranking is the selection rule for candidate links, pre-set considering factors such as communication cost, bandwidth, stability, and latency. The fault identification thresholds can be understood as trigger thresholds for different levels of processing actions, divided into primary and secondary thresholds. The primary threshold corresponds to minor faults, triggering a temporary adjustment and observation mechanism for link priorities, while the secondary threshold corresponds to severe faults, triggering an immediate switching action.
[0068] For example, for the aforementioned severe signal obstruction fault at the physical layer, all wireless and wired communication links are first identified as valid candidate links based on the root cause of the fault. These links are then sorted according to a preset priority order: wireless LAN, 4G, 5G, and wired Ethernet. Finally, the triggering time for immediate handover is determined by combining the secondary fault identification threshold, generating a complete communication link handover strategy. For faults that cannot be resolved by switching communication links, such as connection establishment layer domain name resolution anomalies or application layer protocol version incompatibility, no communication link handover strategy will be generated. Instead, a corresponding non-link handover processing strategy will be generated, fundamentally avoiding the invalid operations and service interruptions caused by blind handover in traditional solutions. When the same fault root cause corresponds to multiple valid candidate links, the link with the highest priority will be selected as the target handover link according to the preset link priority order. When the fault severity is mild, a 30-second observation period will be initiated. If the fault does not recover on its own within the observation period, the handover operation will be executed.
[0069] This embodiment solves the problems of high misjudgment rate and inability to locate the root cause in traditional network diagnosis by adopting a fault root cause identification mechanism based on cross-layer correlation analysis; by adopting a root cause-based hierarchical strategy generation mechanism, it ensures the pertinence and effectiveness of the switching strategy, avoids service interruption caused by blind switching, and shortens the average repair time of network faults from the traditional minutes to the seconds.
[0070] In some embodiments, such as Figure 4 As shown, cross-layer correlation analysis is performed on the state parameters of a multi-layer network to obtain the root cause identifiers of the faults, including: S401 performs time-series alignment and standardization on the state parameters of the multilayer network to obtain a multilayer parameter sequence with a unified time dimension; S402, verify each layer parameter sequence in the multi-layer parameter sequence layer by layer, and identify the network layers with abnormalities and the corresponding abnormal parameters; S403, establish the correlation between abnormal parameters at different network levels, eliminate non-causal concurrent anomalies, and obtain the correlationd abnormal parameter combination; S404, Match the associated abnormal parameter combinations with the preset fault root cause association model to generate fault root cause identifiers. The preset fault root cause association model stores multiple sets of mapping relationships between abnormal parameter combinations and fault root cause identifiers. Each set of mapping relationships corresponds to an abnormal combination consisting of abnormal parameters from at least two different network layers.
[0071] In some embodiments, the process by which the target device performs cross-layer correlation analysis on multi-layer network state parameters to obtain fault root cause identifiers is executed in a progressive manner, consisting of multiple steps: time-series alignment and standardization, initial screening of single-layer anomalies, cross-layer causal correlation, and root cause model matching. These multiple steps are progressive and interconnected, gradually extracting standardized identifiers that can accurately locate the root cause of the fault from the original, multi-source, heterogeneous multi-layer network parameters. This solves the industry pain points of high misjudgment rate and inability to distinguish between causal anomalies and concurrent anomalies in traditional single-threshold judgment.
[0072] The target device first performs time-series alignment and standardization on the collected multi-layer network state parameters to obtain a multi-layer parameter sequence with a unified time dimension. Because there are significant differences in the sampling frequency and units of parameters across different network layers—physical and transport layer parameters are typically collected at second-level frequencies, while application and service layer parameters are statistically analyzed at periods of tens of seconds—direct cross-layer analysis can lead to time misalignment and incomparable units. For example, time-series alignment uses the lowest sampling frequency among all layers as a benchmark, aligning parameter sequences of different frequencies to a unified time granularity through linear interpolation or sliding window statistics; standardization maps parameters with different units to a unified numerical range of 0 to 1, eliminating the unitary differences between parameters.
[0073] For example, the physical layer signal strength is collected every 1 second, the transmission layer packet loss rate is calculated every 5 seconds, and the service layer data reporting success rate is calculated every 30 seconds. All parameters are aligned to a uniform time granularity of 5 seconds, and the signal strength range from -100dBm to -30dBm is standardized to a value of 0 to 1.
[0074] The target device verifies each parameter sequence in the multi-layer parameter sequence layer by layer, identifying the network layers with anomalies and their corresponding abnormal parameters. The parameters are verified sequentially from bottom to top: physical layer, transport layer, connection establishment layer, application layer, and service layer. Each parameter is compared with its corresponding fault identification threshold, and parameters exceeding the threshold and their corresponding network layers are marked. The fault identification thresholds are divided into primary and secondary thresholds; exceeding the primary threshold is marked as a minor anomaly, and exceeding the secondary threshold is marked as a severe anomaly.
[0075] For example, when checking the parameters of each layer in sequence, if the physical layer signal strength is detected to be lower than the secondary threshold of -85dBm, the physical layer is marked as an abnormal layer and the corresponding signal strength is an abnormal parameter; at the same time, if the transport layer packet loss rate is detected to be higher than the secondary threshold of 15%, the transport layer is marked as an abnormal layer and the corresponding packet loss rate is an abnormal parameter; the parameters of the other layers are all within the normal range and are not marked as abnormal.
[0076] The target device establishes correlations between abnormal parameters at different network layers, excluding non-causal concurrent anomalies, and obtains correlated combinations of abnormal parameters. This is the core step of cross-layer correlation analysis, capable of distinguishing between truly causal combinations of anomalies and accidental concurrent anomalies. Through temporal correlation analysis and causal inference algorithms, the chronological order and correlation strength between different abnormal parameters are determined. Only when one anomaly occurs earlier than another, and their correlation coefficient exceeds a preset threshold, is a causal relationship established.
[0077] For example, when both physical layer signal strength anomalies and transport layer packet loss rate anomalies are detected simultaneously, analysis reveals that the signal strength decrease precedes the packet loss rate increase by two time periods, and the correlation coefficient between the two reaches 0.92, indicating a causal relationship between them. However, if both transport layer packet loss rate anomalies and service layer data reporting success rate anomalies are detected simultaneously, but the service layer anomaly occurs before the transport layer anomaly, and the correlation coefficient is only 0.23, then the two are determined to be non-causal concurrent anomalies, excluding the influence of the service layer anomaly on the determination of the root cause of this fault.
[0078] The target device matches the associated abnormal parameter combinations with a preset fault root cause association model to generate a fault root cause identifier. The preset fault root cause association model pre-stores multiple sets of mapping relationships between abnormal parameter combinations and fault root cause identifiers. The characteristic is that each mapping relationship corresponds to an abnormal combination consisting of abnormal parameters from at least two different network layers. This is the essential difference between this solution and traditional single-parameter fault diagnosis solutions.
[0079] For example, the preset model stores multiple mapping relationships, such as "abnormal physical layer signal strength + abnormal transmission layer packet loss rate" corresponding to "wireless signal obstruction," "abnormal transmission layer jitter + abnormal application layer heartbeat response" corresponding to "base station congestion," and "abnormal connection establishment layer TCP connection + abnormal service layer data reporting" corresponding to "router port failure." The associated combination of "abnormal physical layer signal strength + abnormal transmission layer packet loss rate" is matched against the model. If a successful match identifies "wireless signal obstruction" as the root cause, a root cause identifier is generated, containing information in three dimensions: "physical layer," "signal obstruction," and "severity." If a combination of abnormal parameters contains only a single-level abnormal parameter, the model will not match any root cause and can be marked as an anomaly to be confirmed.
[0080] This embodiment achieves accurate location of network fault root causes through a progressive cross-layer correlation analysis process, increasing the accuracy of fault root cause identification from about 60% in the traditional single threshold scheme to over 95%, and reducing the false positive rate and false negative rate, providing a reliable basis for the subsequent generation of targeted communication link switching strategies.
[0081] In some embodiments, the parameter sequence of each layer in the multi-layer parameter sequence is verified layer by layer to identify the network layers with abnormalities and the corresponding abnormal parameters, including: The network state parameters in each layer of the multi-layer parameter sequence are checked layer by layer to see if they exceed the corresponding fault identification threshold, and the abnormal network layers and their corresponding abnormal parameters are identified. The fault identification threshold includes a primary threshold and a secondary threshold. The primary threshold is used to trigger the adjustment of link priority sorting, and the secondary threshold is used to trigger the output communication link switching strategy.
[0082] In some embodiments, the process of the target device verifying multi-layer parameter sequences and identifying anomalies layer by layer is a prerequisite for cross-layer correlation analysis. By setting hierarchical fault identification thresholds for different parameters at each network layer, fine-grained hierarchical detection of network anomalies can be achieved. This not only enables timely detection of serious faults and triggers emergency handling, but also avoids misjudgments and unnecessary link switching caused by instantaneous network fluctuations.
[0083] The target device verifies the network status parameters of each layer in the multi-layer parameter sequence, following the network protocol stack from bottom to top. Each parameter is compared with its corresponding fault identification threshold to identify the network layer with anomalies and its corresponding abnormal parameters. Different parameters at different network layers have independent fault identification thresholds to suit the characteristics of parameters at different layers and their impact on services.
[0084] For example, the fault identification threshold is uniformly divided into two levels: Level 1 threshold and Level 2 threshold. Level 1 threshold is a threshold for minor anomalies, used to trigger a temporary adjustment and observation mechanism for link priority; Level 2 threshold is a threshold for severe anomalies, used to trigger the output of communication link switching strategies.
[0085] For example, the primary threshold for physical layer signal strength is set to -75dBm, and the secondary threshold is set to -85dBm; the primary threshold for transport layer packet loss rate is set to 5%, and the secondary threshold is set to 15%; the primary threshold for transport layer jitter is set to 50ms, and the secondary threshold is set to 100ms; the primary threshold for TCP connection establishment time at the connection establishment layer is set to 200ms, and the secondary threshold is set to 500ms; the primary threshold for application layer heartbeat response time is set to 2s, and the secondary threshold is set to 5s; the primary threshold for business layer data reporting success rate is set to 95%, and the secondary threshold is set to 80%. When the value of a parameter exceeds the corresponding primary threshold but does not exceed the secondary threshold, it is marked as a minor anomaly; when the value of a parameter exceeds the corresponding secondary threshold, it is marked as a severe anomaly.
[0086] When a parameter is detected to exceed the first-level threshold, link switching is not immediately triggered. Instead, a temporary link priority adjustment is performed, temporarily lowering the priority of the current link and initiating a pre-defined observation period. During this period, the parameter's trend is continuously monitored. If the parameter falls back below the first-level threshold during the observation period, the link's original priority is restored, and subsequent processing is canceled. If the parameter continues to deteriorate and exceeds the second-level threshold during the observation period, the communication link switching strategy generation process is immediately triggered. This tiered processing mechanism effectively filters out misjudgments caused by instantaneous network fluctuations, preventing frequent link oscillations from impacting services.
[0087] When a parameter is detected to exceed the secondary threshold, a serious network fault is immediately identified at that level. A cross-layer correlation analysis process is then initiated, combining abnormal parameters from other layers to pinpoint the root cause of the fault and generate a corresponding communication link switching strategy. For example, if the physical layer signal strength is detected to be below the secondary threshold of -85dBm, the parameter status of other layers such as the transport layer and connection establishment layer can be immediately checked to confirm whether there are any correlation anomalies, thereby generating a targeted switching strategy.
[0088] This embodiment employs a tiered fault identification threshold mechanism to achieve refined tiered processing of network anomalies. This ensures rapid response to severe network faults while effectively preventing erroneous handovers caused by momentary fluctuations, thus improving the stability and reliability of network communication. Furthermore, the independent setting of thresholds for different level parameters more accurately reflects the operational status of different network components, providing reliable foundational data for subsequent fault root cause localization.
[0089] In some embodiments, such as Figure 5 As shown, based on the fault root cause identifier, preset link priority ranking, and fault identification threshold, the communication link switching strategy is determined, including: S501, based on the root cause identifier of the fault, determines whether the target device needs to perform a communication link switch; S502, if the target device needs to perform communication link switching, then based on the fault level and fault type corresponding to the fault root cause identifier, the candidate communication link is determined, and based on the fault identification threshold, the triggering time for communication link switching is determined. S503: Sort the candidate communication links according to the link priority and determine the target switching link. The target switching link includes the primary link and the alternative link. S504: Generate a communication link switching strategy based on the triggering time of the communication link switching and the target switching link.
[0090] In some embodiments, the process by which the target device generates a communication link switching strategy based on the root cause identifier of the fault is executed in four logically progressive steps: decision on whether to switch, determination of candidate links and triggering timing, ranking of target links, and generation of a complete strategy. This process changes the blind logic of "switching according to a fixed priority as long as the parameters exceed the limit" in traditional solutions, and realizes accurate decision-making based on the root cause of the fault, fundamentally avoiding invalid switching and service interruption.
[0091] The target device first determines whether a communication link switch is needed based on the root cause identifier. This is the first decision-making step in the entire policy generation process. Only when the root cause of the fault is one that can be resolved by switching the communication link will the subsequent link selection process begin. For faults that cannot be resolved by switching the communication link, such as connection establishment layer domain name resolution anomalies, application layer protocol version incompatibility, or business layer cloud server overload, it will be directly determined that no link switch is needed, and a corresponding non-link switching handling policy will be generated instead.
[0092] For example, when the root cause of the fault is identified as "physical layer, signal obstruction, severe", it is determined that a link switch needs to be performed; when the root cause of the fault is identified as "connection establishment layer, DNS server failure, moderate", it is determined that no link switch needs to be performed, and a processing strategy for switching the DNS server is generated.
[0093] If a communication link switchover is deemed necessary, the target device will filter candidate communication links that can effectively resolve the current fault based on the fault level and type corresponding to the root cause identifier, and determine the triggering time for the communication link switchover in conjunction with the fault identification threshold. The selection of candidate links strictly follows the "fault-solution" correspondence; only links that can cover the scope of the current fault's impact will be included in the candidate set. For example, for a physical layer wireless signal obstruction fault, all wireless and wired communication links are valid candidate links; for a transmission layer base station congestion fault of a certain operator, only mobile communication links and wired Ethernet links of other operators are valid candidate links.
[0094] It should be understood that in the embodiments of this application, the switching triggering time is determined by the level of the fault identification threshold. For example, when the fault severity corresponds to the first-level threshold, the triggering time is set to execute the switching if the fault is not recovered after 30 seconds of observation; when the fault severity corresponds to the second-level threshold, the triggering time is set to execute the switching immediately.
[0095] The target device prioritizes the selected candidate communication links according to a preset link priority, determining a target handover link set including the primary link and backup links. The preset link priority ranking comprehensively considers factors such as communication cost, bandwidth, stability, and latency, with the default priority from highest to lowest being: wireless LAN communication links, 4G mobile communication network links, 5G mobile communication network links, and wired Ethernet links. After prioritization, the candidate link with the highest priority is determined as the primary link, and the remaining candidate links are selected as backup links in order of priority. Even if the candidate link set contains only one link, this ranking process remains valid, and that link will naturally become the sole primary link.
[0096] The target device generates a complete communication link switching strategy based on the determined communication link switching triggering time and the target switching link set. The switching strategy explicitly includes key information such as the specific time point of switching execution, the primary link identifier, the order of alternative links, the fallback triggering conditions, and the length of the verification window period, providing a complete execution basis for the device to perform the switching operation.
[0097] For example, for a severe signal blockage fault at the physical layer, the handover strategy that can be generated is: immediately execute the handover operation, with the primary link being a fourth-generation mobile communication network link, and the alternative links being a fifth-generation mobile communication network link and a wired Ethernet link. The fallback trigger condition is that the new link exceeds the parameter limit more than 3 times within the verification window period, and the verification window period is 60 seconds.
[0098] This embodiment ensures that each link switch is specific and necessary through a hierarchical decision-making mechanism based on the root cause of the fault, increasing the effectiveness of link switching from less than 50% in traditional solutions to over 90%, and significantly reducing the service interruption time caused by invalid switches, thereby significantly improving the stability and reliability of the optical storage equipment communication system.
[0099] In some embodiments, candidate communication links are determined based on the fault level and fault type corresponding to the root cause identifier, including: Based on the fault level and fault type corresponding to the root cause identifier, select available communication links that can resolve the fault. Based on the target device’s historical handover records, candidate communication links are determined from the available communication links.
[0100] In some embodiments, the process of determining candidate communication links by the target device combines initial screening of root causes of faults with optimization based on historical experience. This method ensures that candidate links can fundamentally resolve the current fault, while also avoiding known unstable links based on the actual operating history of the device. It solves the problems of candidate links including invalid links and repeated switching to faulty links in traditional solutions, significantly improving the success rate of link switching.
[0101] The target device first filters out available communication links that can resolve the current fault based on the fault level and type corresponding to the root cause identifier. It should be understood that the filtering process strictly follows the correspondence between faults and solutions; only links that can cover the scope of the current fault's impact are included in the set of available links. For example, for physical layer faults such as signal obstruction and wireless interference affecting all wireless communications, all wireless communication links and wired Ethernet links are available links. As another example, for a transmission layer fault due to congestion at a specific operator's base station, only mobile communication links from other operators and wired Ethernet links are available links. For a connection establishment layer local router port fault, only mobile communication links that do not pass through that router are available links. Furthermore, for faults that cannot be resolved by switching communication links, such as domain name resolution anomalies, application layer protocol incompatibility, and cloud server overload, no communication links will be filtered; instead, a corresponding non-link switching handling strategy will be generated.
[0102] The target device further determines the final candidate communication links from the selected available communication links based on locally stored historical handover records. For example, the historical handover records store key indicators for each link under different times, locations, and fault scenarios, such as handover success rate, average stable runtime, service recovery speed, and data reporting success rate. Links with good historical performance can be prioritized, while links that have repeatedly failed to handover or are unstable under similar fault scenarios can be excluded.
[0103] For example, when the selected available links include Wi-Fi, 4G, and 5G, if historical records show that under the current physical layer signal obstruction fault scenario, the handover success rate of the 5G link is only 60%, and the average stable running time is less than 30 minutes, while the handover success rate of the 4G link reaches 95%, and the average stable running time exceeds 24 hours, the 5G link will be temporarily excluded, and Wi-Fi and 4G links will be determined as the final candidate communication links.
[0104] For example, when the root cause of the fault is identified as transport layer, mobile base station congestion, or moderate, the A mobile communication link, B mobile communication link, and wired Ethernet link are first selected as available links based on the fault type. Then, historical handover records are checked, revealing that the wired Ethernet link has not experienced any faults in the past month, with a 100% handover success rate; the A mobile communication link has a 92% handover success rate in similar base station congestion fault scenarios; and the B mobile communication link has a handover success rate of only 78% in similar scenarios. Based on this, the wired Ethernet link and the A mobile communication link are determined as the final candidate communication links, while the B mobile communication link is excluded.
[0105] This embodiment uses a two-step method that combines initial screening of fault roots with optimization based on historical experience to ensure the effectiveness and reliability of candidate communication links. It increases the success rate of link switching from about 70% in traditional solutions to over 95%, and significantly reduces the probability of failures recurring after switching, thus significantly improving communication stability.
[0106] In some embodiments, determining the policy execution verification result corresponding to the communication link switching policy includes: Within the preset verification window, the communication link switching strategy is executed, and new multi-layer network status parameters of the target switching link are continuously collected. Cross-layer correlation analysis is performed on the new multi-layer network state parameters to generate policy execution verification results. The policy execution verification results include at least one of the following: policy validity identifier, network state stability identifier of target switching link, and jitter identifier.
[0107] In some embodiments, after the target device generates a communication link switching policy, it can enter the policy execution verification stage. This stage is a key component of the network diagnostics and adaptive optimization closed loop, used to verify the actual implementation effect of the switching policy, avoid secondary service interruptions caused by switching to an unstable link, and provide data basis for subsequent link priority adjustment, fault identification threshold optimization, and model updates.
[0108] The target device executes a communication link switching strategy within a preset verification window period and continuously collects new multi-layer network status parameters of the target switching link. The length of the verification window period can be flexibly configured according to the fault type and severity. For example, it can be set from 30 seconds to 5 minutes. A shorter window period is used for minor faults to complete the verification quickly, while a longer window period is used for severe faults to fully verify the long-term stability of the link.
[0109] After the handover operation is performed, a multi-parameter synchronous acquisition mechanism that is completely consistent with the original link is started. According to the physical layer, transport layer, connection establishment layer, application layer and service layer, all network status parameters of the target handover link are acquired at the same acquisition frequency to ensure that the verification data and diagnostic data are comparable.
[0110] For example, for a physical layer severe signal blockage fault, the verification window period is set to 60 seconds. After switching to the fourth-generation mobile communication link, real-time parameters such as signal strength, packet loss rate, and jitter are collected at a frequency of 1 second, and business layer indicators such as statistical reporting success rate and control command confirmation success rate are collected at a frequency of 30 seconds.
[0111] The target device performs cross-layer correlation analysis on the newly collected multi-layer network status parameters, generating policy execution verification results containing multi-dimensional information. Unlike traditional solutions that judge handover effectiveness based on a single service parameter, this embodiment employs the same cross-layer correlation analysis method as the fault diagnosis phase, comprehensively evaluating link operation status from multiple network layers to avoid misjudgments caused by instantaneous fluctuations in individual parameters. The policy execution verification results include policy validity indicators, network status stability indicators of the target handover link, and jitter indicators, reflecting the actual effect of the handover policy from different perspectives.
[0112] In some embodiments, the policy validity flag is used to determine whether the switching policy has achieved the expected fault repair goal. By comparing the multi-layer network state parameters before and after the switch, the original fault is evaluated to determine whether it has been completely resolved. If the parameters of all layers return to the normal range after the switch, and the service layer indicators meet the preset requirements, the policy is marked as valid; if the original fault still exists after the switch, or a new cross-layer anomaly occurs, the policy is marked as invalid.
[0113] For example, before the handover, the physical layer signal strength is -90dBm, the packet loss rate of the transport layer is 20%, and the data reporting success rate of the service layer is 70%; after the handover, the signal strength stabilizes at -65dBm, the packet loss rate drops to 2%, and the data reporting success rate increases to 99%, then the policy validity flag is marked as valid.
[0114] In some embodiments, network state stability indicators are used to determine whether the target handover link can operate stably and continuously. This is evaluated by statistically verifying the fluctuation range and number of out-of-bounds occurrences of each parameter within a specified verification window. If the fluctuation range of all parameters is within a preset allowable range and no out-of-bounds occurrences exceed the first-level fault identification threshold, the network is marked as stable; if the parameter fluctuation range is large, or if there are multiple out-of-bounds occurrences exceeding the first-level threshold, the network is marked as unstable.
[0115] For example, if the signal strength of the new link fluctuates between -60dBm and -70dBm and the packet loss rate fluctuates between 1% and 3% during the verification window period, and no parameters exceed the limits, then the network stability indicator is marked as stable.
[0116] In some embodiments, jitter flags are used to determine whether frequent transient fluctuations exist in the target switching link. This is evaluated by calculating the standard deviation and rate of change per unit time for each parameter. If the standard deviation of a parameter is below a preset threshold and the change trend is gradual, it is marked as jitter-free; if a parameter exhibits frequent large jumps and the standard deviation exceeds the preset threshold, it is marked as jitter-present. For example, if the transport layer jitter value of the new link is stable between 10ms and 20ms with a standard deviation of 3ms, the jitter flag is marked as jitter-free; if the jitter value frequently jumps between 10ms and 100ms with a standard deviation reaching 25ms, it is marked as jitter-present.
[0117] This embodiment achieves comprehensive and accurate verification of the switching strategy's effectiveness through continuous multi-layer parameter collection and cross-layer correlation analysis within the verification window, increasing the accuracy of the verification results from 75% in traditional single-parameter schemes to over 98%. The multi-dimensional verification result identifiers comprehensively reflect the actual operating status of the link, providing a reliable decision-making basis for subsequent link rollback and adaptive parameter adjustments, effectively avoiding the impact of frequent link oscillations on business operations.
[0118] In some embodiments, such as Figure 6 As shown, based on the policy execution verification results, a communication link switching operation is performed, including: S601, when there are at least two target switching links, collect the verification indicators of each target switching link during the verification window period; S602, compare the verification metrics of each target switching link, including service recovery speed, parameter stability rate and data reporting success rate. S603, based on the comparison results, mark the target switching link with the best comprehensive index as the primary link, mark the remaining target switching links as alternative links, and update the link priority ranking of each target switching link. S604 performs communication link switching operations based on the main link.
[0119] In some embodiments, when multiple candidate target switching links exist, the target device can perform a multi-link comparison and verification process. This process is a significant optimization of the traditional fixed-priority link selection mechanism, breaking through the limitation of selecting links solely based on preset rules. It can dynamically select the optimal communication path based on the actual operating performance of each link, ensuring that the best communication quality is always provided for services in complex and ever-changing network environments.
[0120] When at least two target switching links exist, the target device can collect verification metrics for each target switching link within a preset verification window. The verification window can be set, for example, from 30 seconds to 2 minutes, to ensure sufficient statistical data collection for accurate evaluation without causing prolonged service interruptions due to excessively long verification times. For example, verification metrics include three dimensions: service recovery speed, parameter stability rate, and data reporting success rate, comprehensively evaluating the link's overall performance from the perspectives of response speed, operational stability, and service support capabilities.
[0121] Among them, service recovery speed refers to the time required from the execution of link switching operation to the restoration of normal operation of core services; parameter stability rate refers to the percentage of time during which the multi-layer network status parameters of the link are within the normal range within the verification window period; and data reporting success rate refers to the percentage of service data successfully reported to the cloud by the device within the verification window period out of the total reported data. For example, when there are two target switching links, wireless LAN and 4G mobile communication, the device can switch to the two links sequentially and collect the above three indicators for each link within a 60-second verification window period to ensure the objectivity and comparability of the evaluation data.
[0122] The target device compares the verification metrics of each target switching link and calculates the comprehensive performance score of each link through weighted summation. The weights of different verification metrics can be flexibly configured according to service type and requirements. For application scenarios with high data reliability requirements, such as optical storage systems, the data reporting success rate has the highest weight, followed by the parameter stability rate, and the service recovery speed has the lowest weight. For example, the weight of data reporting success rate is set to 0.5, the weight of parameter stability rate is 0.3, and the weight of service recovery speed is 0.2. If the service recovery speed of the wireless LAN link is 3 seconds, the parameter stability rate is 92%, and the data reporting success rate is 95%, then its comprehensive score is 3×0.2+92×0.3+95×0.5=76.7; if the service recovery speed of the fourth-generation mobile communication link is 2 seconds, the parameter stability rate is 98%, and the data reporting success rate is 99%, then its comprehensive score is 2×0.2+98×0.3+99×0.5=79.3.
[0123] Based on the comparison results, the target device marks the target switching link with the best overall performance as the primary link, and marks the remaining target switching links as candidate links in descending order of their overall scores. The link priority ranking of each target switching link is updated simultaneously. This updated priority ranking is not only applied to the current communication link switching but can also be saved to the device's local history as a reference for selecting and ranking candidate links in similar failure scenarios in the future. Even if a link has a high default priority, it will be adjusted to a lower priority position if its actual overall performance is poor.
[0124] For example, in the above example, although the default priority of the wireless LAN is higher than that of the fourth-generation mobile communication, the fourth-generation mobile communication link is marked as the primary link and the wireless LAN link is marked as the alternative link because the fourth-generation mobile communication has a higher overall score. The temporary priority of the fourth-generation mobile communication link is also adjusted to the highest.
[0125] The target device performs the final communication link switching operation based on the marked main link. If multiple target switching links have the same comprehensive index score, they will be selected according to a preset default priority order. After the switch is completed, the operating status of the main link can continue to be monitored. If the main link experiences performance degradation during subsequent operation, it will automatically switch to the backup link to ensure the continuous operation of services.
[0126] This embodiment achieves dynamic optimization of link selection through comparative verification of the actual operation effect of multiple links, solves the problem that the traditional fixed priority mechanism cannot adapt to complex and ever-changing network environments, improves the rationality of link selection to more than 95%, and significantly improves the communication quality and service continuity of optical storage equipment in complex network environments.
[0127] In some embodiments, the method further includes adjusting at least one of the model parameters used for link priority ranking, fault identification threshold, and cross-layer correlation analysis based on the policy execution verification results, in the following manner: When the policy execution verification result indicates that the communication link switching policy is effective and the target switching link meets the preset stability conditions, the link priority ranking of the target switching link is increased. When the strategy execution verification result indicates that the communication link switching strategy is invalid or the target switching link does not meet the preset stability conditions, the link priority ranking of the target switching link is lowered, the link priority ranking of the stable link before switching is raised, and the fault identification threshold of the corresponding level of the fault root cause identifier is adjusted. When the strategy execution verification results indicate the presence of unidentified faults, update the model parameters used in the cross-layer correlation analysis.
[0128] In some embodiments, the target device can perform full-dimensional parameter adaptive adjustment based on the policy execution verification results. This is an evolutionary link in the entire network diagnosis and adaptive optimization closed loop, which enables continuous self-optimization based on actual operating results. This solves the defects of traditional network solutions with fixed parameters that cannot adapt to different deployment environments and network changes. As the operating time increases, the diagnostic accuracy and switching effectiveness will continue to improve.
[0129] When the policy execution verification result indicates that the communication link switching policy is effective and the target switching link meets the preset stability conditions, the target device can increase the link priority ranking of the target switching link. Priority adjustment adopts a gradient adjustment mechanism; each effective and stable switch will increase the priority of the corresponding link by one level, up to the highest priority. The adjusted priority will be saved in the device's local priority configuration table as the basis for link ranking in subsequent similar fault scenarios.
[0130] For example, when a fourth-generation mobile communication link is used as the target handover link, and the verification results after the handover show that the strategy is effective and the link is stable, its priority can be increased from the default second tier to the first tier, surpassing the default highest priority wireless LAN link. Subsequently, when a physical layer signal obstruction failure occurs again, the fourth-generation mobile communication link can be preferentially selected as the primary link.
[0131] When the policy execution verification result indicates that the communication link switching policy is invalid or the target switching link does not meet the preset stability conditions, the target device can perform three linked adjustment operations. First, the link priority ranking of the target switching link is lowered. Each invalid or unstable switch will lower the priority of the corresponding link by one level, down to the lowest priority. At the same time, the link priority ranking of the stable link before the switch is raised, restoring its priority selection right in similar fault scenarios. Finally, the fault identification threshold at the corresponding level of the fault root cause identifier is adjusted. By narrowing the threshold range, the judgment standard for similar faults in the future is improved, avoiding invalid switches caused by misjudgments. For example, when switching to a 5G mobile communication link, if the verification result shows frequent link jitter and a data reporting success rate of less than 80%, the priority of the 5G mobile communication link can be lowered from the third level to the fourth level, and the priority of the wireless LAN link before the switch can be restored from the first level to the highest priority. The secondary threshold of the physical layer signal strength can be narrowed from -85dBm to -80dBm. Subsequent switching operations will only be triggered when the signal strength is below -80dBm.
[0132] When the strategy execution verification result indicates the existence of unidentified faults, the target device can update the model parameters used in the cross-layer correlation analysis. Unidentified faults include two situations: first, the anomalous parameter combinations after correlation do not match the corresponding fault root cause in the preset fault root cause correlation model; second, the fault root cause matched by the model does not match the actual verification result. The anomalous parameter combinations, actual fault phenomena, and verification results can be added to the local model's training dataset, and the model's mapping relationship and matching weights can be updated through incremental learning.
[0133] For example, when a parameter combination is detected where the physical layer signal strength is normal, the transport layer packet loss rate is abnormal, and the application layer heartbeat response is abnormal, the preset model does not match the corresponding root cause of the fault. After verification, it is found that the fault is caused by the router firmware version being too low. The parameter combination of "normal physical layer + abnormal transport layer packet loss + abnormal application layer heartbeat" and the root cause of "abnormal router firmware" can be added to the model, and the matching weights of the relevant parameters can be adjusted so that such faults can be accurately identified in the future.
[0134] This embodiment achieves self-learning and self-evolution capabilities through a multi-dimensional adaptive parameter adjustment mechanism. The longer the running time and the more running data accumulated, the higher the accuracy of fault diagnosis and the effectiveness of link switching. It can continuously adapt to complex network conditions in different deployment environments, significantly reducing the cost and workload of manual operation and maintenance.
[0135] In some embodiments, the method further includes: When a target device is detected to be disconnected, all network links that can be successfully connected are given the highest priority. Based on the preset link priority, the network link with the highest priority is determined as the primary link, and the remaining switching links are marked as alternative links. Based on the main link, perform a communication link switching operation.
[0136] In some embodiments, the target device has a built-in cloud-based emergency response mechanism for network outages. This serves as a fallback solution to ensure the continuity of communication for optical storage devices in extreme network environments. It addresses the technical problem that traditional centralized cloud-based network diagnostic solutions result in complete device loss of control and inability to autonomously restore communication when the cloud link is lost. In cases where the cloud server is unreachable, the target device can autonomously complete available link detection, priority adjustment, link switching, and subsequent full-process network management, ensuring the normal operation of the device's control functions.
[0137] For example, the target device monitors its communication status with the cloud server in real time through a multi-dimensional heartbeat detection mechanism. When a persistent disconnection with the cloud server is detected, an emergency handling procedure is immediately triggered. The disconnection determination employs a multi-verification mechanism to avoid false triggers caused by momentary network fluctuations. Simultaneously, the heartbeat response status, data reporting success rate, and control command reception status are monitored. A cloud disconnection is only formally determined when three consecutive heartbeats fail to receive a response from the cloud, five consecutive business data reports fail, and no control commands are received. For instance, the target device sends a heartbeat to the cloud every 10 seconds. When three consecutive heartbeats time out and the most recent five data reports fail, a cloud disconnection is determined, and the emergency handling procedure is immediately initiated.
[0138] The target device immediately initiates full-link probing, testing the connectivity of all network links supported by the device. It then identifies all network links capable of successfully connecting to the cloud server and sets the priority of these connectable links to the highest priority. In emergency scenarios involving link outages, communication restoration takes precedence over all other factors such as communication cost and bandwidth; therefore, all links capable of establishing communication with the cloud have the highest selection priority. Link connectivity is tested by sending probe packets to the cloud server. Links that receive a probe response within 3 seconds are considered connectable. For example, when the cloud connection is lost, the device sequentially probes Wi-Fi, 4G, 5G, and wired Ethernet links. It finds that only 4G and 5G links can successfully connect to the cloud and immediately sets the priority of these two links to the highest.
[0139] The target device sorts the links according to a preset default priority. From all the highest-priority connectable links, it identifies the highest-priority network link as the emergency primary link, and marks the remaining connectable links as backup links in descending order of default priority. The preset default priority sorting takes into account factors such as link stability, reliability, and emergency response speed, and defaults to wired Ethernet links, fourth-generation mobile communication network links, fifth-generation mobile communication network links, and wireless LAN communication links from highest to lowest priority. For example, in the above example, the default priority of the fourth-generation mobile communication link is higher than that of the fifth-generation mobile communication link, therefore the fourth-generation mobile communication link is identified as the emergency primary link, and the fifth-generation mobile communication link is marked as a backup link.
[0140] The target device immediately performs a communication link switchover operation based on the identified emergency primary link. After the switchover is complete, the target device enters an emergency verification window period to continuously monitor the communication status of the primary link and verify whether communication has been successfully restored. During the cloud-based link outage, the target device will operate completely autonomously, completing all network diagnostics, link switching, policy verification, and parameter adaptive adjustment operations locally without relying on any instructions from the cloud server. Once communication with the cloud server is restored, the device can automatically synchronize all operational data, diagnostic results, and parameter adjustment records from the outage period to the cloud and receive the latest global parameters from the cloud server, resuming normal cloud-edge collaborative operation.
[0141] This embodiment reduces the communication recovery time after a cloud link failure from hours in traditional solutions to seconds by using an autonomous cloud link failure emergency handling mechanism on the device side. This solves the problem of network outages when the cloud link fails, ensures the basic monitoring and control functions of the optical storage equipment in extreme network environments, and significantly improves the overall reliability and security of the optical storage system.
[0142] This application provides an example of a network diagnostic processing method. Please refer to [link / reference]. Figure 7 As shown, Figure 7 A schematic flowchart of a network diagnostic processing method provided in this application is shown. This is an example and not a limitation; the method can be applied to or run on a server, such as a cloud server. The method includes: S701, Obtain the multi-layer network status parameters of the current communication link of the target device; S702 performs cross-layer correlation analysis on the status parameters of multi-layer networks and generates network diagnostic results. The network diagnostic results include fault root cause identification and communication link switching strategy. The communication link switching strategy is determined based on fault root cause identification, preset link priority ranking and fault identification threshold. S703, Determine the policy execution verification result corresponding to the communication link switching policy; S704, based on the policy execution verification results, adjust at least one of the following model parameters: link priority sorting, fault identification threshold, and cross-layer correlation analysis.
[0143] This embodiment provides a network diagnostic processing method running on the server side. Specifically, it can be executed by a proxy service engine (such as a cloud-based agent diagnostic engine) deployed on the server. In the fields of industrial IoT, optical storage cloud platform, and IoT operation and maintenance, the proxy service engine refers to the terminal access proxy service deployed on the cloud server. It is suitable for centralized network operation and maintenance and intelligent optimization of multiple edge devices. Relying on the sufficient computing power and storage capacity of the server (the proxy service engine of the server, which will not be described again below), it can realize unified monitoring of the network status of all devices, fault diagnosis, policy distribution and parameter iterative optimization, and build a closed loop of network operation and maintenance under cloud-based overall management.
[0144] In some embodiments, the server continuously acquires multi-layer network status parameters corresponding to the current communication link of each target device. The collected parameters cover one or more layers of the physical layer, transport layer, connection establishment layer, application layer, and service layer, comprehensively capturing all-dimensional network operation data of the device from the underlying signal transmission to the upper layer service interaction, providing complete data support for subsequent fault analysis.
[0145] In some embodiments, the server performs cross-layer correlation analysis on the aggregated multi-layer network state parameters to generate corresponding network diagnostic results. These results mainly include two parts: fault root cause identification and communication link switching strategy. The communication link switching strategy is determined by combining the fault root cause identification, preset link priority ranking, and fault identification threshold. Leveraging cross-layer correlation analysis capabilities, the server can penetrate surface-level network anomalies, accurately locate the fault occurrence level, fault type, and fault severity, and then match an appropriate link switching scheme using preset rules, avoiding blindly issuing switching commands.
[0146] In some embodiments, the server further determines the policy execution verification result corresponding to the communication link switching policy. After the target device performs the link switching operation, it can send the network operation data of the new link back to the server. The server combines the sent-back data to comprehensively evaluate the actual implementation effect of the switching policy and forms a verification result that includes information such as policy effectiveness, link stability, and link jitter, providing a complete feedback on the actual status of this policy execution.
[0147] Based on the obtained policy execution verification results, the server performs targeted adaptive optimization adjustments, and can choose to update any one or more of the following: link priority ranking, fault identification threshold, and model parameters used for cross-layer correlation analysis. When the switching policy is implemented effectively, the server optimizes the priority of the corresponding link; when the policy execution fails or the link operation is abnormal, the fault identification threshold and link priority are adjusted simultaneously; if a new type of fault is encountered that the existing model cannot identify, the model parameters for cross-layer correlation analysis are iteratively updated.
[0148] This method, by leveraging a centralized server processing model, can manage the network operation status of multiple target devices in a unified manner, and complete fault diagnosis, policy control, and model self-optimization in a unified way. This not only reduces the computing load of edge devices, but also continuously optimizes the entire set of network diagnosis and scheduling rules based on the operation data of devices across the entire domain, thereby improving the network operation and maintenance capabilities and fault handling level of the optical storage system.
[0149] In some embodiments, cross-layer correlation analysis is performed on the state parameters of multi-layer networks to generate network diagnostic results, including: Cross-layer correlation analysis is performed on the state parameters of multi-layer networks to obtain fault root cause identifiers; Based on the root cause identification, link priority ranking, and fault identification threshold, determine the communication link switching strategy; The root cause identifier includes information on the level at which the fault occurred, the type of fault, and the severity of the fault.
[0150] In some embodiments, the server performs cross-layer correlation analysis on the collected multi-layer network status parameters and generates network diagnostic results, which is divided into two coherent steps: fault root cause identification and switching strategy formulation. The entire process of analysis and decision-making is completed by relying on the centralized computing power of the cloud.
[0151] The server first performs cross-layer correlation analysis on multi-layer network state parameters to generate standardized root cause identifiers. This analysis method is no longer limited to parameter judgment at a single network layer, but combines the linkage and change patterns of data at different layers to distinguish between occasional parameter anomalies and systemic faults with causal relationships, accurately locating the essence of network anomalies. For example, the generated root cause identifier integrates three types of key information: fault occurrence level information, fault type information, and fault severity information. Among them, the level information clarifies the network architecture location of the problem, the type information defines the specific manifestation of the fault, and the severity information classifies the impact of the fault on normal device communication and service operation.
[0152] The server then combines the obtained root cause identifiers, preset link priority rankings, and fault identification thresholds to determine the communication link switching strategy. The root cause identifier serves as the basis for strategy formulation, determining whether the current fault is suitable for resolution by switching communication links and narrowing down the range of feasible alternative links. The link priority ranking serves as the selection rule for alternative links, defining the order of selection for different links based on long-term operational experience. The fault identification threshold serves as the action trigger threshold, distinguishing the appropriate handling timing for different fault levels. The communication link switching strategy generated by the cooperation of these three elements is highly targeted, effectively avoiding meaningless link switching behavior and ensuring stable operation of equipment communication and services.
[0153] It should be noted that the above analysis and decision-making process runs on the server side, which can simultaneously connect to the network data of multiple target devices, rely on cloud computing power to complete batch fault diagnosis and policy generation, and realize centralized management and unified operation and maintenance of the network status of all terminal devices in the region.
[0154] In some embodiments, cross-layer correlation analysis is performed on the state parameters of the multi-layer network to obtain fault root cause identifiers, including: The state parameters of the multilayer network are time-series aligned and standardized to obtain a multilayer parameter sequence with a unified time dimension; The parameter sequence of each layer in the multi-layer parameter sequence is verified layer by layer to identify the network layers with abnormalities and the corresponding abnormal parameters. Establish the correlation between abnormal parameters at different network levels, eliminate non-causal concurrent anomalies, and obtain the correlationd abnormal parameter combination; The associated abnormal parameter combinations are matched with the preset fault root cause association model to generate fault root cause identifiers.
[0155] The root cause association model stores multiple sets of abnormal parameter combinations and root cause identifiers. Each set of mappings corresponds to an abnormal combination consisting of abnormal parameters from at least two different network layers.
[0156] In some embodiments, the server completes cross-layer correlation analysis and generates fault root cause identifiers through a multi-step progressive processing flow. Relying on cloud computing power, it performs unified calculations on network parameters uploaded by multiple target devices, gradually completing data regularization, anomaly screening, causal identification and root cause matching, thereby achieving accurate location of network faults.
[0157] The server first performs time-series alignment and standardization on the received multi-layer network state parameters to form a multi-layer parameter sequence with a unified time dimension. Different network layers and different target devices have different parameter acquisition frequencies and data units. Time-series alignment can organize scattered parameter data to the same time granularity and eliminate time misalignment problems; standardization process unifies the numerical range of various parameters and avoids interference caused by inconsistent units to subsequent analysis.
[0158] The server performs verification on multiple parameter sequences according to network layers, checking network status parameters layer by layer to identify the network layers experiencing anomalies and the corresponding abnormal parameters. This step completes the initial screening of anomalies at a single layer, visually marking the specific location and parameter items where network operation problems occur, and completing the initial capture of fault phenomena.
[0159] After identifying single-level anomalous parameters, the server further establishes correlations between anomalous parameters at different network levels, distinguishing between anomalous combinations with inherent causal relationships and concurrent anomalies that only occur synchronously. Irrelevant, accidental anomalous data is eliminated, ultimately yielding anomalous parameter combinations with causal logic. This step effectively filters out interfering information, preventing the misclassification of transient, irrelevant parameter fluctuations as systemic faults and improving the accuracy of fault analysis.
[0160] The server processes the abnormal parameter combinations and matches them against a locally pre-defined fault root cause association model to generate a fault root cause identifier. This model pre-stores a large number of mapping relationships between abnormal parameter combinations and fault root cause identifiers, with each mapping consisting of at least two abnormal parameter combinations from different network layers, aligning with the design logic of cross-layer analysis. The server relies on mature mapping rules to perform intelligent matching, combining the characteristics of abnormal combinations to determine the nature of the fault, and outputting a fault root cause identifier containing information such as layer, type, and severity, providing a reliable basis for subsequent communication link switching strategies. The centralized cloud processing mode also aggregates abnormal data from all network devices to continuously optimize the model and improve overall fault identification capabilities.
[0161] In some embodiments, a communication link switching strategy is determined based on the root cause identifier, link priority ranking, and fault identification threshold, including: Based on the root cause identifier, determine whether the target device needs to perform a communication link switch. If the target device needs to perform a communication link switch, then candidate communication links are determined based on the fault level and fault type corresponding to the fault root cause identifier, and the triggering time for communication link switch is determined based on the preset fault identification threshold. Candidate communication links are sorted according to link priority to determine the target switching link, which includes the primary link and alternative links. Based on the triggering time of the communication link switching and the target switching link, a communication link switching strategy is generated.
[0162] In some embodiments, the server combines fault root cause identification, link priority ranking, and fault identification threshold to formulate communication link switching strategies. Specifically, it can determine the fault layer by layer and execute it step by step. Relying on the centralized decision-making capabilities of the cloud, it outputs standardized and implementable execution strategies for the target device, thereby achieving standardization and precision in fault handling.
[0163] The server first determines whether the target device needs to perform a communication link switching operation based on the fault root cause identifier obtained from the parsing. The fault root cause identifier clarifies the fault level and specific fault type. The server uses this as a basis to distinguish the fault category. For problems that can be solved by switching links, such as abnormal wireless signals or transmission link congestion, the server proceeds to the subsequent policy development process. For faults that cannot be repaired by switching links, such as domain name resolution errors or application protocol incompatibility, the server directly terminates the link switching process and generates a corresponding handling plan to avoid issuing invalid commands from the source.
[0164] After determining that a link switch is necessary, the server combines the fault level and type corresponding to the root cause identifier to filter out candidate communication links that can effectively eliminate the current fault. It then defines the specific triggering time for the communication link switch based on a preset fault identification threshold. Different fault levels and types correspond to different compatible communication links, thus defining a reasonable range of candidate links. The fault identification threshold categorizes the severity of the fault, and different execution times are set accordingly, such as immediate switchover or switchover after a delay for observation, balancing fault handling efficiency with the need to prevent false positives.
[0165] The server reorders the selected candidate communication links according to a preset link priority ranking rule, dividing them into a primary link and a backup link set. The preset priority is set by taking into account factors such as link stability, transmission performance, and usage cost. The link with the highest priority after ranking is designated as the primary link, and the remaining candidate links are designated as backup links in order, providing multi-level link backup protection for the main device.
[0166] The server integrates the determined handover triggering time and target handover link information to generate a complete communication link handover policy. This policy includes complete execution elements such as execution time, primary / backup link selection, and exception fallback rules. The server then distributes the policy to the corresponding target devices, guiding them to complete the link handover action, thus forming a complete business process from cloud-based diagnostics and policy generation to terminal execution.
[0167] In some embodiments, based on the policy execution verification results, at least one of the model parameters used in link priority ranking, fault identification threshold, and cross-layer correlation analysis is adjusted, including: When the policy execution verification result indicates that the communication link switching policy is effective and the target switching link meets the preset stability conditions, the link priority ranking of the target switching link is increased. When the strategy execution verification result indicates that the communication link switching strategy is invalid or the target switching link does not meet the preset stability conditions, the link priority ranking of the target switching link is lowered, the link priority ranking of the stable link before switching is raised, and the fault identification threshold of the corresponding level of the fault root cause identifier is adjusted. When the strategy execution verification results indicate the presence of unidentified faults, update the model parameters used in the cross-layer correlation analysis.
[0168] In some embodiments, the server performs verification based on the strategy returned by the target device, and adaptively adjusts the link priority ranking, fault identification threshold, and cross-layer correlation analysis model parameters. It relies on the operation data of all devices to complete rule iteration, so that the entire network diagnostic system can continuously adapt to changes in the field network environment and form a cloud-based self-optimization closed loop.
[0169] When verification results indicate that the communication link switching strategy is effective and the target switching link's operating status meets preset stability requirements, the server will increase the priority ranking of that target switching link. This adjustment will be synchronously recorded in the system configuration. When formulating strategies for similar faults in the future, this high-performing link will receive a higher selection weight and will be preferentially recommended for use by the target device.
[0170] When the verification results indicate that the switching strategy has not achieved the expected results, or the target switching link is unstable, the server will perform several coordinated adjustment operations. On the one hand, the priority of the current target switching link will be lowered, reducing its probability of being selected subsequently; on the other hand, the priority of the original stable link before the switch will be raised, restoring its priority usage rights. At the same time, for the network layer corresponding to the root cause of the fault, the fault identification threshold of that layer will be adjusted, the anomaly judgment criteria will be optimized, and the recurrence of similar problems will be reduced.
[0171] When unidentified faults that cannot be matched by existing rules are found during the verification process, the server updates the model parameters on which the cross-layer correlation analysis is based. The server incorporates the abnormal parameter combinations and actual fault phenomena collected this time into the model samples, optimizes the parameter correlation logic and matching rules, expands the fault identification capability, enables the model to identify new network faults, and gradually improves the overall accuracy of fault diagnosis across the entire domain.
[0172] The following provides several specific embodiments to illustrate the network diagnostic processing method provided in this application. Specifically, the target edge device is the edge device of the photovoltaic-storage system, such as a photovoltaic inverter. In the first embodiment, this embodiment provides a scenario of abnormal Wi-Fi communication for edge devices in a photovoltaic storage system, such as a scenario where the Wi-Fi link fails to switch to a 4G link. The edge device completes standardized layered parameter collection throughout the process according to the physical layer, transport layer, connection establishment layer, application layer, and service layer. Relying on the cross-layer correlation analysis capability of the cloud-based Agent diagnostic engine, it achieves accurate fault location and policy distribution. With the verification, rollback, and cloud-based adaptive parameter tuning mechanism on the edge device side, a complete self-healing closed loop is completed.
[0173] Edge devices typically operate on Wi-Fi communication links, and their data acquisition modules continuously collect multi-layered network status parameters. Specifically, the physical layer collects Wi-Fi RSSI signal strength and Wi-Fi signal quality; the transport layer collects network latency, packet loss rate, and data jitter; the connection establishment layer collects DNS resolution time, TCP connection establishment time, and TLS handshake time; the application layer collects the edge device's private protocol connection status and cloud heartbeat response time; and the service layer collects service data reporting success rate and control command ACK confirmation success rate.
[0174] When issues such as wireless interference and signal obstruction occur on-site, edge devices collect multi-layer parameter synchronization anomalies: Physical layer Wi-Fi signal strength continuously weakens and signal quality deteriorates; transmission layer packet loss rate continuously increases, and latency and jitter exceed the first-level fault identification threshold; connection establishment layer TCP connection establishment time significantly increases; application layer heartbeat response timeout and protocol connection status is unstable; business layer data reporting success rate and command confirmation success rate drop significantly. The cloud-based Agent diagnostic engine performs time-series alignment, standardization, and cross-layer causal correlation analysis on multi-time-series multi-layer parameter sequences, eliminating interference from instantaneous fluctuations of single parameters. It determines that the current anomaly is not a single point of accidental occurrence, but a complex multi-layer fault due to insufficient overall stability of the Wi-Fi link, generating corresponding severe fault root cause identifiers.
[0175] Based on the root cause identification of the fault, the preset secondary fault threshold, and the global link priority ranking, the cloud determines that the current Wi-Fi link cannot guarantee the stable operation of the service, generates a link switching policy, determines to switch the main communication link from Wi-Fi to 4G mobile communication link, and simultaneously marks the switching trigger time, the order of the main and backup links, and the verification window period configuration.
[0176] After receiving the policy and completing the link switch, the edge device immediately initiates the policy execution verification process, entering a preset verification window period. During this window, it continuously collects network parameters across all layers of the 4G link, covering core indicators such as 4G physical signal strength, transmission latency, packet loss rate, jitter, TCP establishment time, heartbeat response, and service reporting success rate. It also performs cross-layer stability verification on parameters from multiple consecutive sampling periods. If all layer parameters of the 4G link fall back to normal threshold ranges within the window period, and packet loss rate, latency, and jitter are stable and controllable, and the service reporting success rate recovers to standard levels without parameter out-of-bounds or drastic fluctuations, then the policy execution verification result is marked as policy effective, link stable, and jitter-free.
[0177] If, during the verification window, the 4G link still experiences issues such as continuous parameter fluctuations, frequent connection establishment failures, and unrecoverable service metrics, the handover strategy is deemed invalid. The edge device automatically triggers the link fallback mechanism to restore the stable link before the handover or switch to another alternative link in order of priority, effectively suppressing frequent link oscillations.
[0178] Finally, the edge device sends the complete policy execution verification results back to the cloud-based Agent diagnostic engine. Based on valid handover results, the cloud increases the priority ranking of 4G links in the current scenario; based on invalid handover results, it decreases the priority of the corresponding links and corrects the fault identification threshold of the corresponding fault level; if a new anomaly with model mismatch occurs, the cross-layer correlation analysis model parameters are updated synchronously to achieve adaptive iterative optimization of the global policy.
[0179] This embodiment achieves seamless link self-healing in Wi-Fi failure scenarios through multi-layer parameter cross-layer diagnosis, precise strategy switching, window period closed-loop verification, and cloud-based adaptive parameter tuning, ensuring the continuous and stable data acquisition, status reporting, and remote control services of optical storage edge devices.
[0180] In the second embodiment, this embodiment provides a scenario for the normal operation of 4G and self-healing recovery of Wi-Fi link failures in optical storage edge devices. Based on the cross-layer comprehensive evaluation of multi-layer network parameters, it realizes the optimal switching of multi-link quality, and avoids frequent link switching through jitter suppression verification mechanism. With the iterative optimization of cloud parameters, it improves the rationality of long-term link scheduling.
[0181] The initial main link of the edge device is a 4G mobile communication link. The edge device side acquisition module continuously collects the dual-link multi-layer network status parameters of the 4G link and the backup Wi-Fi link in layers, covering physical layer signal quality, transmission layer latency, packet loss and jitter, connection establishment layer handshake time of various types, application layer heartbeat and protocol status, and service layer reporting and command confirmation success rate.
[0182] As the on-site wireless environment recovered, the backup Wi-Fi link's multi-layer parameters were simultaneously repaired: the physical layer's Wi-Fi signal strength and quality returned to the excellent range; the transmission layer's latency, packet loss rate, and jitter all fell back to normal threshold ranges; the connection establishment layer's DNS resolution, TCP connection, and TLS handshake time were significantly reduced; the application layer's heartbeat response was timely, and private protocol connections were stable; the service layer's data reporting success rate and control command ACK success rate consistently met standards. The cloud-based agent diagnostic engine, through cross-layer correlation analysis, determined that the Wi-Fi link had completely recovered to a stable and usable state, and through multi-link indicator comparison, confirmed that the overall Wi-Fi communication quality was superior to the current 4G main link.
[0183] The cloud combines preset link priority ranking and multi-link verification indicators to generate an optimal switching strategy, instructing edge devices to smoothly switch the main communication link from 4G link to Wi-Fi link, and specifying the timing of the switch and subsequent verification rules.
[0184] After the edge device completes the link switchback, it immediately enters the verification window period, continuously collecting network parameters at all levels of the newly switched Wi-Fi link to conduct closed-loop stability verification. If, within multiple consecutive sampling periods, the parameters at each level of the Wi-Fi link remain stable and within acceptable limits, and the service indicators remain reliable, the switchback strategy is deemed effective, and the current Wi-Fi main link status is solidified.
[0185] If intermittent fluctuations, insufficient parameter recovery, or unstable service metrics are found in the Wi-Fi link during the verification window, the system will activate the jitter suppression and anti-oscillation mechanism, not confirm the effectiveness of this switchback, automatically maintain the original 4G main link status, and suspend the link switching to avoid service jitter and communication interruption caused by frequent switching between 4G and Wi-Fi.
[0186] The edge device transmits the complete results of this back-switch verification back to the cloud. The cloud adaptively optimizes the link priority ranking, link switching trigger threshold, and cross-layer correlation model parameters based on the verification results. For unstable Wi-Fi recovery scenarios, the Wi-Fi back-switch judgment conditions are appropriately increased to avoid subsequent false back-switches; for stable back-switch scenarios, the link selection weights are optimized to improve the accuracy of subsequent multi-link scheduling.
[0187] This embodiment realizes intelligent optimal switching back after link failure recovery. Through hierarchical and multi-dimensional link quality verification and jitter suppression mechanisms, it takes into account both communication economy and stability, and effectively improves the smoothness of long-term communication operation of the optical storage system.
[0188] In the third embodiment, this embodiment provides a fallback scenario for cloud-edge collaboration anomalies. In the extreme case where the cloud normally issues a switching strategy but the target link suddenly breaks down after the switch, the edge device relies on a pre-set emergency mechanism for link breakage and fallback rules to achieve local autonomous recovery without cloud involvement, thus building a dual-insurance communication guarantee system of cloud decision-making + local fallback.
[0189] The edge device initially runs stably on the 4G main link. The cloud-based Agent diagnostic engine continuously monitors the multi-layer parameters of the dual links and determines that the backup Wi-Fi link has returned to normal and has better overall quality. It then issues a link switching policy, instructing the edge device to switch the main link from 4G to Wi-Fi.
[0190] After performing a handover operation, the edge device briefly switches to Wi-Fi operation and immediately initiates verification window monitoring. During the verification window, the edge device detects a sudden and extreme abnormal disconnection of the Wi-Fi link, specifically manifested as: complete interruption of the physical layer Wi-Fi signal, continuous packet loss at the transport layer, failure to establish a TCP connection at the connection establishment layer, application layer heartbeat timeout, disconnection of the private protocol connection, and complete failure of data reporting and command confirmation at the service layer.
[0191] At this point, with no effective intervention available in the cloud, the edge device triggers its locally pre-configured emergency recovery strategy for network outages. This local strategy has a higher execution priority than the regular strategies subsequently issued by the cloud. The edge device retrieves the locally cached link priority ranking and fault rollback rules, and autonomously and quickly rolls back the communication link from the faulty Wi-Fi link to a 4G link, ensuring that core services such as data acquisition, remote reporting, and scheduling command interaction of the optical storage edge device are not interrupted.
[0192] After the local handover is completed, the edge device immediately verifies the multi-layer network parameters of the 4G link. If the parameters of each layer of the 4G link are stable and meet the standards, and the service returns to normal, the local fallback self-healing strategy is confirmed to be effective. The edge device will record the abnormal event of this "cloud handover failure + local self-rollback", the full process parameters and verification results completely and send them back to the cloud Agent diagnostic engine.
[0193] Based on this abnormal sample, the cloud platform completed adaptive optimization iterations and made targeted adjustments to the strategy delivery logic and judgment rules: reducing the switching priority of Wi-Fi links in the current scenario, tightening the Wi-Fi link recovery judgment threshold, optimizing the link switching verification window duration, and updating the abnormal matching weight of the cross-layer fault identification model to avoid the problem of instantaneous link loss after switching recurring in the future.
[0194] This embodiment compensates for the blind spot of instantaneous link failure after the cloud policy is issued by the local backup and self-healing mechanism of the edge device, forming a two-layer protection architecture of cloud intelligent decision-making and local extreme backup, which greatly improves the communication reliability, business continuity and risk resistance of the optical storage system in complex and volatile network environments.
[0195] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0196] It should be understood that the technical features of the various embodiments described above in this application can be combined arbitrarily without conflict. For the sake of brevity, this specification does not describe all possible combinations, but as long as these combinations do not violate the technical spirit of this application, they should all be considered within the scope of this application. Based on the content disclosed in this application, those skilled in the art can reasonably combine, delete, or replace the technical features of the above embodiments according to actual needs, and such modifications and variations all fall within the protection scope of this application.
[0197] Corresponding to the network diagnostic processing method implemented by the target device in the above embodiments, Figure 8 This is a schematic diagram of the structure of a network diagnostic processing device provided in an embodiment of this application. The device can be implemented as part or all of a computer device by software, hardware, or a combination of both. The computer device can be... Figure 10 The electronic device shown.
[0198] Reference Figure 8 The network diagnostic processing device is applied to the target device, and the device includes: The acquisition unit 801 is used to acquire the multi-layer network status parameters of the current communication link of the target device; Analysis unit 802 is used to perform cross-layer correlation analysis on multi-layer network status parameters and generate network diagnostic results, which include fault root cause identification and communication link switching strategy. The determining unit 803 is used to determine the policy execution verification result corresponding to the communication link switching policy; The execution unit 804 is used to perform communication link switching operations based on the policy execution verification results.
[0199] Corresponding to the network diagnostic processing method implemented by the server in the above embodiments, Figure 9 This is a schematic diagram of the structure of a network diagnostic processing device provided in an embodiment of this application. The device can be implemented as part or all of a computer device by software, hardware, or a combination of both. The computer device can be... Figure 10 The electronic device shown.
[0200] Reference Figure 9 The network diagnostic processing device is applied to a server, and the device includes: The acquisition module 901 acquires the multi-layer network status parameters of the current communication link of the target device; Analysis module 902 is used to perform cross-layer correlation analysis on multi-layer network status parameters and generate network diagnostic results. The network diagnostic results include fault root cause identification and communication link switching strategy. The communication link switching strategy is determined based on fault root cause identification, preset link priority ranking and fault identification threshold. The determination module 903 is used to determine the policy execution verification result corresponding to the communication link switching policy; The optimization module 904 is used to adjust at least one of the model parameters used in link priority sorting, fault identification threshold and cross-layer correlation analysis based on the strategy execution verification results.
[0201] It is understood that the embodiments of the network diagnostic processing device and any implementation thereof correspond to the embodiments of the network diagnostic processing method and any implementation thereof. The technical effects corresponding to the embodiments of the network diagnostic processing device and any implementation thereof can be found in the technical effects corresponding to the above-mentioned embodiments of the network diagnostic processing method and any implementation thereof, and will not be repeated here.
[0202] It should be noted that the network diagnostic processing device provided in the above embodiments is only an example of the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0203] The functional units and modules in the above embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of the embodiments of this application.
[0204] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0205] This application also provides an electronic device, which includes one or more processors and a memory; The memory is coupled to one or more processors. The memory is used to store computer program code, which includes computer instructions. One or more processors invoke the computer instructions to cause the electronic device to perform the network diagnostic processing method described above.
[0206] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 1000 can be a mobile phone, smart screen, tablet computer, wearable electronic device, in-vehicle electronic device, augmented reality (AR) device, virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), projector, or a communication device such as a server, storage device, or base station, or a smart car, etc. This application embodiment does not impose any limitations on the specific type of electronic device.
[0207] The memory 1001 can be used to store computer programs 1002 and modules. The processor 1003 executes various functional applications and data processing of the electronic device by running the software programs and modules stored in the memory 1001. The memory 1001 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device (such as audio data, telephone directory, etc.). In addition, the memory 1001 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0208] The processor 1003 may include one or more processors such as a central processing unit (CPU), an application processor (AP), and a baseband processor. The processor can serve as the nerve center and command center of the wireless router. The processor 1003 can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution. The memory 1001 can be used to store executable program code, including instructions. The processor 1003 executes various functional applications and data processing of the network device by running the instructions stored in the memory. The memory 1001 may include a program storage area and a data storage area, such as storing data for audio signals to be played. For example, the memory may be Double Data Rate Synchronous Dynamic Random Access Memory (DDR) or Flash memory.
[0209] This application also provides a computer-readable storage medium storing computer instructions; when the computer-readable storage medium is used on an electronic device, it causes the electronic device to perform the network diagnostic processing method described above.
[0210] The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or can include one or more data storage devices such as servers or data centers that can be integrated with media. The available medium can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media, or semiconductor media (e.g., solid-state disks (SSDs)).
[0211] This application also provides a computer program product containing computer instructions, which, when run on an electronic device, enables the electronic device to execute the aforementioned network diagnostic processing method.
[0212] The computer storage medium and computer program product provided in the embodiments of this application are used to execute the methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects corresponding to the methods provided above, and will not be repeated here.
[0213] In the above embodiments, implementation can also be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, Digital Subscriber Line, DSL) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk drive (HDD), or solid-state drive (SSD), etc., and the storage medium can also include combinations of the above types of memory.
[0214] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0215] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments claimed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0216] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0217] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0218] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A network diagnostic processing method, characterized in that, Applied to target devices, including: Obtain the multi-layer network status parameters of the current communication link of the target device; Cross-layer correlation analysis is performed on the multi-layer network status parameters to generate network diagnostic results, wherein the network diagnostic results include fault root cause identification and communication link switching strategy; Determine the policy execution verification result corresponding to the communication link switching policy; Based on the verification results of the strategy, a communication link switching operation is performed.
2. The method according to claim 1, characterized in that, The step of performing cross-layer correlation analysis on the state parameters of the multi-layer network to generate network diagnostic results includes: Cross-layer correlation analysis is performed on the state parameters of the multi-layer network to obtain the root cause identifier of the fault; The communication link switching strategy is determined based on the fault root cause identifier, the preset link priority ranking, and the fault identification threshold. The root cause identifier includes information on the level of the fault occurrence, the type of fault, and the severity of the fault.
3. The method according to claim 2, characterized in that, The step of performing cross-layer correlation analysis on the state parameters of the multi-layer network to obtain the root cause identifier of the fault includes: The state parameters of the multi-layer network are time-aligned and standardized to obtain a multi-layer parameter sequence with a unified time dimension. The parameter sequence of each layer in the multi-layer parameter sequence is verified layer by layer to identify the network layers with abnormalities and the corresponding abnormal parameters; Establish the correlation between abnormal parameters at different network levels, eliminate non-causal concurrent anomalies, and obtain the correlationd abnormal parameter combination; The associated abnormal parameter combinations are matched with a preset fault root cause association model to generate the fault root cause identifier. The preset fault root cause association model stores multiple sets of mapping relationships between abnormal parameter combinations and fault root cause identifiers. Each set of mapping relationships corresponds to an abnormal combination consisting of abnormal parameters from at least two different network layers.
4. The method according to claim 3, characterized in that, The step-by-step verification of each layer of the multi-layer parameter sequence identifies abnormal network layers and their corresponding abnormal parameters, including: The network state parameters in each layer of the multi-layer parameter sequence are checked layer by layer to see if they exceed the corresponding fault identification threshold, thereby identifying the abnormal network layers and their corresponding abnormal parameters. The fault identification threshold includes a primary threshold and a secondary threshold. The primary threshold is used to trigger the adjustment of the link priority sorting, and the secondary threshold is used to trigger the output of the communication link switching strategy.
5. The method according to claim 2, characterized in that, The step of determining the communication link switching strategy based on the fault root cause identifier, the preset link priority ranking, and the fault identification threshold includes: Based on the fault root cause identifier, determine whether the target device needs to perform a communication link switch; If the target device needs to perform a communication link switch, then based on the fault level and fault type corresponding to the fault root cause identifier, a candidate communication link is determined, and based on the fault identification threshold, the triggering time for the communication link switch is determined. The candidate communication links are sorted according to the link priority to determine the target switching link, which includes the primary link and the alternative links; The communication link switching strategy is generated based on the triggering time of the communication link switching and the target switching link.
6. The method according to claim 5, characterized in that, The step of determining candidate communication links based on the fault level and fault type corresponding to the fault root cause identifier includes: Based on the fault level and fault type corresponding to the fault root cause identifier, select available communication links that can resolve the fault. Based on the historical handover records of the target device, the candidate communication link is determined from the available communication links.
7. The method according to claim 5, characterized in that, The determination of the policy execution verification result corresponding to the communication link switching policy includes: Within the preset verification window period, the communication link switching strategy is executed, and new multi-layer network status parameters of the target switching link are continuously collected. Cross-layer correlation analysis is performed on the new multi-layer network state parameters to generate policy execution verification results, wherein the policy execution verification results include at least one of policy validity identifier, network state stability identifier of target switching link, and jitter identifier.
8. The method according to claim 7, characterized in that, The step of performing a communication link switching operation based on the verification result of the strategy includes: When there are at least two target switching links, the verification indicators of each target switching link are collected during the verification window period. Compare the verification metrics of each target switching link, wherein the verification metrics include service recovery speed, parameter stability rate and data reporting success rate; Based on the comparison results, the target switching link with the best comprehensive index is marked as the primary link, and the remaining target switching links are marked as alternative links. The link priority ranking of each target switching link is then updated. Based on the main link, a communication link switching operation is performed.
9. The method according to claim 5, characterized in that, The method further includes: Based on the verification results of the strategy, adjust at least one of the model parameters used in the link priority ranking, fault identification threshold, and cross-layer correlation analysis in the following manner: When the strategy execution verification result indicates that the communication link switching strategy is effective and the target switching link meets the preset stability conditions, the link priority ranking of the target switching link is increased; When the strategy execution verification result indicates that the communication link switching strategy is invalid or the target switching link does not meet the preset stability conditions, the link priority ranking of the target switching link is lowered, the link priority ranking of the stable link before switching is raised, and the fault identification threshold of the corresponding level of the fault root cause identifier is adjusted. When the strategy execution verification result indicates the presence of an unidentified fault, the model parameters used in the cross-layer correlation analysis are updated.
10. The method according to any one of claims 1 to 8, characterized in that, The method further includes: When the target device is detected to be disconnected, all network links that can be successfully connected are set to the highest priority. Based on the preset link priority, the network link with the highest priority is determined as the primary link, and the remaining switching links are marked as alternative links. Based on the main link, a communication link switching operation is performed.
11. The method according to any one of claims 1 to 8, characterized in that, The multi-layer network state parameters include one or more of the following: physical layer parameters, transport layer parameters, connection establishment layer parameters, application layer parameters, and service layer parameters. The physical layer parameters include at least one of signal strength, wireless LAN signal quality, and mobile communication signal strength. The transport layer parameters include at least one of network latency, packet loss rate, and jitter. The connection establishment layer parameters include at least one of the following: domain name resolution time, transmission control protocol connection establishment time, and secure transport layer protocol handshake time. The application layer parameters include at least one of the following: private protocol connection status and heartbeat response time; The business layer parameters include at least one of the following: data reporting success rate and control command confirmation success rate.
12. A network diagnostic processing method, applied to a server, characterized in that, include: Obtain the multi-layer network status parameters of the current communication link of the target device; Cross-layer correlation analysis is performed on the multi-layer network status parameters to generate network diagnostic results. The network diagnostic results include fault root cause identification and communication link switching strategy. The communication link switching strategy is determined based on the fault root cause identification, preset link priority ranking, and fault identification threshold. Determine the policy execution verification result corresponding to the communication link switching policy; Based on the verification results of the strategy, adjust at least one of the model parameters used in the link priority sorting, fault identification threshold, and cross-layer correlation analysis.
13. The method according to claim 12, characterized in that, The step of performing cross-layer correlation analysis on the state parameters of the multi-layer network to generate network diagnostic results includes: Cross-layer correlation analysis is performed on the state parameters of the multi-layer network to obtain the root cause identifier of the fault; The communication link switching strategy is determined based on the fault root cause identifier, the link priority sorting, and the fault identification threshold. The root cause identifier includes information on the level of the fault occurrence, the type of fault, and the severity of the fault.
14. The method according to claim 13, characterized in that, The step of performing cross-layer correlation analysis on the state parameters of the multi-layer network to obtain the root cause identifier of the fault includes: The state parameters of the multi-layer network are time-aligned and standardized to obtain a multi-layer parameter sequence with a unified time dimension. The parameter sequence of each layer in the multi-layer parameter sequence is verified layer by layer to identify the network layers with abnormalities and the corresponding abnormal parameters; Establish the correlation between abnormal parameters at different network levels, eliminate non-causal concurrent anomalies, and obtain the correlationd abnormal parameter combination; The associated abnormal parameter combinations are matched with a preset fault root cause association model to generate the fault root cause identifier. The fault root cause association model stores multiple sets of mapping relationships between abnormal parameter combinations and fault root cause identifiers. Each set of mapping relationships corresponds to an abnormal combination consisting of abnormal parameters from at least two different network layers.
15. The method according to claim 13, characterized in that, The step of determining the communication link switching strategy based on the fault root cause identifier, the link priority ranking, and the fault identification threshold includes: Based on the fault root cause identifier, determine whether the target device needs to perform a communication link switch; If the target device needs to perform a communication link switch, then based on the fault level and fault type corresponding to the fault root cause identifier, a candidate communication link is determined, and based on the preset fault identification threshold, the triggering time for the communication link switch is determined. The candidate communication links are sorted according to the link priority to determine the target switching link, which includes the primary link and the alternative links; The communication link switching strategy is generated based on the triggering time of the communication link switching and the target switching link.
16. The method according to claim 15, characterized in that, The step of adjusting at least one of the model parameters used in the link priority ranking, fault identification threshold, and cross-layer correlation analysis based on the verification results of the strategy includes: When the strategy execution verification result indicates that the communication link switching strategy is effective and the target switching link meets the preset stability conditions, the link priority ranking of the target switching link is increased; When the strategy execution verification result indicates that the communication link switching strategy is invalid or the target switching link does not meet the preset stability conditions, the link priority ranking of the target switching link is lowered, the link priority ranking of the stable link before switching is raised, and the fault identification threshold of the corresponding level of the fault root cause identifier is adjusted. When the strategy execution verification result indicates the presence of an unidentified fault, the model parameters used in the cross-layer correlation analysis are updated.
17. A network diagnostic processing system, characterized in that, include: The target device is used to obtain the multi-layer network status parameters of the current communication link of the target device; The server connects to the target device and is used to perform cross-layer correlation analysis on the multi-layer network status parameters and generate network diagnostic results. The network diagnostic results include fault root cause identification and communication link switching strategy. The communication link switching strategy is determined based on the fault root cause identification, preset link priority ranking, and fault identification threshold. The target device is further configured to determine the policy execution verification result corresponding to the communication link switching policy, and to perform a communication link switching operation based on the policy execution verification result; The server is also used to perform verification results according to the strategy and adjust at least one of the model parameters used in the link priority sorting, fault identification threshold and cross-layer correlation analysis.
18. A network diagnostic processing device, characterized in that, Applied to a target device, the device includes: The acquisition unit is used to acquire the multi-layer network status parameters of the current communication link of the target device; The analysis unit is used to perform cross-layer correlation analysis on the multi-layer network status parameters and generate network diagnostic results, wherein the network diagnostic results include fault root cause identification and communication link switching strategy. A determining unit is used to determine the policy execution verification result corresponding to the communication link switching policy; The execution unit is used to perform communication link switching operations based on the verification results of the strategy.
19. A network diagnostic processing device, applied to a server, characterized in that, The device includes: The acquisition module acquires the multi-layer network status parameters of the current communication link of the target device. The analysis module is used to perform cross-layer correlation analysis on the multi-layer network status parameters and generate network diagnostic results. The network diagnostic results include fault root cause identification and communication link switching strategy. The communication link switching strategy is determined based on the fault root cause identification, preset link priority ranking and fault identification threshold. The determining module is used to determine the policy execution verification result corresponding to the communication link switching policy; An optimization module is used to adjust at least one of the model parameters used in the link priority sorting, fault identification threshold, and cross-layer correlation analysis based on the verification results of the strategy.
20. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it causes the electronic device to implement the method as described in any one of claims 1 to 16.
21. A computer program product, characterized in that, Includes a computer program, which, when run, causes the method as described in any one of claims 1 to 16 to be performed.
22. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 16.