A fault diagnosis processing method and device, network equipment and storage medium

By acquiring and analyzing business requirements, network metrics, and second-level network metrics, the system automatically locates the root cause of faults in 5G industry virtual private networks, solving the problems of coarse-grained fault diagnosis and reliance on manual methods in existing technologies, and achieving efficient fault location and low-cost operation and maintenance.

CN118802469BActive Publication Date: 2026-01-16CHINA MOBILE COMM LTD RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410501243.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-24
Publication Date
2026-01-16
Estimated Expiration
2044-04-24

AI Technical Summary

Technical Problem

Existing 5G industry virtual private network fault diagnosis solutions suffer from coarse-grained fault diagnosis, difficulty in data acquisition, lack of standardized interfaces, and reliance on manual data parsing, resulting in low location efficiency and high costs, especially when dealing with occasional and difficult problems.

Method used

By acquiring relevant business requirements, network metrics, and second-level network metrics through the first network device, hierarchical analysis is performed to automatically identify abnormal metrics and locate the root cause of problems, providing a unified interface and an automated fault diagnosis mechanism.

Benefits of technology

It enables more precise fault diagnosis, improves operation and maintenance efficiency, reduces labor costs, and enhances the accuracy and efficiency of fault location.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118802469B_ABST
    Figure CN118802469B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a fault diagnosis processing method, device and equipment and a storage medium. The method comprises: a first network device obtaining a first type of index and a second type of index, and obtaining a third type of index from a second network device; the first type of index is a service demand related index, the second type of index comprises a threshold value of a network index and / or a service index, and the third type of index is a second-level network index; when the third type of index does not satisfy the first type of index, it is determined that an index is abnormal; based on the third type of index and the second type of index, hierarchical analysis is performed on the index abnormality to obtain an analysis result; the analysis result comprises a problem root cause corresponding to an abnormal index.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of communication, in particular to a fault diagnosis processing method and device, network equipment and storage medium. BACKGROUND

[0002] Industry customers require that the 5G industry virtual private network provide deterministic guarantee capabilities, that is, the network needs to meet the business continuity operation requirements, such as the requirement of 24-hour operation of the factory automated guided vehicle (AGV) business. However, due to the air interface wireless environment of the 5G network, the occurrence of environmental influences is large, such as sudden shielding or interference or resource limitation, etc., which leads to the problems of coarse fault diagnosis granularity (such as the index collection time granularity being mostly minute level) and great difficulty in obtaining fault data in the current fault diagnosis scheme. Moreover, there is no standard interface for index data output, data analysis relies on manufacturers, index analysis relies on experts, the manual threshold is high, and the problem positioning efficiency is low; the fault diagnosis mainly relies on manual retesting to find the problem root cause, especially for occasional difficult problems, the retesting problem takes a long time, and needs to coordinate multiple departments, which is high in manual cost. SUMMARY

[0003] To solve the existing technical problems, the embodiments of the present application provide a fault diagnosis processing method, device, network equipment and storage medium.

[0004] To achieve the above-mentioned purpose, the technical scheme of the embodiments of the present application is as follows:

[0005] In a first aspect, the embodiments of the present application provide a fault diagnosis processing method, which is applied to a first network equipment, and the method comprises:

[0006] The first network equipment obtains a first type of index and a second type of index, and obtains a third type of index from a second network equipment; the first type of index is a business requirement related index, the second type of index includes a threshold value of a network index and / or a business index, and the third type of index is a second-level network index;

[0007] When the third type of index does not meet the first type of index, it is determined that the index is abnormal;

[0008] Based on the third type of index and the second type of index, a hierarchical analysis is performed on the index abnormality to obtain an analysis result; the analysis result includes a problem root cause corresponding to the abnormal index.

[0009] In the above-mentioned scheme, the method further comprises: the first network equipment determines a first business type corresponding to the third type of index based on the characteristics of the third type of index.

[0010] In the scheme, when the third type of index does not meet the first type of index, the first network device determines that the index is abnormal.

[0011] In the scheme, the third type of index is a third type of index of the first service; and the obtaining of the third type of index from the second network device includes:

[0012] The first network device receives the third type of index of the first service sent by the second network device; the third type of index of the first service is obtained by the second network device based on the first identifier of the first service; or the third type of index of the first service is obtained by the second network device based on the first service model corresponding to the first service.

[0013] Alternatively, the first network device receives the third type of index sent by the second network device, identifies the third type of index based on the first service model corresponding to the first service obtained in advance, and obtains the third type of index of the first service.

[0014] In the scheme, the third type of index is a third type of index of the first user; and the obtaining of the third type of index from the second network device includes:

[0015] The first network device receives the third type of index of the first user sent by the second network device; the third type of index of the first user is obtained by the second network device based on the identifier of the first user; or

[0016] The first network device receives the third type of index sent by the second network device, identifies the third type of index based on the identifier of the first user, and obtains the third type of index of the first user.

[0017] In the scheme, the hierarchical analysis of the index abnormality based on the third type of index and the second type of index includes:

[0018] The first network device determines abnormal time information corresponding to the abnormal index, determines a first type to which a problem belongs based on the abnormal time information, obtains first processing logic corresponding to the first type, and the first type is an occasional problem, a periodic problem or a long-term problem.

[0019] The third type of index and the second type of index are hierarchically analyzed based on the first processing logic.

[0020] In the scheme, the hierarchical analysis of the index abnormality based on the third type of index and the second type of index to obtain an analysis result comprises: the first network device compares the third type of index and the second type of index corresponding to each level in turn, and obtains the analysis result based on the comparison result; wherein the third type of index compared in the next level is associated with the third type of index with abnormality in the comparison result of the previous level; the problem range indicated by the comparison result of the next level is reduced compared with the previous level.

[0021] In the scheme, the hierarchical analysis of the index abnormality based on the third type of index and the second type of index to obtain an analysis result comprises:

[0022] The first network device compares the third type of index and the second type of index corresponding to the first level to obtain a first comparison result, and the first comparison result comprises end-to-end delay abnormality information.

[0023] The second comparison result is obtained by comparing the third type of index and the second type of index corresponding to the second level, and the second comparison result comprises segment delay abnormality information; the segment delay comprises one or more of the following: uplink and / or downlink user equipment (UE) delay, uplink and / or downlink medium access control (MAC) layer delay, uplink and / or downlink packet data convergence protocol (PDCP) layer delay, uplink and / or downlink radio link control (RLC) delay.

[0024] According to the segment delay abnormality information, the third comparison result is obtained by comparing the third type of index and the second type of index corresponding to the third level, and the third comparison result comprises a first problem category to which the index abnormality belongs, and the first problem category comprises one or more of the following: user state problem, channel condition problem, scheduling opportunity problem, and resource limitation problem.

[0025] According to the first problem category, the fourth comparison result is obtained by comparing the third type of index and the second type of index corresponding to at least one fourth level based on the comparison result, and the root cause of the problem is determined based on the fourth comparison result.

[0026] In the scheme, the method further comprises: the first network device determines the location information corresponding to the abnormal problem based on the network topology obtained in advance, the mobile path of the terminal, and the abnormal time information corresponding to the abnormal problem.

[0027] In the scheme, the method further comprises: the first network device receives the location information of the terminal sent by the second network device.

[0028] In the scheme, the method further comprises: the first network device determines the solution corresponding to the root cause of the problem according to a first mapping relationship; wherein the first mapping relationship comprises a mapping relationship between a plurality of root causes of the problem and solutions.

[0029] In the solution, the method further includes: the first network device stores a first index in the third type of indexes in a way of covering storage, and only retains the first index of a specified time length or a specified data volume; and / or,

[0030] The first network device stores a second index in the third type of indexes in a way of continuous storage.

[0031] The importance of the first index is lower than that of the second index, and the importance is related to a service or a service type.

[0032] In the solution, the third type of indexes includes one or more of the following: a dynamic air interface performance index, an air interface state index, a static index, a reference signal received power (RSRP) of a demodulation reference signal (DMRS) and / or a channel sounding reference signal (SRS), a signal to interference plus noise ratio (SINR) of the DMRS and / or the SRS, and a SINR of each data stream in a multi-data stream case.

[0033] In a second aspect, an embodiment of the present application further provides a fault diagnosis processing method, which is applied to a second network device, and includes the following steps:

[0034] The second network device sends a third type of indexes to the first network device, the third type of indexes being second-level network indexes; the third type of indexes are used by the first network device to determine an index exception and perform hierarchical analysis on the index exception, to obtain an analysis result, the analysis result including a problem root cause corresponding to an abnormal index.

[0035] In the solution, the third type of indexes are third type of indexes of a first service; before the second network device sends the third type of indexes to the first network device, the method further includes:

[0036] The second network device identifies indexes based on a first identifier of the first service, to obtain the third type of indexes of the first service; or,

[0037] The second network device identifies indexes based on a first service model corresponding to the first service, to obtain the third type of indexes of the first service.

[0038] In the solution, the third type of indexes are third type of indexes of a first user; before the second network device sends the third type of indexes to the first network device, the method further includes:

[0039] The second network device identifies indexes based on an identifier of the first user, to obtain the third type of indexes of the first user.

[0040] In the solution, the method further comprises: the second network device sends the location information of the terminal to the first network device.

[0041] In a third aspect, the embodiments of the present application further provide a fault diagnosis processing device, which is applied to a first network device, and comprises an acquisition unit and a diagnosis unit.

[0042] The acquisition unit is configured to obtain a first type of index and a second type of index, and obtain a third type of index from a second network device; the first type of index is a service demand related index, the second type of index comprises a threshold value of a network index and / or a service index, and the third type of index is a second-level network index.

[0043] The diagnosis unit is configured to determine that an index is abnormal when the third type of index does not meet the first type of index, perform hierarchical analysis on the index abnormality based on the third type of index and the second type of index, and obtain an analysis result; the analysis result comprises a problem root cause corresponding to an abnormal index.

[0044] In a fourth aspect, the embodiments of the present application further provide a fault diagnosis processing device, which is applied to a second network device, and comprises a communication unit configured to send a third type of index to the first network device; the third type of index is a second-level network index; the third type of index is used for the first network device to determine an index abnormality, perform hierarchical analysis on the index abnormality, obtain an analysis result, and the analysis result comprises a problem root cause corresponding to an abnormal index.

[0045] In a fifth aspect, the embodiments of the present application further provide a computer readable storage medium, which stores a computer program; the program is executed by a processor to implement the steps of the fault diagnosis processing method in the first aspect or the second aspect.

[0046] In a sixth aspect, the embodiments of the present application further provide a network device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor; the processor implements the steps of the fault diagnosis processing method in the first aspect or the second aspect when executing the program.

[0047] In a seventh aspect, the embodiments of the present application further provide a computer program product, which comprises computer program instructions; the computer program instructions enable a computer to execute the steps of the fault diagnosis processing method in the first aspect or the second aspect.

[0048] This invention provides a fault diagnosis and processing method, apparatus, network device, and storage medium. The method includes: a first network device obtaining a first type of indicator and a second type of indicator, and obtaining a third type of indicator from a second network device; the first type of indicator is a service requirement-related indicator, the second type of indicator includes threshold values ​​for network indicators and / or service indicators, and the third type of indicator is a second-level network indicator; when the third type of indicator does not meet the first type of indicator, an indicator anomaly is determined; hierarchical analysis of the indicator anomaly is performed based on the third type of indicator and the second type of indicator to obtain analysis results; the analysis results include the root cause of the problem corresponding to the abnormal indicator. This invention proposes an automated fault diagnosis and recovery mechanism. The first network device obtains second-level network indicators from the second network device through a unified interface and has indicator analysis capabilities. Through hierarchical analysis, the root cause of the problem corresponding to the abnormal indicator is determined, and finally, a solution corresponding to the root cause is obtained. Compared with traditional fault diagnosis, this invention collects more refined indicators, resulting in more accurate fault diagnosis; through automated hierarchical analysis, it greatly improves operation and maintenance efficiency and reduces labor costs. Attached Figure Description

[0049] Figure 1 This is a flowchart illustrating the fault diagnosis and processing method according to an embodiment of the present invention. Figure 1 ;

[0050] Figure 2 This is a schematic diagram illustrating the collection of the third type of indicator in the fault diagnosis and processing method of this invention.

[0051] Figure 3a and Figure 3b This is a schematic diagram of two-stream SINR analysis;

[0052] Figure 4 This is a schematic diagram of the diagnostic logic in the fault diagnosis and processing method of this invention.

[0053] Figure 5 This is a schematic diagram of the indicator hierarchy analysis process in the fault diagnosis and processing method of this invention.

[0054] Figure 6 This is a schematic diagram of the algorithm logic in the fault diagnosis and processing method of this invention.

[0055] Figure 7 This is a schematic diagram of index collection in the fault diagnosis and processing method of this invention.

[0056] Figure 8 This is a flowchart illustrating the fault diagnosis and processing method according to an embodiment of the present invention. Figure 2 ;

[0057] Figure 9Interface diagram of fault diagnosis and recovery mechanism in the fault diagnosis processing method of the embodiment of the present application

[0058] Figure 10 Composition structure diagram of the fault diagnosis processing device of the embodiment of the present application Figure 1

[0059] Figure 11 Composition structure diagram of the fault diagnosis processing device of the embodiment of the present application Figure 2

[0060] Figure 12 Hardware composition structure diagram of the network device of the embodiment of the present application. DETAILED DESCRIPTION

[0061] The present application will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0062] The technical solution of the embodiment of the present application can be applied to various communication systems, for example: Global System of Mobile communication (GSM) system, Long Term Evolution (LTE) system or 5G system, etc. Optionally, the 5G system or 5G network can also be referred to as New Radio (NR) system or NR network.

[0063] For example, the communication system to which the embodiment of the present application is applied can include a network device and a terminal device (also referred to as a terminal, a communication terminal, etc.); the network device can be a device that communicates with the terminal device. Among them, the network device can provide communication coverage in a certain area range, and can communicate with terminals located in the area. Optionally, the network device can be a base station in each communication system, for example, an Evolutional Node B (eNB) in the LTE system, and for example, a base station (gNB) in the 5G system or the NR system.

[0064] It should be understood that the devices with communication functions in the network / system in the embodiments of the present application can be referred to as communication devices. The communication devices can include network devices and terminals with communication functions, and the network devices and terminal devices can be the specific devices described above, which will not be described here; the communication devices can also include other devices in the communication system, such as network controllers, mobile management entities and other network entities, which are not limited in the embodiments of the present application.

[0065] ​​It should be understood that the terms "system" and "network" are often used interchangeably herein. The term "and / or", used herein only describes an associated relationship, which means that there can be three relationships, for example, A and / or B, which means that A exists alone, A and B exist together, B exists alone. In addition, the character " / " generally represents an "or" relationship between the front and rear associated objects.

[0066] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be exchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0067] The embodiments of the present application provide a fault diagnosis processing method. Figure 1 The flowchart of the fault diagnosis processing method of the embodiments of the present application is shown in Figure 1 As shown in Figure 1 The method comprises the following steps:

[0068] Step 101: The first network device obtains the first type of index and the second type of index, and obtains the third type of index from the second network device; the first type of index is a service demand related index, the second type of index includes a threshold value of a network index and / or a service index, and the third type of index is a second-level network index;

[0069] Step 102: When the third type of index does not meet the first type of index, it is determined that the index is abnormal;

[0070] Step 103: Based on the third type of index and the second type of index, the index abnormality is analyzed to obtain an analysis result; the analysis result includes a problem root cause corresponding to the abnormal index.

[0071] In this embodiment, the first network device can be a first network element, a first network function, etc. The first network device has an automated fault diagnosis module, which is connected with the second network device through an interface to obtain the third type of indicators. The second network device can be an access network device, such as a base station, etc. In some optional embodiments, the first network device can be a network device independently arranged and specially used for automated fault diagnosis, or can be arranged together with other network devices, such as a network management device, which includes the automated fault diagnosis module.

[0072] Referring to Figure 2 As shown in the figure, the base station (i.e. the second network device) collects or monitors the third type of indicators, and can send the obtained third type of indicators to the automated fault diagnosis module of the first network device through interface targeting metric information. When the automated fault diagnosis module of the first network device detects abnormal indicators through the obtained third type of indicators and the service type indicators (first type of indicators), it obtains the problem root cause corresponding to the abnormal indicators and the corresponding solution through hierarchical analysis. In other optional embodiments, the base station (i.e. the second network device) can also determine abnormal services through interaction with the automated fault diagnosis module of the first network device, and the base station (i.e. the second network device) can automatically configure tracking of abnormal services to automatically obtain second-level network indicators.

[0073] In this embodiment, the input data for fault diagnosis includes the first type of indicators, the second type of indicators and the third type of indicators; wherein, except that the third type of indicators are obtained from the second network device, the first type of indicators and the second type of indicators are statically set, i.e. pre-set or obtained before fault diagnosis.

[0074] In this embodiment, the first type of indicators are related to service requirements, and are used to determine whether the network indicators (i.e. the third type of indicators) meet the service requirements; if not, it can be indicated that the indicators are abnormal, and fault diagnosis needs to be performed based on the network indicators (i.e. the third type of indicators). In some optional embodiments, the first type of indicators can be related to service types, and different service types can correspond to different first type of indicators. Exemplarily, the service types can include control type services, video type services, etc., and the service types can be divided based on preset rules to obtain multiple service types.

[0075] In this embodiment, the second type of indicators include threshold values of network indicators and / or service indicators, which are used as reference values for fault hierarchical diagnosis. If a certain network indicator (third type of indicator) exceeds the corresponding threshold value (second type of indicator), further diagnosis of the next level is performed until the diagnosis result of the problem root cause is obtained.

[0076] In this embodiment, the third type of index is a second-level network index, which can be a second-level air interface index. The second network device can periodically send the third type of index to the first network device for analyzing network conditions, diagnosing specific abnormal indexes, and locating problems.

[0077] For example, Table 1 is an input index definition table of an embodiment of the present application. As shown in Table 1, the service class index (i.e., the first type of index) can include service types and the first type of index corresponding to each service type, and the first type of index corresponding to each service type can include end-to-end ping delay, latency, rate requirement, and the like. The threshold class index (i.e., the second type of index) can include at least an index empirical value and a constant index. The network class index (i.e., the third type of index) can include at least a latency rate index and an air interface state index, and the like. For example, the latency rate index includes at least protocol stack segment latency, such as uplink and / or downlink user equipment (UE) latency, uplink and / or downlink medium / medium access control (MAC) layer latency, uplink and / or downlink packet data convergence protocol (PDCP) layer latency, uplink and / or downlink radio link control (RLC) latency, and the like. It should be noted that the types of indexes listed in Table 1 are only examples, and the present embodiment is not limited to the types of indexes shown in Table 1.

[0078] Table 1

[0079]

[0080]

[0081] In some optional embodiments, the third type of index includes one or more of the following: a dynamic air interface performance index, an air interface state index, a static index, a reference signal receiving power (RSRP) of air interface signaling, a demodulation reference signal (DMRS), and / or a sounding reference signal (SRS), a signal to interference plus noise ratio (SINR) of the DMRS and / or the SRS, and a SINR of each data stream in a multi-data stream case.

[0082] In this embodiment, the third type of indicators can include one or more of dynamic air interface performance indicators, air interface state indicators, static indicators, and air interface signaling; the air interface performance indicators are used to reflect air interface performance and can include various indicators related to latency and rate. The air interface state indicators can include dynamic and static configuration indicators and are used to reflect air interface state. The static indicators can be specifically indicators related to pre-scheduling. The air interface signaling can be signaling related to the air interface, such as radio resource control (RRC) signaling.

[0083] For example, Table 2 is part of the third type of indicators. As can be seen from Table 2, the output time granularity of the third type of table is 1 second.

[0084] Table 2

[0085]

[0086]

[0087] On this basis, the third type of indicators can further include one or more of RSRP of DMRS and / or SRS, SINR of DMRS and / or SRS, and SINR of each data stream in the case of multiple data streams. By positioning and solving the problem of occasional disconnection of AGV services, it is found through analysis of the above dynamic air interface performance indicators, air interface state indicators, static indicators, and air interface signaling that the channel quality related indicators RSRP and SINR need to be optimized, and therefore the related indicators reflecting channel quality are added to the third type of indicators. On the one hand, the RSRP and SINR indicators corresponding to DMRS and SRS signals are different. On the other hand, the problem analysis involves imbalance between streams (at least between two data streams), and the average SINR indicator cannot reflect such typical problems.

[0088] Figure 3a and Figure 3b is a schematic diagram of two-stream SINR analysis; as shown in Figure 3a and Figure 3b , the AGV disconnection abnormal time point corresponds to 0 for 2-8 s of consecutive PDCP received packet numbers and a rank (RANK) of 2, as shown in Figure 3a , the RANK is 2, the uplink two-stream DMRS SINR difference is more than 30 dB, as shown in Figure 3b , when the RANK changes from 1 to 2, the bit error rate suddenly increases from 0 to 100%. The root cause is high error caused by the terminal sending two streams with large difference in uplink signal quality. The introduction of RSRP and SINR corresponding to DMRS and SRS, and the SINR indicator of each stream in the case of multiple streams can more accurately locate the problems of channel quality and stream imbalance, perfect the existing scheme indicators, and realize deep and comprehensive positioning of occasional difficult problems in the existing network.

[0089] In this embodiment, the first network device compares the obtained third type of index with the first type of index related to the service requirement, and if the third type of index of the same type does not meet the service requirement of the first type of index, it indicates that the index is abnormal. Further, hierarchical analysis is performed according to the third type of index and the second type of index to determine the root cause of the fault. In each level of fault analysis, different third type of indexes and corresponding second type of indexes can be used. The problem range represented by the analysis results of each level is successively reduced, that is, when the third type of index and the second type of index of a certain level are compared and it is determined that the index is abnormal, the third type of index and the second type of index of the next level associated with the abnormal index are compared, the range is gradually reduced, and finally the root cause of the problem is located.

[0090] In some optional embodiments of the present application, the method further comprises: the first network device determining, based on the characteristics of the third type of index, the first service type corresponding to the third type of index.

[0091] Correspondingly, in some optional embodiments, when the third type of index does not meet the first type of index, determining that the index is abnormal comprises: the first network device determining that the index is abnormal when the third type of index does not meet the first type of index corresponding to the first service type.

[0092] In this embodiment, the third type of index can be a network index under various service types; for example, the service types can include control type services, video type services, and the like. Different service types can correspond to different first type of indexes and / or second type of indexes. In other words, the service requirements corresponding to different service types can be different, and / or the threshold values of the network indexes and / or service indexes corresponding to different service types can be different.

[0093] In this embodiment, the first network device determines the first service type corresponding to the third type of index based on the characteristics of the third type of index. In some optional embodiments, the characteristics can be packet sending characteristics, such as the size of the data packet, and the first network device can determine the service type according to the size of the data packet, for example, large packets can be video type services, small packets can be control type services, and the like.

[0094] In some optional embodiments of the present application, the hierarchical analysis of the index abnormality based on the third type of index and the second type of index comprises: the first network device determining abnormal time information corresponding to the abnormal index, determining the first type to which the problem belongs based on the abnormal time information, obtaining the first processing logic corresponding to the first type; the first type is an occasional problem, a periodic problem or a long-term problem; and the hierarchical analysis of the index abnormality is performed based on the third type of index and the second type of index according to the first processing logic.

[0095] In this embodiment, the first network device analyzes based on the third type of indicators and the second type of indicators, determines that the third type of indicators do not meet the first type of indicators, determines that the indicators are abnormal, and further obtains the time corresponding to the abnormal indicators (i.e., abnormal time information), determines the problem law according to the abnormal time information, for example, whether it is an occasional problem, a periodic problem or a long-term problem, different types of problems correspond to different processing logics. The processing logic includes the processing strategy of each level, for example, which indicators are analyzed in each level, the correlation of indicators between levels, and the problem root cause corresponding to each indicator in the last level, and the like.

[0096] In some optional embodiments of the present application, the hierarchical analysis of the indicators based on the third type of indicators and the second type of indicators to obtain the analysis result includes: the first network device compares the third type of indicators and the second type of indicators corresponding to each level in turn, and obtains the analysis result based on the comparison result; wherein the third type of indicators compared by the next level are associated with the third type of indicators that appear abnormal in the comparison result of the previous level; the problem range represented by the comparison result of the next level compared with the previous level is reduced.

[0097] In this embodiment, the first network device gradually narrows down the problem range by performing hierarchical analysis of the indicators based on the third type of indicators and the second type of indicators after determining that the indicators are abnormal, and finally locates the problem root cause. In some optional embodiments, each level can correspond to different or at least partially different third type of indicators; the third type of indicators between levels have correlation; through the abnormal analysis of each level, when the third type of indicators of a certain level appear abnormal compared with the second type of indicators, the indicator analysis of the next level is performed; at this time, only the third type of indicators of the next level associated with the abnormal indicators of the previous level that appear abnormal are compared and analyzed, and the third type of indicators of the next level associated with the normal indicators of the previous level are not analyzed. For example, the control class service delay indicator of a certain level appears abnormal, the next level needs to analyze the delay of each layer associated with the control class service delay indicator, if the delay of each layer is found to be normal compared with the relevant threshold value (second type of indicator), then the delay related indicators of the next layer associated with it are not analyzed, and the service delay is abnormal, and the abnormal reason is irrelevant to the air interface delay.

[0098] In some optional embodiments, the hierarchical analysis of the index anomaly based on the third type of index and the second type of index obtains an analysis result, including: the first network device compares the third type of index and the second type of index corresponding to a first level to obtain a first comparison result, the first comparison result including end-to-end delay anomaly information; compares the third type of index and the second type of index corresponding to a second level to obtain a second comparison result, the second comparison result including segment delay anomaly information; the segment delay includes one or more of: uplink and / or downlink UE delay, uplink and / or downlink MAC layer delay, uplink and / or downlink PDCP layer delay, uplink and / or downlink RLC delay; according to the segment delay anomaly information, the third type of index and the second type of index corresponding to a third level are compared to obtain a third comparison result, the third comparison result including a first problem category to which the index anomaly belongs, the first problem category including one or more of: user state problem, channel condition problem, scheduling opportunity problem, and resource limitation problem; according to the first problem category, the third type of index and the second type of index corresponding to at least one fourth level are compared, and a problem root cause is determined based on the comparison result.

[0099] In this embodiment, the first network device first performs analysis based on the end-to-end delay index, determines the end-to-end delay anomaly, further analyzes the segment delay index, determines the position of the segment delay index anomaly, and then performs the next level of index analysis according to the segment delay index anomaly, determines which one or more of the user state problem, the channel condition problem, the scheduling opportunity problem, and the resource limitation problem, focuses on one or more problem categories, and further performs the next level of index analysis according to the problem category, further narrows down the range, and determines the problem root cause. For example, according to the problem category, whether the related air interface state index is abnormal can be checked, the positioning unit is gradually narrowed down, and the problem root cause is determined.

[0100] In some optional embodiments of the application, the method further includes: the first network device determines a solution corresponding to the problem root cause according to a first mapping relationship, wherein the first mapping relationship includes a mapping relationship between a plurality of problem root causes and solutions.

[0101] In this embodiment, the corresponding solutions are set in advance according to each problem root cause to form a first mapping relationship. Specifically, the specific problems in the wireless network user state problem, the channel condition problem, the scheduling opportunity problem, and the resource limited problem can be located based on the second-level dynamic air interface performance class and state class indicators, the static parameter class indicators, and a series of solutions such as the corresponding scheduling optimization algorithm, the control resource allocation optimization algorithm, the interference avoidance, the network expansion, the network optimization, and the service level parameter function configuration are associated. Then, after the first network device determines the problem root cause, the problem root cause corresponding solution can be determined by searching the first mapping relationship. In other optional embodiments, the first network device can output the related problem root cause and solution through an interface to provide a reference for the operation and maintenance personnel. For example, the first mapping relationship can be as shown in Table 3.

[0102] Table 3

[0103]

[0104]

[0105] Figure 4 The diagnostic logic in the fault diagnosis processing method of the embodiment of the application is shown in FIG. 1. As shown in FIG. 1, the fault diagnosis processing flow of the embodiment of the application mainly includes inputting positioning indicators, abnormality judgment logic, and outputting positioning results. Wherein, Figure 4

[0106] The input positioning indicators can include service class indicators (i.e., first class indicators), threshold class indicators (i.e., second class indicators), and network class indicators (i.e., third class indicators). Specifically, the service class indicators can include end-to-end ping delay, service type and delay, and rate requirement as shown in Table 1. The threshold class indicators can include index empirical values and constant class. The network class indicators can include delay rate class indicators and air interface state class indicators. Wherein,

[0107] The service class indicators are service demand related indicators, which are set based on different services or service types. The service class indicators are used to discover problems and obtain problem rules. Specifically, based on the network class indicators and the service class indicators, it is determined that the network class indicators are abnormal when the network class indicators do not meet the service demand of the service class indicators, and subsequent abnormality judgment is needed. In addition, based on the time point of the abnormal indicators, the abnormal rules are analyzed to determine whether it is a periodic problem, a long-term problem, or an occasional problem.

[0108] The threshold class indicators are threshold values of network indicators and / or service indicators set according to different service types or service scenarios or service demands. The threshold class indicators are used as a reference value to judge the abnormality. When the network class indicators exceed the threshold class indicators, it is judged that the indicators are abnormal, and the next step of diagnosis is performed.

[0109] ​Network type index is a second network device (such as a base station) periodically sent to the first network base station. Network type index is used to analyze network conditions, determine specific abnormal indicators and locate the problem.

[0110] Abnormal judgment logic, based on the input positioning index, judges the abnormal index step by step. Specifically, the service type is first distinguished, and the abnormal judgment index is different for different service types. For example, the service type can be distinguished based on large packets or small packets, such as control type services, video type services, etc. Further, the problem rule can be divided into occasional problems, periodic problems and long-term problems, and the judgment conditions (or judgment logic) are different. Further, the abnormal delay segment is judged as a preliminary screening of the problem, which excludes some problems; wherein, the delay segment can include: uplink and / or downlink UE delay, uplink and / or downlink MAC layer delay, uplink and / or downlink PDCP layer delay, uplink and / or downlink RLC delay, etc. Finally, the abnormal judgment index: combined with the index, the problem range is gradually narrowed down to determine the specific problem. The specific problem can include dynamic type problems, static type problems, etc.

[0111] The input positioning result can specifically include the problem root cause and the corresponding solution.

[0112] Figure 5 The figure shows the index level analysis process in the fault diagnosis processing method of the embodiment of the application. Figure 6 The figure shows the algorithm logic in the fault diagnosis processing method of the embodiment of the application. The embodiment sets the judgment order (i.e. processing logic) according to the descending principle of the range and depth of the influence judgment logic, determines the abnormal judgment direction in combination with the service type and end-to-end delay, etc. For example, if the control type small packet service index is abnormal, the delay related network index is analyzed; in combination with the air interface segment delay index and the state type index, the problem range is gradually narrowed down, for example Figure 5 As shown in the figure, the terminal D1 delay (such as uplink UE delay) index is abnormal, only the uplink UE delay related network index is analyzed, and other layer delay related logic algorithm does not need to be executed, and the air interface problem root cause is finally determined.

[0113] Specifically, step 1: confirm the service type. The first network device can determine the first service type corresponding to the network type index based on the characteristics of the network type index. For example, according to the packet characteristics of the network type index, large packet services or small packet services are distinguished, generally small packets are control type services, and large packets are video type services. For different service types, the network type index used for abnormal analysis may be different, and accordingly, the service type index and the threshold type index may also be different.

[0114] Step 2: Problem rule determination. For the first service type, corresponding network type indicators and service type indicators are used for analysis, and when the network type indicators do not meet the service type indicators, it is determined that the indicators are abnormal, and problem rule analysis is performed according to the abnormal indicators corresponding to the abnormal practice to determine whether it is an occasional problem, a long-term problem or a periodic problem. The embodiment of the application mainly aims at the occasional problem.

[0115] Step 3: Abnormal determination algorithm confirms problem root cause. For the occasional problem, corresponding network type indicators and threshold type indicators are used for hierarchical judgment, and when the indicators of a certain level are abnormal, the indicators of the next level are judged, and the positioning range is gradually narrowed, and finally the problem root cause is located. Among them, only when the indicators of each layer are analyzed abnormally, the next layer of indicators is analyzed, such as when the service indicators are normal, the network indicators are not analyzed; if the service indicators control the abnormal service delay, there is no problem in the analysis of the delay of each layer, and the subsequent delay related indicators are not analyzed, the abnormal service delay is output, and the abnormal reason is irrelevant to the air interface delay.

[0116] First, end-to-end delay abnormality determination is performed, and based on the abnormal time point, the segment delay before and after the abnormal time point (such as 15 minutes) is analyzed to determine the location of the problem. Secondly, segment delay abnormality determination is performed, such as Figure 5 As shown, by analyzing the uplink and / or downlink UE delay, uplink and / or downlink MAC layer delay, uplink and / or downlink PDCP layer delay, uplink and / or downlink RLC delay, etc. Delay, determine the air interface position where the problem occurs, and exclude some problems in the initial screening, and then analyze the related indicators from four aspects of user state, channel condition, scheduling opportunity and resource utilization, focus on one or several problem categories, and further narrow the positioning range. Thirdly, air interface state indicator abnormality determination, according to the problem category determined in the last step, the related air interface state indicators are viewed, the positioning unit is gradually narrowed, and the problem root cause is confirmed. For details, refer to Figure 6 As shown.

[0117] In some optional embodiments of the application, the method further comprises: the first network device determines the location information corresponding to the abnormal problem based on the pre-obtained network topology, the mobile path of the terminal and the abnormal time information corresponding to the abnormal problem.

[0118] In the embodiment, the first network device can determine the location information of the abnormal problem according to the pre-obtained network topology, in combination with the mobile path of the terminal and the abnormal time information corresponding to the abnormal problem. The mobile path of the terminal can also be referred to as a service motion path, that is, the mobile path of the terminal in the service use process. In some application scenarios, for example, the terminal is an AGV, and in general, the AGV moves along a fixed path in the service use process. Therefore, the mobile path of the terminal can be pre-obtained by the first network device. For example, according to the air interface index positioning, the AGV service will suddenly have a coverage deterioration problem. Therefore, according to the vehicle movement trajectory, the location where the RSRP and SINR deterioration occurs can be determined, so that the network optimization can be performed in a targeted manner.

[0119] In some optional embodiments, the method further includes that the first network device receives the location information of the terminal sent by the second network device.

[0120] In the embodiment, the first network device can obtain the location information of the terminal from the second network device through an interface. Therefore, the first network device can determine the location information corresponding to the abnormal problem in combination with the location information of the terminal.

[0121] In some optional embodiments of the present application, the method further includes that the first network device stores the first index in the third type of index in a coverage storage manner, and only retains the first index of a specified time length or a specified data volume; and / or, the first network device stores the second index in the third type of index in a continuous storage manner; wherein the importance of the first index is lower than the importance of the second index, and the importance is related to the service or the service type.

[0122] In the embodiment, since the first network device obtains the second-level network index from the second network device, in order to optimize the memory space, the full amount of index needs to be tracked and collected in the early stage. When the tracked specific service is more, the storage capacity requirement of the device is higher. Therefore, based on the service requirement, the index can be divided into a regular index (such as the first index) and an important index (such as the second index). In the data processing process, if the regular index (such as the first index) is normal, the newly obtained data covers the data collected in the early stage, only the data in a period of time (the time length is also pre-set) or a certain amount of data is retained, and the important index (such as the second index) is continuously counted and saved.

[0123] In some optional embodiments, the third type of index is a third type of index of the first service; and the obtaining of the third type of index from the second network device comprises: receiving, by the first network device, the third type of index of the first service sent by the second network device; the third type of index of the first service is obtained by the second network device based on the first identifier of the first service; or the third type of index of the first service is obtained by the second network device based on the first service model corresponding to the first service; or the first network device receives the third type of index sent by the second network device, identifies the third type of index based on the first service model corresponding to the first service obtained in advance, and obtains the third type of index of the first service.

[0124] In some optional embodiments, the third type of index is a third type of index of the first user; and the obtaining of the third type of index from the second network device comprises: receiving, by the first network device, the third type of index of the first user sent by the second network device; the third type of index of the first user is obtained by the second network device based on the identifier of the first user; or the first network device receives the third type of index sent by the second network device, identifies the third type of index based on the identifier of the first user, and obtains the third type of index of the first user.

[0125] In the embodiment, the wireless network faults in the existing network are mainly divided into long-term problems, periodic problems and occasional problems, and the problem objects include all service users at the cell level, all users of a specific service, and users of a specific terminal. At present, the problem objects that are difficult to diagnose are mainly all users of a specific service and users of a specific terminal. For example, all services of a certain type are in a cell, and all services of this type have problems due to network abnormalities of the cell. For example, the users involved in a certain type of service are mobile, and the user has a problem because the user moves to a place with network abnormalities, or the user has no problem with the wireless network, and the terminal has a problem. Therefore, the automatic fault diagnosis can first identify a specific service, diagnose the network condition of the specific service, or first identify a specific user, identify the related network index of the specific user for fault analysis, or first identify a specific service, diagnose the network condition of the specific service, identify a specific user in the specific service when the specific user has a problem, and collect related indexes of the specific user.

[0126] For the identification of a specific service, the first network device can identify the service (denoted as the first service) through a first identifier of the service. For example, the first identifier can be a 5G QoS identifier (5QI) or a slice identifier (slice ID). As an implementation, the second network device can identify the service based on the first identifier of the first service, and then obtain the third type of index of the first service (the index of all users). As a second implementation, the second network device can identify a certain type of service based on a first service model (such as packet size, packet period, etc.) of the first service, and collect the index. As a third implementation, the first network device can obtain the service model (such as packet size, packet period, etc.) from the terminal in advance, and identify the specific service (such as the first service) based on the service model. The second network device transmits the full amount of the third type of index to the first network device. The first network device can calculate the relevant features (such as packet size, packet period, etc.) based on the received full amount of the third type of index, determine the corresponding service model, and determine the third type of index of the specific service. For example, the first network device can calculate the packet size and packet period based on the user-level second-level PDCP received packet number, PDCP or RLC throughput, etc. reported by the base station. For example, if the PDCP received packet number in one second is known, and the average throughput of PDCP or RLC is known, then the packet size is “throughput / received packet number”, and the packet period is “received packet number 60 / s”. The third type of index of the specific service can be obtained by screening the third type of index obtained from the base station. Figure 7

[0127] For the identification of a specific user, as an implementation, the second network device has the ability to unpack, and the third type of index can be associated with the identifier of the first user. The second network device determines the first user based on the identifier of the first user, and then identifies and obtains the third type of index of the first user. As another implementation, the first network device has the ability to unpack, and the third type of index can be associated with the identifier of the first user. After receiving the full amount of the third type of index sent by the second network device, the first network device can screen the full amount of the third type of index based on the identifier of the first user, and obtain the third type of index of the first user. For example, the identifier of the first user can be one or more of the following identifiers: short-term mobile subscriber identity (S-TMSI), UE tunnel endpoint identifier (TEID), NG APID, IP address, etc.

[0128] ​Based on the above embodiment, the embodiment of the application further provides a fault diagnosis processing method. Figure 8 A flowchart of the fault diagnosis processing method of the embodiment of the application is shown in Figure 2 Figure 8 As shown in the figure, the method comprises the following steps.

[0129] Step 201: The second network device sends a third type of index to the first network device, the third type of index being a second-level network index; the third type of index is used by the first network device to determine an index anomaly and perform hierarchical analysis on the index anomaly, to obtain an analysis result, the analysis result including a problem root cause corresponding to the abnormal index.

[0130] In some optional embodiments, the third type of index is a third type of index of a first service; before the second network device sends the third type of index to the first network device, the method further comprises: the second network device identifies the index based on a first identifier of the first service, to obtain the third type of index of the first service; or, the second network device identifies the index based on a first service model corresponding to the first service, to obtain the third type of index of the first service.

[0131] In other optional embodiments, the third type of index is a third type of index of a first user; before the second network device sends the third type of index to the first network device, the method further comprises: the second network device identifies the index based on an identifier of the first user, to obtain the third type of index of the first user.

[0132] In some optional embodiments, the method further comprises: the second network device sends location information of a terminal to the first network device.

[0133] The fault diagnosis processing method of the embodiment of the application is described below in combination with a specific example.

[0134] 1. Input of service demand and index of a certain port unmanned truck

[0135] Unmanned trucks are used for cargo transportation in a port factory based on a 5G network, the service network demand and rate demand being 16 kbps-6 Mbps, and the time delay reliability requirement being 100 ms@99.99%. The service model is that the discovery unmanned truck service packet sending period is 1 s, the packet sending size is 30 Bytes, and the average time delay is about 20 ms, and there are occasional 100 ms or more large time delay points.

[0136] Therefore, the service type index (i.e., the first type of index) comprises:

[0137] Terminal information: card number, IP address, trance id, etc. ​

[0138] Service type: Video_PLC Biz (control type service)

[0139] Service requirement index: De_Delay = 200 ms; Ul_throughput = 2 Mbps; Dl_throughput = 2 Mbps.

[0140] The threshold type index (i.e., the second type index) can be referred to Table 4.

[0141] Table 4

[0142]

[0143]

[0144]

[0145] The network type index (i.e., the third type index) includes: the first network device connects the port park base station device, and the park base station transmits the second-level air interface index through the interface.

[0146] 2. Fault diagnosis and recovery

[0147] 1) Fault diagnosis: the first network device analyzes the static parameter configuration, RRC signaling, service channel, and control channel type index from the first level based on the second-level air interface index reported by the base station. The delay anomaly is caused by limited control resources, and the single-user control resource occupation and terminal downlink control information (DCI, Downlink Control Information) missing detection cause the control resource to be limited. The analysis steps are as follows:

[0148] Step 1: First-level service requirement index analysis, based on the service index, diagnose that at 2023-10-1011:26:42, Ping_Delay (end-to-end delay) = 412.1 ms, which is greater than the service requirement index De_Delay = 200 ms. Determine that the service delay is abnormal and is an occasional problem, and perform the next layer of data index analysis;

[0149] Step 2: Second-level segmented delay index analysis, based on the threshold index, analyze the terminal D1 delay, uplink and downlink MAC layer delay, uplink and downlink RLC layer delay, and uplink and downlink PDCP layer delay within 15 min before and after the abnormal time point, determine that only the terminal delay D1 is abnormal (BSR index is 11:26:41~11:26:42, up to 451402192), and other layer delays are normal, then further analyze the next layer of terminal D1 delay abnormality related index;

[0150] Step 3: Third layer static index and RRC signaling analysis, static parameters such as pre-scheduling start, no SR period too large problem, discontinuous reception (DRX) function off, DRX sleep stop scheduling impact; RRC signaling link and switching situation is normal, further analysis of dynamic network index data;

[0151] Step 4: Fourth layer dynamic second-level network index analysis:

[0152] Control channel data analysis, UE level CCE allocation failure rate is high (up to 334 at 11:26:41-11:26:42) -> UE level CCE aggregation level is high (up to 13 at 11:26:41-11:26:42) -> Downlink RSRP / SINR is normal, UE level DTX is abnormal (up to 30% at 11:26:41-11:26:42), positioning as terminal DCI missed detection and single user control resource restriction caused by terminal side stack caused by large delay problem.

[0153] Service channel data analysis, within 15min before and after the abnormal time point, UE level PDCP layer rate is normal, MCS and UE occupied RB number is normal, it is judged that there is no problem in service channel transmission.

[0154] Output problem root cause and solution, including two problems: Problem 1: Occasional terminal missed detection leads to high single terminal CCE aggregation level, leading to control channel resource restriction. Fault root cause: Occasional terminal missed detection DCI. Solutions include: 1. Upgrade terminal version; 2. Single user CCE aggregation level optimization algorithm; 3. Based on 5QI / slice specific business priority lifting; 4. Cell CCE available symbol number improvement; 5. SPS / CGI technology; 6. Control resource allocation optimization algorithm. Problem 2: Occasional uplink interference / noise increase leads to service channel transmission rate reduction. Fault root cause: Occasional uplink interference / noise, service rate reduction. Solutions include: 1. Adopt series scheduling optimization technology to reduce interference influence; 2. Dual path technology: double sending and single receiving. Specific reference can be made to Figure 9 .

[0155] 2) Solution: Control resource restriction can be alleviated on the one hand by expanding the available control resources of the network, and on the other hand by reducing the occupation of single user control resources, including terminal solutions and network avoidance solutions.

[0156] a. After expanding the available control resources of the network, the end-to-end delay performance is obviously improved, the delay reliability is improved from 100ms@99.9% to 100ms@99.99%, and the maximum delay is reduced from 412ms to 130ms; At the same time, the control resource allocation failure rate is also alleviated.

[0157] b. After upgrading the chip and adjusting the control resource enhancement algorithm on the network side, the proportion of control resources used by the terminal is significantly reduced, the end-to-end latency is further improved, the latency reliability reaches 100ms@99.99%, and the maximum latency is 73ms.

[0158] Based on the above embodiments, this invention also provides a fault diagnosis and processing device, which is applied to a first network device. Figure 10 This is a schematic diagram of the composition and structure of the fault diagnosis and processing device according to an embodiment of the present invention. Figure 1 ;like Figure 10 As shown, the device includes: an acquisition unit 11 and a diagnostic unit 12; wherein,

[0159] The acquisition unit 11 is used to acquire a first type of indicator and a second type of indicator, and to acquire a third type of indicator from the second network device; the first type of indicator is a service requirement-related indicator, the second type of indicator includes threshold values ​​of network indicators and / or service indicators, and the third type of indicator is a second-level network indicator.

[0160] The diagnostic unit 12 is used to determine that an indicator is abnormal when the third type of indicator does not meet the first type of indicator; to perform hierarchical analysis on the indicator abnormality based on the third type of indicator and the second type of indicator to obtain analysis results; the analysis results include the root causes of the problem corresponding to the abnormal indicator.

[0161] In some optional embodiments of the present invention, the diagnostic unit 12 is further configured to determine the first business type corresponding to the third type of indicator based on the characteristics of the third type of indicator.

[0162] In some optional embodiments of the present invention, the diagnostic unit 12 is used to determine that the indicator is abnormal when the third type of indicator does not meet the first type of indicator corresponding to the first business type.

[0163] In some optional embodiments of the present invention, the third type of indicator is a third type of indicator of the first service; the acquisition unit 11 is used to receive the third type of indicator of the first service sent by the second network device; the third type of indicator of the first service is obtained by the second network device by identifying the indicator based on the first identifier of the first service, or the third type of indicator of the first service is obtained by the second network device by identifying the indicator based on the first service model corresponding to the first service; or, it is used to receive the third type of indicator sent by the second network device, identify the third type of indicator based on the first service model corresponding to the first service obtained in advance, and obtain the third type of indicator of the first service.

[0164] In some optional embodiments of the present application, the third type of index is the third type of index of the first user; the obtaining unit 11 is configured to receive the third type of index of the first user sent by the second network device; the third type of index of the first user is obtained by the second network device based on the identification of the first user; or, configured to receive the third type of index sent by the second network device, identify the third type of index based on the identification of the first user, and obtain the third type of index of the first user.

[0165] In some optional embodiments of the present application, the diagnosis unit 12 is configured to determine abnormal time information corresponding to the abnormal index, determine the first type to which the problem belongs based on the abnormal time information, obtain the first processing logic corresponding to the first type, and perform hierarchical analysis on the index abnormality based on the third type of index and the second type of index according to the first processing logic, wherein the first type is an occasional problem, a periodic problem, or a long-term problem.

[0166] In some optional embodiments of the present application, the diagnosis unit 12 is configured to sequentially compare the third type of index and the second type of index corresponding to each level, and obtain an analysis result based on the comparison result, wherein the third type of index compared in the next level is associated with the third type of index that appears abnormal in the comparison result of the previous level, and the problem range indicated by the comparison result of the next level is reduced compared with the previous level.

[0167] In some optional embodiments of the present application, the diagnosis unit 12 is configured to compare the third type of index and the second type of index corresponding to the first level to obtain a first comparison result, wherein the first comparison result includes end-to-end delay abnormal information; compare the third type of index and the second type of index corresponding to the second level to obtain a second comparison result, wherein the second comparison result includes segmented delay abnormal information; the segmented delay includes one or more of the following: uplink and / or downlink UE delay, uplink and / or downlink MAC layer delay, uplink and / or downlink PDCP layer delay, uplink and / or downlink RLC delay; compare the third type of index and the second type of index corresponding to the third level based on the segmented delay abnormal information to obtain a third comparison result, wherein the third comparison result includes a first problem category to which the index abnormality belongs, and the first problem category includes one or more of the following: user state problem, channel condition problem, scheduling opportunity problem, and resource limitation problem; compare the third type of index and the second type of index corresponding to at least one fourth level based on the first problem category, and determine the root cause of the problem based on the comparison result.

[0168] In some optional embodiments of the present application, the diagnostic unit 12 is further configured to determine the location information corresponding to the abnormal problem based on the pre-obtained network topology, the moving path of the terminal and the abnormal time information corresponding to the abnormal problem.

[0169] In some optional embodiments of the present application, the obtaining unit 11 is further configured to receive the location information of the terminal sent by the second network device.

[0170] In some optional embodiments of the present application, the apparatus further comprises a determining unit 13 configured to determine the solution corresponding to the problem root cause according to a first mapping relationship, wherein the first mapping relationship comprises a mapping relationship between a plurality of problem root causes and solutions.

[0171] In some optional embodiments of the present application, the apparatus further comprises a storage unit configured to store the first indicator in the third type of indicators in a way of overlay storage, and only retain the first indicator of a specified time length or a specified data volume; and / or store the second indicator in the third type of indicators in a way of continuous storage; wherein the importance degree of the first indicator is lower than that of the second indicator, and the importance degree is related to a service or a service type.

[0172] In some optional embodiments of the present application, the third type of indicators comprises one or more of the following: dynamic air interface performance indicators, air interface state indicators, static indicators, air interface signaling, RSRP of DMRS and / or SRS, SINR of DMRS and / or SRS, SINR of each data stream in a multi-data stream case.

[0173] In the embodiments of the present application, the diagnostic unit 12 and the determining unit 13 in the apparatus can be implemented by a central processing unit (CPU), a digital signal processor (DSP), a microcontroller unit (MCU) or a field-programmable gate array (FPGA) in actual application; and the obtaining unit 11 in the apparatus can be implemented by the CPU, the DSP, the MCU or the FPGA in combination with a communication module (including a basic communication suite, an operating system, a communication module, a standardized interface and a protocol, etc.) and a transceiving antenna in actual application.

[0174] The embodiments of the present application further provide a fault diagnosis processing apparatus, which is applied to a second network device. Figure 11 The composition structure of the fault diagnosis processing apparatus of the embodiments of the present application is shown in Fig. 1. Figure 2 The composition structure of the fault diagnosis processing apparatus of the embodiments of the present application is shown in Fig. 1. Figure 11As shown, the apparatus comprises a communication unit 21 configured to send a third type of index to the first network device, the third type of index being a second-level network index; the third type of index is used by the first network device to determine an index anomaly and perform hierarchical analysis on the index anomaly to obtain an analysis result, the analysis result including a problem root cause corresponding to the abnormal index.

[0175] In some optional embodiments of the present application, the third type of index is a third type of index of a first service; the apparatus further comprises a processing unit 22 configured to, before the communication unit 21 sends the third type of index to the first network device, identify the index based on a first identifier of the first service to obtain the third type of index of the first service; or identify the index based on a first service model corresponding to the first service to obtain the third type of index of the first service.

[0176] In some optional embodiments of the present application, the third type of index is a third type of index of a first user; the apparatus further comprises a processing unit 22 configured to, before the communication unit 21 sends the third type of index to the first network device, identify the index based on an identifier of the first user to obtain the third type of index of the first user.

[0177] In some optional embodiments of the present application, the communication unit 21 is further configured to send position information of a terminal to the first network device.

[0178] In the embodiments of the present application, the processing unit 22 in the apparatus can be implemented by a CPU, a DSP, a MCU or an FPGA in actual application; and the communication unit 21 in the apparatus can be implemented by a communication module (including a basic communication suite, an operating system, a communication module, a standardized interface and a protocol, etc.) and a transceiving antenna in actual application.

[0179] It should be noted that the fault diagnosis processing apparatus provided in the above embodiments is only taken as an example for the division of the above program modules in the fault diagnosis processing, and in actual application, the above processing can be completed by different program modules according to needs, that is, the internal structure of the apparatus is divided into different program modules to complete all or part of the above processing. In addition, the fault diagnosis processing apparatus and the fault diagnosis processing method provided in the above embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be described here.

[0180] The embodiments of the present application further provide a network device, which is a first network device or a second network device. Figure 12 The hardware composition structure of the network device of the embodiments of the present application is shown in FIG. 2. Figure 12As shown, the communication device comprises a memory 32, a processor 31, and a computer program stored in the memory 32 and executable on the processor 31, wherein the processor 31 implements the steps of the fault diagnosis processing method applied to the first network device or the second network device according to the embodiments of the present application when executing the program.

[0181] Optionally, the network device can further comprise at least one network interface 33. In the network device, various components are coupled together through a bus system 34. It can be understood that the bus system 34 is used to realize the connection communication between the components. In addition to the data bus, the bus system 34 also includes a power bus, a control bus, and a status signal bus. However, in order to clearly illustrate, all kinds of buses are marked as the bus system 34 in the following description. Figure 12

[0182] ​It can be appreciated that the memory 32 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a Read Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electrically Erasable Programmable Read-Only Memory (EEPROM), a Ferromagnetic Random Access Memory (FRAM), a Flash Memory, a magnetic surface memory, an optical disc, or a Compact Disc Read-Only Memory (CD-ROM). The magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a Random Access Memory (RAM) used as an external cache. By way of example but not limitation, many forms of RAM can be used, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), Sync Link Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memory 32 described in the embodiments of the present application is intended to include, but not limited to, these and any other suitable type of memory.

[0183] The method disclosed in the embodiments of the present application can be applied in the processor 31 or implemented by the processor 31. The processor 31 can be an integrated circuit chip having a processing capability of signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of hardware in the processor 31 or the instruction in the form of software. The processor 31 described above can be a general processor, a DSP, or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc. The processor 31 can implement or execute the disclosed methods, steps and logic block diagrams in the embodiments of the present application. The general processor can be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiments of the present application, the execution can be directly completed by the hardware decoding processor or by the combination of hardware and software modules in the decoding processor. The software module can be located in the storage medium, which is located in the memory 32. The processor 31 reads the information in the memory 32 and combines the hardware to complete the steps of the above method.

[0184] In the exemplary embodiments, the network device can be implemented by one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), FPGAs, general-purpose processors, controllers, MCUs, microprocessors (Microprocessors), or other electronic elements, for executing the above-mentioned methods.

[0185] In the exemplary embodiments, the embodiments of the present application also provide a computer readable storage medium, such as the memory 32 including a computer program, which can be executed by the processor 31 of the network device to complete the steps of the above-mentioned method. The computer readable storage medium can be FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM, etc. The computer readable storage medium can also be various devices including one or any combination of the above-mentioned memories.

[0186] The computer readable storage medium provided by the embodiments of the present application has a computer program stored thereon, which is executed by the processor to implement the steps of the fault diagnosis processing method applied in the first network device or the second network device in the embodiments of the present application.

[0187] The embodiment of the present application further provides a computer program product, comprising a computer program, which can be executed by a computer device (such as the processor 41 of the network device) to complete the steps of any of the preceding fault diagnosis processing methods.

[0188] The methods disclosed in the several method embodiments provided by the present application can be combined arbitrarily without conflict to obtain new method embodiments.

[0189] The features disclosed in the several product embodiments provided by the present application can be combined arbitrarily without conflict to obtain new product embodiments.

[0190] The features disclosed in the several method or device embodiments provided by the present application can be combined arbitrarily without conflict to obtain new method embodiments or device embodiments.

[0191] In the several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the various components shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.

[0192] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or distributed on a plurality of network units; some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0193] In addition, each functional unit in each embodiment of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be realized in the form of hardware, or in the form of hardware plus software functional unit.

[0194] Those skilled in the art can understand that all or part of the steps of the above-mentioned method embodiments can be completed by program instruction related hardware, and the foregoing program can be stored in a computer readable storage medium, and the program is executed to execute the steps of the above-mentioned method embodiments; and the foregoing storage medium includes mobile storage equipment, ROM, RAM, magnetic disc or optical disc and various storage program codes.

[0195] Alternatively, the above-mentioned integrated unit of the present application, if realized in the form of a software function module and sold or used as an independent product, can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solutions of the embodiments of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: mobile storage devices, ROM, RAM, magnetic disks or optical disks, and various media that can store program codes.

[0196] The above describes only the specific implementation of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A failure diagnosis processing method characterized by comprising: The method is applied to a first network device, and the method comprises: The first network device obtains a first type of index and a second type of index, and obtains a third type of index from a second network device; the first type of index is a service demand related index, the second type of index comprises a threshold value of a network index and / or a service index, and the third type of index is a second-level network index; When the third type of index does not meet the first type of index, it is determined that the index is abnormal; Based on the third type of index and the second type of index, hierarchical analysis is performed on the index abnormality to obtain an analysis result; the analysis result comprises a problem root cause corresponding to the abnormal index; The first network device determines abnormal time information corresponding to the abnormal index, determines a first type to which the problem belongs based on the abnormal time information, obtains a first processing logic corresponding to the first type, and performs hierarchical analysis on the index abnormality based on the third type of index and the second type of index according to the first processing logic; the first type is an occasional problem, a periodic problem or a long-term problem; The first network device sequentially compares the third type of index and the second type of index corresponding to each level, and obtains an analysis result based on the comparison result; the third type of index compared in the next level is associated with the third type of index whose comparison result in the previous level is abnormal; the problem range indicated by the comparison result in the next level is reduced compared with the comparison result in the previous level. The method further comprises: The first network device determines a first service type corresponding to the third type of index based on the characteristics of the third type of index.

2. The method of claim 1, wherein, When the third type of index does not meet the first type of index, the first network device determines that the index is abnormal. The third type of index is a third type of index of a first service; the first network device receives the third type of index of the first service sent by the second network device; the third type of index of the first service is obtained by the second network device based on a first identifier of the first service, or the third type of index of the first service is obtained by the second network device based on a first service model corresponding to the first service; 3. The method of claim 2, wherein, Or, the first network device receives the third type of index sent by the second network device, identifies the third type of index based on a first service model corresponding to the first service obtained in advance, and obtains the third type of index of the first service. The third type of index is a third type of index of a first user; the first network device receives the third type of index sent by the second network device, identifies the third type of index based on a first service model corresponding to the first service obtained in advance, and obtains the third type of index of the first service.

4. The method of claim 1, wherein, ​ ​ ​ 5. The method of claim 1, wherein, ​ The first network device receives the third type of index of the first user sent by the second network device; the third type of index of the first user is obtained by the second network device based on identification of the first user; or, The first network device receives the third type of index sent by the second network device, identifies the third type of index based on the identification of the first user, and obtains the third type of index of the first user.

6. The method of claim 1, wherein, The third type of index and the second type of index are compared based on the third type of index and the second type of index, and an analysis result is obtained, including: The first network device compares the third type of index and the second type of index corresponding to the first level, and obtains a first comparison result, which includes end-to-end delay anomaly information; The third type of index and the second type of index corresponding to the second level are compared, and a second comparison result is obtained, which includes segment delay anomaly information; the segment delay includes one or more of the following: uplink and / or downlink user equipment (UE) delay, uplink and / or downlink medium access control (MAC) layer delay, uplink and / or downlink packet data convergence protocol (PDCP) layer delay, and uplink and / or downlink radio link control (RLC) delay; According to the segment delay anomaly information, the third type of index and the second type of index corresponding to the third level are compared, and a third comparison result is obtained, which includes a first problem category to which the index anomaly belongs, and the first problem category includes one or more of the following: user state problem, channel condition problem, scheduling opportunity problem, and resource limitation problem; According to the first problem category, the third type of index and the second type of index corresponding to at least one fourth level are compared, and a problem root cause is determined based on the comparison result.

7. The method of claim 1, wherein, The method further includes: The first network device determines location information corresponding to the abnormal problem based on the pre-obtained network topology, the mobile path of the terminal, and the abnormal time information corresponding to the abnormal problem.

8. The method of claim 7, wherein, The method further includes: The first network device receives the location information of the terminal sent by the second network device.

9. The method of claim 1, wherein, The method further includes: The first network device determines a solution corresponding to the problem root cause according to a first mapping relationship; wherein the first mapping relationship includes a mapping relationship between a plurality of problem root causes and solutions.

10. The method of claim 1, wherein, The method further includes: The first network device stores the first index in the third type of index in a coverage storage manner, and only retains the first index with a specified time length or a specified data volume; and / or, The first network device stores the second index in the third type of index in a continuous storage manner; The importance of the first index is lower than that of the second index, and the importance is related to a service or a service type.

11. The method of claim 1, wherein, The third type of index includes one or more of the following: dynamic air interface performance index, air interface state index, static index, air interface signaling, reference signal received power (RSRP) of demodulation reference signal (DMRS) and / or channel sounding reference signal (SRS), signal to interference plus noise ratio (SINR) of DMRS and / or SRS, and SINR of each data stream in a multi-data stream case.

12. A failure diagnosis processing method characterized by comprising: The method is applied to a second network device, and the method includes: The second network device sends a third type of index to the first network device, the third type of index being a second-level network index; the third type of index is used by the first network device to perform the method of any one of claims 1 to 11 to obtain a problem root cause corresponding to an abnormal index.

13. The method of claim 12, wherein, The third type of index is a third type of index of a first service; before the second network device sends the third type of index to the first network device, the method further includes: The second network device identifies the index based on a first identifier of the first service to obtain the third type of index of the first service; or The second network device identifies the index based on a first service model corresponding to the first service to obtain the third type of index of the first service.

14. The method of claim 12, wherein, The third type of index is a third type of index of a first user; before the second network device sends the third type of index to the first network device, the method further includes: The second network device identifies the index based on an identifier of the first user to obtain the third type of index of the first user.

15. The method of claim 12, wherein, The method further includes: The second network device sends location information of a terminal to the first network device.

16. A failure diagnosis processing apparatus characterized by comprising: The apparatus is applied to a first network device, and the apparatus includes an acquisition unit and a diagnosis unit; wherein The acquisition unit is configured to obtain a first type of index and a second type of index, and to obtain a third type of index from a second network device; the first type of index is a service demand related index, the second type of index includes a threshold value of a network index and / or a service index, and the third type of index is a second-level network index; The diagnosis unit is configured to determine that an index is abnormal when the third type of index does not meet the first type of index, to perform hierarchical analysis on the index abnormality based on the third type of index and the second type of index to obtain an analysis result, and to include a problem root cause corresponding to an abnormal index in the analysis result; The diagnosis unit is configured to determine abnormal time information corresponding to an abnormal index, to determine a first type to which a problem belongs based on the abnormal time information, to obtain a first processing logic corresponding to the first type, to determine the first type as an occasional problem, a periodic problem, or a long-term problem, and to perform hierarchical analysis on the index abnormality based on the third type of index and the second type of index according to the first processing logic. The diagnosis unit is configured to compare the third type of index and the second type of index corresponding to each level in sequence, and to obtain the analysis result based on a comparison result; a third type of index compared in a next level is associated with a third type of index that appears abnormal in a comparison result of a previous level; a problem range indicated by the comparison result of the next level is reduced compared to the previous level.

17. A failure diagnosis processing apparatus characterized by comprising: The device is applied to a second network device, and the device comprises a communication unit configured to send a third type of index to the first network device, the third type of index being a second-level network index; the third type of index is used for the first network device to perform the method of any one of claims 1 to 11 to obtain an abnormal index corresponding problem root cause.

18. A computer readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the steps of the method of any one of claims 1 to 11; or, The program is executed by the processor to implement the steps of the method of any one of claims 12 to 15.

19. A network device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, The processor executes the program to implement the steps of the method of any one of claims 1 to 11; or, The processor executes the program to implement the steps of the method of any one of claims 12 to 15.

20. A computer program product, characterised in that, The computer program instructions cause the computer to perform the steps of the method of any one of claims 1 to 11; or, The computer program instructions cause the computer to perform the steps of the method of any one of claims 12 to 15.

Citation Information

Patent Citations

  • Fault positioning method and system, and computer readable storage medium

    CN115988243A

  • Micro-service system fault diagnosis and root cause positioning method

    CN116450399A