5g network control plane monitoring
By incorporating trace context information elements with trace IDs into 5G network monitoring, the method addresses the challenges of real-time monitoring in complex 5G networks, enhancing scalability and reducing storage needs.
Patent Information
- Application Number
- JP2023183135
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-25
- Publication Date
- 2025-05-12
AI Technical Summary
In 5G networks, real-time monitoring and analysis of microservices become challenging due to the increasing complexity and volume of application logs, leading to difficulties in detecting application failures and performance degradation.
A new monitoring method that utilizes trace context information elements with trace IDs in messages between UE, (R)AN, or UPF and control plane network functions, allowing for real-time processing and analysis of monitoring data without relying on communication session details.
This approach reduces storage costs by collecting only the minimum necessary data, enables immediate processing and collection, and facilitates real-time service monitoring, improving scalability and granularity of service monitoring.
Smart Images

Figure 2025072796000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to terminal-level monitoring in the control plane of a 5G network. [Background technology]
[0002] 5G (fifth generation mobile communication system) network services began in March 2020. 5G is characterized by high speed and large capacity, multiple simultaneous connections, and low latency. Taking advantage of these features, it is predicted that not only mobile phone users but also machines, objects (e.g., automobiles), and devices will be connected to the 5G network. Summary of the Invention [Problem to be solved by the invention]
[0003] Figure 1 shows communication from a terminal through a 5G network to the system domain. A thick arrow represents a single transaction. A single transaction passes through multiple applications; for example, app #1 calls app #2. Figure 1 shows a situation where a failure occurs in app #2.
[0004] As microservices become more widespread in the future, it is expected that E2E (End to End) processes and databases will become more complexly interconnected. Also, the volume of monitoring data (application logs, etc.) is expected to increase, making real-time monitoring and analysis more difficult.
[0005] The first problem is that real-time processing becomes difficult due to the growing volume of application logs. As the volume of application logs increases and the processing load increases, real-time detection processing becomes difficult. This is expected to make it more difficult to detect application failures and performance degradation.
[0006] The second problem is that it becomes difficult to detect faults and check normality from an E2E perspective, including the network. Existing monitoring is achieved through external monitoring and log analysis. However, the increasing complexity of configurations and the bloat of monitoring data creates scalability issues. In other words, it is not possible to fine-tune the granularity of service monitoring (service or terminal unit, etc.).
[0007] Service monitoring in a microservices environment requires monitoring data from both the platform and the services / applications, as well as a mechanism for efficiently analyzing massive amounts of data.
[0008] Current monitoring relies on detailed communication session information (CDL: Call Detail Log), which is distributed across control plane components. For E2E normality checks and fault analysis, it is necessary to combine CDLs to create detailed information for the entire session. Another problem is that the CDL contains a large amount of information that is unnecessary for fault analysis.
[0009] Therefore, the current monitoring method using CDL has the following two problems: Collecting and analyzing huge amounts of CDL requires a large amount of storage. - Log merging and creating merging logic takes time, making real-time performance difficult.
[0010] The present invention proposes a new monitoring method that does not rely on CDL. The new monitoring method uses tracing to monitor the E2E control plane processing path and the relationships between each component. Each span is processed and sent immediately after processing is complete, allowing for real-time analysis and monitoring. Only the minimum amount of data necessary for service monitoring is provided.
[0011] Due to the above features, the present invention improves the following points. By collecting the minimum amount of monitoring data required for service monitoring, storage costs can be reduced. - Real-time service monitoring is possible at all times because data is processed and collected immediately. [Means for solving the problem]
[0012] The inventors discovered that a field for a trace context information element containing a trace ID can be provided in a message between a UE, (R)AN, or UPF and a control plane network function, and thereby completed the present invention.
[0013] (1) A 5G system in which a trace context information element field containing a trace ID is provided in messages between a UE, (R)AN, or UPF and a control plane network function. (2) A 5G system according to (1), in which a trace context information element field including a trace ID is provided in an NGAP protocol and / or PFCP protocol message. (3) A 5G system (1) that assigns SUPI to span attributes in AMF or UDM. (4) A monitoring method in a 5G system, comprising providing a field of a trace context information element including a trace ID in a message between a UE, (R)AN, or UPF and a network function of a control plane. (5) A monitoring method according to (4), in which a field of a trace context information element including a trace ID is provided in a message of the NGAP protocol and / or the PFCP protocol. (6) In AMF or UDM, a monitoring method of (4) in which SUPI is assigned to the span attribute. [Effects of the Invention]
[0014] According to the present invention, monitoring can be achieved for each terminal in the control plane. [Brief explanation of the drawings]
[0015] [Figure 1] This figure shows one transaction from the terminal / network domain to the system domain. [Figure 2] FIG. 1 is a diagram illustrating the configuration of a 5G system. [Figure 3] FIG. 10 is a diagram illustrating a definition of a trace context IE according to an embodiment of the present invention. [Figure 4] FIG. 2 is a diagram illustrating an operation sequence according to an embodiment of the present invention. [Figure 5] FIG. 10 is a diagram illustrating an example of transaction information according to the embodiment of the present invention. [Figure 6] FIG. 10 illustrates a spike in API failure rates in an embodiment of the present invention. [Figure 7] FIG. 10 is a diagram illustrating a search for a cause of a failure in an embodiment of the present invention. [Figure 8] FIG. 10 is a diagram showing a procedure for investigating the cause of a failure in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0016] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. FIG. 2 is a diagram showing the configuration of the 5G system 100. In the 5G system 100, each node is called a network function (NF).
[0017] Here, we will briefly explain the official names of the main NFs and their functions. AF (Application Function: external application server) AMF (Access and Mobility Management Function: subscriber authentication, security, terminal location management) AUSF (Authentication Server Function: subscriber authentication server) NRF (Network Repository Function: a function that registers the services of each NF) NSSF (Network Slice Selection Function: SMF selection for each slice) ·PCF (Policy Control Function) SMF (Session Management Function) ·UDM (Unified Data Management: retention of subscriber-related information) ·UDR (Unified Data Repository) UE (User Equipment) (R)AN (Radio Access Network) UPF (User Plane Function: packet transfer of user data) DN (Data Network: Data network outside the 5G system (Internet, etc.))
[0018] The area enclosed by the closed curve marked with the reference numeral 120 in Figure 2 is called the "5G core network." Within the 5G core network 120, the UPF is called the user plane (U-plane), and the part other than the UPF is called the control plane (C-plane). For example, the UPF of the user plane only forwards user data packets, while the SMF of the control plane establishes and disconnects sessions. In this way, the NF of the user plane processes data forwarding, and the NF of the control plane performs various control processes.
[0019] Communications in the 5G system 100 are divided into user plane communications and control plane communications. Communications between UEs, (R)ANs, UPFs, and DNs are communications for carrying user data and therefore can be considered user plane communications. In contrast, communications between NFs in the 5G core network 120 and between AMFs and UEs or (R)ANs are communications for control processing and therefore can be considered control plane communications.
[0020] Processing in the control plane is realized by repeating requests and responses between each NF. HTTP (Hypertext Transfer Protocol) is used as the protocol between NFs in the control plane.
[0021] However, for example, the (R)AN is connected to the AMF of the control plane via the N2 interface, which uses a protocol called NGAP (NG Application Protocol). The UPF of the user plane is connected to the SMF of the control plane via the N4 interface, which uses a protocol called PFCP (Packet Forwarding Control Protocol).
[0022] The W3C (World Wide Web Consortium) defines TraceContext, and has standardized embedding it in HTTP headers. According to the W3C definition of TraceContext, it consists of TraceParent and TraceState. TraceParent has the following four fields:
[0023] (1) Version This is a 1-byte field. The current specification specifies that the value should be "00" in hexadecimal. (2) Trace ID (trace-id) This is a 16-byte field. The system is configured so that messages from the same transaction have the same trace ID value. Therefore, messages with the same trace ID value can be determined to be messages from the same transaction. (3) Parent ID (parent-id) It is an 8-byte field. When creating a new span, this parent ID is set as the parent span ID (Span-id). (4) Trace flags This is a 1-byte field that contains information such as whether to sample or not and what the trace level is.
[0024] Tracestate contains vendor specific information. Since it is standardized to embed trace context in the HTTP header, for messages between NFs in the control plane, it is possible to determine which transaction the message belongs to by referring to the trace ID in the HTTP header.
[0025] However, NGAP on the N2 interface and PFCP on the N4 interface are not HTTP, so trace IDs are not embedded in the messages. Therefore, in the prior art, messages via the N2 interface and the N4 interface and messages between NFs in the control plane cannot be combined into one transaction.
[0026] In this embodiment, a field for a trace context information element (IE) is provided for messages sent via the N2 interface or the N4 interface, and a trace ID (trace-id) is included in the trace context IE.
[0027] First, let's explain transaction information (Trace). Transaction information (Trace) consists of multiple spans (Spans) that are related to each other within a transaction. A new span may be generated from a span, in which case the original "span" becomes the "parent" of the newly generated span. Each span contains the following information:
[0028] (a) Span ID: An ID that uniquely identifies a span (i) Trace ID: An ID that uniquely identifies the trace to which the span belongs. (c) Parent ID: ID that uniquely identifies the parent span of the span in question (iv) Attributes: A list of pairs of keys and their values (in this embodiment, the SUPI (IMSI) is stored here.) (E) Start Timestamp: The start time of the span. (f) End Timestamp: End time of the span
[0029] Each NF in the control plane is instrumented with OTEL (OpenTelemetry). By instrumenting each NF in this way, when there is communication between NFs in the control plane, each NF that supports trace context propagation immediately sends the processing time and an ID to identify the processing to the backend. Therefore, it is possible to monitor the processing time of each NF and where the processing is stopped in real time and end-to-end.
[0030] FIG. 3 shows a definition of a trace context IE (Trace Context Information Element) in this embodiment. The Trace Context IE has the following fields:
[0031] (a) Version (b) Trace ID (c) Parent ID (d) Trace Flags
[0032] The important thing here is that the trace context IE has a trace ID. Unlike the W3C trace context, the trace context IE does not have a trace state field. In this embodiment, the NGAP protocol message and the PFCP protocol message have a trace context IE field.
[0033] FIG. 4 is a diagram for explaining that a Subscription Permanent Identifier (SUPI) (International Mobile Subscriber Identity (IMSI)) is assigned to the attribute of a span in this embodiment. SUPI is a subscriber ID for 5G, equivalent to the IMSI (International Mobile Subscriber ID) used in 4G and earlier.
[0034] SUCI (Subscription Concealed Identifier) is the MSIN (Mobile Subscriber Identification Number) portion of SUPI encrypted with the public key of the home network. The Globally Unique Temporary Identity (GUTI) is used to uniquely identify a terminal (UE) when it communicates with the network. GUTI is used instead of IMSI between the terminal and AMF, protecting the terminal's privacy.
[0035] In step S1 of FIG. 4, the terminal (UE) encrypts the SUPI (IMSI) to generate the SUCI. In step S2, the base station (gNB) sends a registration request including the SUCI. In step S3, the AMF sends an authentication request including the SUCI. In step S4, the UDM decrypts the SUCI to obtain the SUPI.
[0036] In step S5, the UDM sends an authentication response including the SUPI. In step S6, the AMF generates a GUTI and maintains a map showing the correspondence between the GUTI and the SUPI. In steps S7 and S8, the GUTI is used instead of the IMSI in the communication between the terminal and the AMF. In step S9, the AMF sends an authentication request including the SUPI. In step S10, the UDM sends an authentication response that includes the SUPI.
[0037] Since the AMF and UDM can obtain the SUPI, in this embodiment, the AMF or UDM assigns the SUPI (IMSI) to the span attributes.
[0038] FIG. 5 is a diagram illustrating a state in which one transaction includes spans of multiple parent-child relationships. Reference numeral 400 is the data analysis platform for one transaction. A span named "App #1" exists as a child of a span named "Client." Furthermore, a span named "App #2" exists as a child of "App #1." Furthermore, a span named "App #3DB" exists as a child of "App #2."
[0039] The trace ID is set to the hexadecimal number "4bf92f." Here, for ease of viewing, the trace ID is set to six hexadecimal digits, but in reality, the trace ID is 32 hexadecimal digits.
[0040] The span ID of the span called "Client" is "00f067" in hexadecimal. Here, for ease of viewing, the span ID is shown as a six-digit hexadecimal number, but in reality, the span ID is 16 digits. Similarly, below, the span ID is shown as a six-digit hexadecimal number, but in reality, the span ID is 16 digits in hexadecimal.
[0041] The span ID for the span "App #1" is "10fb9f" in hexadecimal. The span ID for the span "App #2" is "30f71e" in hexadecimal. The span ID for the span "App #3" is "6e42f2" in hexadecimal.
[0042] Reference numeral 500 indicates information held by a span called "client." Since the span called "client" does not have a parent span, the value of the parent ID is "null." As mentioned above, in AMF or UDM, the SUPI (IMSI) is assigned to the span attribute, and therefore, although not shown in FIG. 5, each span has a SUPI (IMSI) that can identify the terminal (UE) as an attribute.
[0043] Reference numeral 600 denotes transaction information under the above assumptions. For the span named "Client", span information similar to that shown in reference numeral 500 is shown. For the span named "API#1", the parent span is the client, so the client's span ID "00f067" is displayed as the parent ID.
[0044] For the span named "API#2", the parent span is API#1, so the span ID of API#1, "10fb9f", is displayed as the parent ID. The trace ID of each span is "4bf92f". In the transaction information 600 of FIG. 5, the span API #3 is not displayed.
[0045] Figure 6 shows that when a failure occurs in one NF in the 5G core network, the API failure rate increases sharply. Processing in the control plane is achieved by repeatedly exchanging requests and responses between each NF. If an error occurs during the execution of a span, the attribute of that span is marked as "error." If an error occurs during the execution of the span in the NF that received the request, the NF that received the request returns an error response to the NF that made the request.
[0046] The denominator of the API failure rate is the total number of requests exchanged within a Kubernetes namespace. The numerator of the API failure rate is the number of spans with the attribute "Error".
[0047] As mentioned above, if an error occurs during the execution of a span in the NF (NF1) that received the request, the NF (NF1) that received the request will return an error response to the NF (NF2) that made the request. However, since the NF (NF2) that made the request actually received a request from a third NF (NF3), the NF that made the request (NF2) will return an error response to the third NF (NF3). In this way, because the error response is traced back to the source of the request, the API failure rate will rise sharply when a failure occurs in a certain NF, as shown in Figure 6.
[0048] FIG. 7 is a diagram illustrating a method for investigating the cause of a failure based on the failure occurrence state of each NF. Reference numeral 700 indicates that the process did not complete successfully in the network domain at the NF marked with an asterisk.
[0049] Reference numeral 800 shows a trace of a transaction. Each line corresponds to one span. The dashed lines show the name, start time, and end time of each span. "AMF," "AUSF," etc. indicate the NF where each span occurred. An exclamation point (!") indicates an error occurred while processing the span. The star in reference number 700 is based on the exclamation mark "!" in trace 800.
[0050] The fact that the opening of line 3 is slightly to the right of the opening of line 2 indicates that the span of line 3 is a child of the span of line 2. Similarly, the opening of line 5 is slightly to the right of the opening of line 2, so the span of line 5 is also a child of the span of line 2.
[0051] Looking at lines 2 to 4, we can see that there is no exclamation mark "!" in lines 3 and 4. This means that the request sent from AUSF to NRF was processed successfully. Looking at lines 5 to 8, we can see that a request is sent from AUSF to UDM, and then from UDM to UDR. The exclamation marks "!" on lines 5, 7, and 8 indicate that an error occurred in the processing of the lowest-level UDR, which is why the exclamation marks "!" were also added to lines 5 and 7. From the above, we can see that the exclamation marks "!" on lines 1 and 2 were originally caused by a processing error in the UDR on line 8.
[0052] The above relationship is illustrated by reference numeral 900 in Figure 7. Errors propagate in the direction indicated by the arrow. Since a trace ID is embedded in the log, by investigating the log based on the transaction trace ID and the information that the cause of the failure is a UDR, it is possible to find out the more specific cause of the failure.
[0053] FIG. 8 shows the procedure for investigating the log based on the trace ID of the transaction and the information that the cause of the failure is a UDR. We can see that the transaction trace ID is "4bf92f." Furthermore, we know from the processing in Figure 7 that the cause of the failure is UDR. Therefore, we investigate the UDR error details log and search for the trace ID "4bf92f." The trace ID in the UDR error details log is "4bf92f," which records "imsi-xxxxxx not found in database," revealing that the more specific cause of the problem was that a certain subscriber ID did not exist in the subscriber database.
[0054] In this embodiment, if there is at least one span in the trace that includes the SUPI (IMSI), a search can be performed in units of SUPI (terminal), such as SUPI ⇒ that span ⇒ trace representing a specific transaction.
[0055] In this way, in this embodiment, since the SUPI (IMSI) and span of a terminal where a failure in communication processing has occurred can be identified, it is also possible to grasp, for example, the SUPI (IMSI) of a terminal where authentication has failed. Therefore, a list of the number of terminals where authentication has failed and the SUPI (IMSI) of the terminal may be displayed on the monitor's dashboard.
[0056] As described above, in this embodiment, in AMF or UDM, SUPI (IMSI) is assigned to the span attribute, making it possible to monitor each terminal. In addition, in this embodiment, a trace context IE field is provided in NGAP and PFCP messages, and a trace ID is provided in the trace context IE, so that it is possible to determine which transaction a message belongs to, even when it is between the user plane and the control plane.
[0057] This will, for example, enable faster recovery from disruptions in wireless communication networks, thereby contributing to Goal 9 of the United Nations-led Sustainable Development Goals (SDGs), which is to "Build resilient infrastructure, promote inclusive and sustainable industrialization, and foster innovation." [Explanation of symbols]
[0058] 100 5G systems 120 5G Core Network 400 transaction data analysis platform 500 span information 600 Transaction Information 700 Network Domains 800 Trace 850 UDR error detail log 900 Relationship between each function
Claims
1. A 5G system in which a trace context information element field including a trace ID is provided in messages between a UE, (R)AN, or UPF and a control plane network function.
2. The 5G system according to claim 1, wherein a field of a trace context information element including a trace ID is provided in a message of the NGAP protocol and / or the PFCP protocol.
3. The 5G system according to claim 1, wherein SUPI is assigned to an attribute of a span in an AMF or UDM.
4. A monitoring method in a 5G system, comprising: A monitoring method comprising providing a field of a trace context information element including a trace ID in a message between a UE, an (R)AN, or a UPF and a network function of a control plane.
5. 5. The monitoring method according to claim 4, further comprising providing a field of a trace context information element including a trace ID in a message of the NGAP protocol and / or the PFCP protocol.
6. The monitoring method according to claim 4, further comprising the step of assigning a SUPI to an attribute of a span in AMF or UDM.
Citation Information
Cited By
Game machine
JP2025105824A
Game machine
JP2025105825A
Open telemetry enabled network device management
US12574310B2