Systems and methods for coordinated generation and collection of observability data in distributed systems

Through the observability management system, multiple observability data consumer requests are sorted out and configuration and data aggregation models are generated, which solves the problems of overhead of resource generation and collection of observability data in mobile networks and the duplicate processing of data, realizing resource efficiency and simplifying data generation.

CN120051979APending Publication Date: 2025-05-27TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280101192.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-11-04
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In mobile networks, generating and collecting observable data increases the overhead of computing, storing and network resources, and different network operations may require the same or similar observable data, resulting in competitive or conflicting configuration and repeated processing of data on observable data producers.

Method used

Through the observability management system, multiple requests submitted by observability data consumers are sorted out, observability configuration and data aggregation models are generated, observability data producers of distributed systems are configured, and data is aggregated, and finally reported to each consumer.

Benefits of technology

Simplify observable data generation, reduce overhead on observable data producer resources, avoid repeated processing of the same data, and improve resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120051979A_ABST
    Figure CN120051979A_ABST
Patent Text Reader

Abstract

A method performed by an observability management system for collecting observability data about a distributed system is disclosed. The method comprises the following steps: obtaining observability data requests submitted by a plurality of observability data consumers; generating an observability configuration and an observability data aggregation model based on sorting the observability data request; configuring one or more observability data producers of the distributed system according to the observability configuration, such that the one or more observability data producers generate observability data; collecting observability data generated by the one or more observability data producers; aggregating the collected observability data according to an observability data aggregation model to generate an aggregation result for each observability data request; and reporting, to each of the plurality of observability data consumers, an aggregated result of the observability data requests submitted by the observability data consumer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of distributed systems and, more particularly, to generating and collecting observability data in a distributed system. Background Art

[0002] In a mobile network, particularly a mobile network having cloud-native network functions deployed in a distributed cloud environment, being able to generate and collect observability data (e.g., performance measurements, logs, traces, etc.) is crucial for enabling automated and intelligent network operations such as fault analysis, workload orchestration, and service assurance. Generating and collecting observability data adds a non-negligible overhead to the mobile network (e.g., these operations consume computing resources, data storage resources, and / or network resources), and thus may have a negative impact on the performance of the mobile network. Network functions are designed to generate various types of observability data and can generate a large amount of observability data when deployed. Effectively configuring network functions to generate an appropriate range of observability data to support network operations while introducing minimal overhead can be challenging.

[0003] In some cases, for different purposes, different network operations may require the same or similar observability data generated by a set of the same network functions. For example, in the context of a mobile network, both a network optimization service and a network slice SLA (service level agreement) management service may require performance measurements from a set of network functions. The network optimization service and the network slice SLA management service may provide different services with different goals but require similar observability data from a set of the same network functions for their operations. Generally, the network optimization service aims to improve network efficiency, while the network slice SLA management service aims to improve the end-user experience. These two observability data consumers may have slightly different requirements in terms of the observability data they need (e.g., different measurement intervals and / or durations), which requires two separate jobs to be executed on the set of network functions to generate the observability data.

[0004] Different observability data consumers that need the same or similar observability data typically configure the observability data producers to generate the observability data separately and also to aggregate / process the collected observability data separately, especially when the observability data consumers are operated and / or managed by different entities. This can lead to competing (or even conflicting) observability data generation configurations being imposed on the observability data producers and / or duplicate processing of the collected observability data. Summary of the Invention

[0005] Disclosed is a method for collecting observability data about a distributed system, which is executed by an observability management system. The method includes: obtaining observability data requests submitted by multiple observability data consumers; generating an observability configuration and an observability data aggregation model based on collating the observability data requests; configuring one or more observability data producers of the distributed system according to the observability configuration to enable the one or more observability data producers to generate observability data; collecting the observability data generated by the one or more observability data producers; aggregating the collected observability data according to the observability data aggregation model to generate an aggregation result for each observability data request; and reporting the aggregation result for the observability data request submitted by each of the multiple observability data consumers to each of the multiple observability data consumers.

[0006] Disclosed is a non-transitory machine-readable medium, which includes computer program code that, when executed by one or more computing devices, causes the one or more computing devices to perform operations for collecting observability data about a distributed system. The operations include: obtaining observability data requests submitted by multiple observability data consumers; generating an observability configuration and an observability data aggregation model based on collating the observability data requests; configuring one or more observability data producers of the distributed system according to the observability configuration to enable the one or more observability data producers to generate observability data; collecting the observability data generated by the one or more observability data producers; aggregating the collected observability data according to the observability data aggregation model to generate an aggregation result for each observability data request; and reporting the aggregation result for the observability data request submitted by each of the multiple observability data consumers to each of the multiple observability data consumers.

[0007] A computing device for implementing an observability management system is disclosed. The computing device includes one or more processors and a non-transitory machine-readable storage medium storing computer instructions that, if executed by the one or more processors, cause the computing device to perform operations for collecting observability data about a distributed system. These operations include: obtaining observability data requests submitted by a plurality of observability data consumers; generating an observability configuration and an observability data aggregation model based on collating the observability data requests; configuring one or more observability data producers of the distributed system according to the observability configuration to cause the one or more observability data producers to generate observability data; collecting the observability data generated by the one or more observability data producers; aggregating the collected observability data according to the observability data aggregation model to generate an aggregation result for each observability data request; and reporting the aggregation result for the observability data request submitted by each of the plurality of observability data consumers to each of the plurality of observability data consumers. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The present invention can be best understood by referring to the following description and drawings that illustrate embodiments of the invention. In the drawings:

[0009] Figure 1 is an environmental diagram showing an environment capable of generating and collecting observability data in a coordinated manner according to some embodiments.

[0010] Figure 2 is a flowchart showing a method for generating and collecting observability data in a coordinated manner according to some embodiments.

[0011] Figure 3 is a component interaction diagram showing components for generating and collecting observability data in a coordinated manner according to some embodiments.

[0012] Figure 4 is a component interaction diagram showing components for updating an observability configuration and an observability data aggregation model according to some embodiments.

[0013] Figure 5 is a flowchart showing a method for collecting observability data about a distributed system according to some embodiments.

[0014] Figure 6A shows the connectivity between network devices (NDs) within an exemplary network and three exemplary implementations of the NDs according to some embodiments of the present invention.

[0015] Figure 6B shows an example manner for implementing a dedicated network device according to some embodiments of the present invention. DETAILED DESCRIPTION

[0016] The following description describes methods and apparatuses for generating and collecting observability data for a distributed system. In the following description, numerous specific details are set forth, such as logical implementation, operation codes (opcodes), means for specifying operands, resource partitioning / sharing / copying implementations, types and interrelationships of system components, and logical partitioning / integration choices, to provide a more thorough understanding of the present invention. However, those skilled in the art will recognize that the present invention may be practiced without these specific details. In other instances, control structures, gate-level circuits, and full software instruction sequences are not shown in detail so as not to obscure the present invention. Using the included description, one of ordinary skill in the art will be able to implement the appropriate functionality without undue experimentation.

[0017] References in the specification to "one embodiment", "an embodiment", "example embodiment", etc., indicate that the described embodiment may include a particular feature, structure, or characteristic, but each embodiment may not necessarily include that particular feature, structure, or characteristic. Moreover, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in connection with an embodiment, it should be considered within the knowledge of one of ordinary skill in the art to implement such feature, structure, or characteristic in connection with other embodiments (whether explicitly described or not).

[0018] In this document, text in parentheses and boxes with dashed boundaries (e.g., long-dashed dotted line, short-dashed dotted line, dotted line, and dots) may be used to illustrate optional operations for adding additional features to embodiments of the present invention. However, such annotation should not be construed to mean that in certain embodiments of the present invention, they are the only options or optional operations, and / or that boxes with solid boundaries are not optional.

[0019] In the following description and claims, the terms "coupled" and "connected" and their derivatives may be used. It should be understood that these terms are not intended as synonyms for each other. "Coupled" is used to indicate that two or more elements may or may not be in direct physical or electrical contact with each other, and may cooperate or interact with each other. "Connected" is used to indicate the establishment of communication between two or more elements that are coupled to each other.

[0020] An electronic device uses machine-readable media (also referred to as computer-readable media) to store and transmit code (which consists of software instructions and is sometimes referred to as computer program code or a computer program) and / or data (internally and / or in conjunction with other electronic devices on a network). Machine-readable media include, for example, machine-readable storage media (e.g., magnetic disks, optical disks, solid-state drives, read-only memory (ROM), flash memory devices, phase change memory) and machine-readable transmission media (also referred to as carriers) (e.g., electrical, optical, radio, acoustic, or other forms of propagated signals - such as carrier waves, infrared signals). Thus, an electronic device (e.g., a computer) includes hardware and software, such as a collection of one or more processors (e.g., where the processor is a microprocessor, controller, microcontroller, central processing unit, digital signal processor, application-specific integrated circuit, field-programmable gate array, other electronic circuits, a combination of one or more of the foregoing), which are coupled to one or more machine-readable storage media to store code for execution on the collection of processors and / or to store data. For example, an electronic device may include non-volatile memory that contains code, since non-volatile memory can retain the code / data even when the electronic device is turned off (when power is lost), and when the electronic device is turned on, that portion of the code that is typically to be executed by the processor of the electronic device is copied from the slower non-volatile memory to volatile memory (e.g., dynamic random access memory (DRAM), static random access memory (SRAM)) of the electronic device. A typical electronic device also includes a collection of one or more physical network interfaces (NIs) for establishing a network connection with other electronic devices (to send and / or receive code and / or data using propagated signals). For example, the collection of physical NIs (or a combination of the collection of physical NIs and the collection of processors that execute code) may perform any formatting, encoding, or conversion to allow the electronic device to send and receive data (via wired and / or wireless connections). In some embodiments, the physical NI may include radio circuitry capable of receiving data from other electronic devices via a wireless connection and / or sending data out to other devices via a wireless connection. The radio circuitry may include a transmitter, receiver, and / or transceiver suitable for radio frequency communication. The radio circuitry may convert digital data into a radio signal with appropriate parameters (e.g., frequency, timing, channel, bandwidth, etc.). The radio signal may then be transmitted via an antenna to an appropriate recipient. In some embodiments, the collection of physical NIs may include a network interface controller (NIC), also referred to as a network interface card, network adapter, or local area network (LAN) adapter. By inserting a cable into a physical port connected to the NIC, the NIC can facilitate connecting the electronic device to other electronic devices, thereby allowing them to communicate wired. One or more portions of embodiments of the present invention may be implemented using different combinations of software, firmware, and / or hardware.

[0021] A network device (ND) is an electronic device that communicates and interconnects other electronic devices on a network (e.g., other network devices, end-user devices). Some network devices are "multi-service network devices" that support multiple network functions (e.g., routing, bridging, switching, layer 2 aggregation, session border control, quality of service, and / or subscriber management) and / or support multi-application services (e.g., data, voice, and video).

[0022] As described above, different observability data consumers that require the same or similar observability data typically configure the observability data producer to separately generate the observability data and also separately aggregate / process the collected observability data, especially when the observability data consumers are operated and / or managed by different entities. This can lead to competing (or even conflicting) observability data generation configurations being imposed on the observability data producer and / or duplicate processing of the collected observability data.

[0023] As used herein, an observability data consumer is an entity that consumes observability data (e.g., edge service assurance, network and service orchestration, or analytics systems in a mobile network). As used herein, an observability data producer is an entity that produces observability data (e.g., a network function in a mobile network). As used herein, observability data can be any type of data regarding the operation, performance, or state of a component or system. Observability data can take the form of logs, metrics, and / or traces.

[0024] An embodiment provides an observability management system that accepts requests for observability data (also referred to as observability data requests) from observability data consumers, which may be associated with different entities / companies / enterprises, and generates an observability data configuration based on collating the observability data requests. The observability data requests may indicate observability data requested at a high level of specificity or a lower level of specificity. If indicated at a high level of specificity (coarse-grained), the observability management system may translate the requirements of the observability data request into lower-level requirements (more fine-grained). The observability data configuration may indicate the configuration to be imposed on the observability data producers to cause the observability data producers to generate the observability data required to satisfy the requirements of the observability request. The observability data configuration may be generated such that a minimal superset of the observability data that satisfies the requirements of the observability data request is generated. Then, the observability management system may configure the observability data producers according to the observability data configuration to cause the observability data producers to generate observability data. The observability management system may also generate an observability data aggregation model based on collating the observability data requests. The observability data aggregation model may indicate how to aggregate the observability data collected from the observability data producers that satisfies the requirements of the corresponding observability data request. The observability management system may collect the observability data generated by the observability data producers and aggregate the collected observability data according to the observability data aggregation model to generate an aggregation result for the corresponding observability data request. Then, the observability management system may report the aggregation result to the corresponding observability data consumer that submitted the observability data request.

[0025] An embodiment is a method for collecting observability data about a distributed system, performed by an observability management system implemented by one or more computing devices. The method includes: obtaining observability data requests submitted by a plurality of observability data consumers; generating an observability configuration and an observability data aggregation model based on collating the observability data requests; configuring one or more observability data producers of the distributed system according to the observability configuration to cause the one or more observability data producers to generate observability data; collecting the observability data generated by the one or more observability data producers; aggregating the collected observability data according to the observability data aggregation model to generate an aggregation result for each observability data request; and reporting the aggregation result of the observability data request submitted by each observability data consumer to each of the plurality of observability data consumers.

[0026] Through the embodiments, the observability management system can consolidate observability data requests for observability data related to the same target component submitted by multiple different observability data consumers and intelligently determine the exact set of required observability data. Additionally, the observability management system can aggregate the collected observability data on behalf of the observability data consumers to avoid duplicate processing of the same raw observability data, which would otherwise have to be separately performed by different observability data consumers. In this regard, the observability management system can act as a mediator between the observability data consumers and the observability data producers. The observability management system can dynamically configure the observability data producers (when the distributed system is in operation) according to the actual needs of the observability data consumers. Further, when generating an observability data configuration to be imposed on the observability data producers, the observability management system can consider the observability data generation capabilities of the observability data producers (e.g., as reported by the observability data producers).

[0027] In one example use case, the observability management system is used to generate and collect observability data in a mobile network having components deployed in a distributed edge / cloud environment. In such a case, the observability data consumers can be an edge service assurance system, a network and service orchestration system, or a network analytics system. Additionally, the observability data producers can be network functions deployed in the mobile network. For illustrative purposes, specific embodiments are described herein in the context of a mobile network. However, it should be understood that the embodiments can be used to generate and collect observability data regarding other types of distributed systems.

[0028] Embodiments can provide one or more advantages over conventional observability solutions. Embodiments collate observability data requests and configure observability data producers in a coordinated manner such that they generate only the observability data required to satisfy the observability data requests. This can simplify observability data generation and reduce potential overheads for observability data producers (e.g., if there are multiple observability data consumers requesting observability data about the same target component, the observability management system can configure the observability data producer to perform a single job that generates observability data about the target component that can be used to satisfy all requests (whereas conventional observability solutions may start a separate job for each request)). Coordinating the configuration of observability data producers and taking into account the observability data generation capabilities of observability data producers can help avoid imposing conflicting / competing observability configurations on observability data producers. The collation of observability data requests also allows common observability data aggregation operations to be performed in an intermediate layer (e.g., by the observability management system), which avoids duplicating the same processing of the same observability data. This helps save overall resource consumption in large-scale deployments. Aggregation can involve sampling the observability data and / or combining observability data to generate an aggregated result that meets the requirements of the observability data requests. While certain advantages have been mentioned above, embodiments can have other advantages that will be apparent to those of ordinary skill in the art in view of the present disclosure.

[0029] Figure 1 is an environmental diagram showing an environment capable of generating and collecting observability data in a coordinated manner according to some embodiments.

[0030] As shown, the environment includes observability data consumers 110A and 110B, observability data producers 160A and 160B, and an observability management system 120. For simplicity of illustration, the figure shows two observability data consumers 110 and two observability data producers 160. It should be understood that the environment can include a different number of observability data consumers 110 and / or observability data producers 160 than shown in the figure.

[0031] An observability data consumer 110 may submit a request for observability data regarding a distributed system to an observability management system 120. In the context where the distributed system is a mobile network, the observability data consumer 110 may be a component of a mobile edge service operating system that requires observability data about the mobile network to perform service orchestration, service performance analysis, and / or service assurance. The observability data consumer 110 may submit a request for observability data related to a specific component of the distributed system to the observability management system 120. For example, the observability data consumer 110 may submit a request for observability data related to a specific application or service deployed in the distributed system (e.g., a network function or an edge service deployed in a distributed mobile edge). The observability data may be any type of data regarding the operation, performance, or state of a component or system. The observability data may take the form of logs, metrics, and / or traces.

[0032] An observability data producer 160 may be configured (e.g., by the observability management system) to generate a specific type of observability data (e.g., logs, metrics, and / or traces) with a specific granularity (e.g., the specificity / range of the observability data) and frequency (e.g., every five minutes). In an embodiment, the observability data producer 160 may be dynamically configured (e.g., while the observability data producer 160 is in operation without having to perform a restart / reboot). The observability data producer 160 may send the observability data they generate to the observability management system 120. In an embodiment, the observability data producer 160 sends an observability capabilities report to the observability management system 120 indicating their capabilities regarding generating observability data (e.g., indicating what types of observability data they are capable of generating, the granularity of the observability data they are capable of generating, and / or the frequency at which they are capable of generating observability data).

[0033] An observability management system 120 may manage the generation and collection of observability data in a distributed system. The observability management system 120 may receive an observability data request from an observability data consumer 110, generate an observability configuration based on collating the observability data request, and configure one or more observability data producers 160 according to the observability configuration so that the one or more observability data producers 160 generate observability data. The observability configuration may indicate how to configure the observability data producers 160 such that they generate the observability data required to meet the requirements of the observability data request submitted by the observability data consumer 110. The observability management system 120 may also generate an observability data aggregation model based on collating the observability request. The observability data aggregation model may indicate how to aggregate the observability data collected from the observability data producers 160 to meet the requirements of the corresponding observability data request. Aggregation may involve sampling the collected observability data (e.g., extracting every fifth measurement) and / or combining the observability data (e.g., summing the measurement results). The observability management system 120 may collect the observability data generated by the observability data producers 160, aggregate the collected observability data if needed to generate an aggregated result that meets the requirements of the corresponding observability data request, and report the aggregated result to the corresponding observability data consumers 110 that submitted those observability data requests.

[0034] As shown in the figure, the observability management system 120 may include various components for managing the generation and collection of observability data, such as an observability data subscription component 130, an observability configuration component 140, and an observability data aggregation component 150. Example message flows and data flows involving these components are shown in the figure. Message flows are shown using solid arrows in the figure, while data flows are shown using thick arrows in the figure. The message flows and data flows are described in further detail below.

[0035] Referring to the message flow, an observability data consumer 110 may send an observability data request to the observability data subscription component 130. The observability data request may specify the observability data requested at a fine-grained level (e.g., as a specific metric on a specific target component) or at a coarse-grained level (e.g., as a general performance metric of a subsystem). The observability data subscription component 130 may merge / summarize the requirements of the observability data request into an observability data requirement (in a format that can be interpreted and used by the observability configuration component 140), and send the observability data requirement to the observability configuration component 140.

[0036] The observability configuration component 140 can generate an observability configuration and an observability data aggregation model based on collating observability data requirements. The observability configuration can indicate how to configure the observability data producers 160 such that they generate the observability data required to meet the requirements of the observability data requests submitted by the observability data consumers 110. The observability configuration component 140 can configure one or more observability data producers 160 according to the observability configuration so that the one or more observability data producers 160 generate observability data. In an embodiment, as shown in the figure with dashed lines, the observability configuration component 140 receives an observability capability report from the observability data producers 160. The observability capability report can indicate what types of observability data the observability data producers are capable of generating and related information. The observability configuration component 140 considers the capabilities of the observability data producers 160 when generating the observability configuration. The observability data aggregation model can indicate how to aggregate the observability data collected from the observability data producers 160 to meet the requirements of the corresponding observability data requests. The observability configuration component 140 can send the observability data aggregation model to the observability data aggregation component 150.

[0037] Referring now to the data flow, the observability data producers 160 can generate observability data (according to the configuration imposed by the observability configuration component 140) and send the generated observability data to the observability data aggregation component 150. The observability data aggregation component 150 can collect and aggregate the observability data generated by the observability data producers 160 according to the observability data aggregation model to generate an aggregation result that meets the requirements of the corresponding observability data requests. Then, the observability data aggregation component 150 can report the aggregation result to the corresponding observability data consumers 110.

[0038] Figure 2 is a flowchart showing a method for generating and collecting observability data in a coordinated manner according to some embodiments.

[0039] The operations in the flowchart will be described with reference to the exemplary embodiments of other figures. However, it should be understood that the operations of the flowchart can be performed by embodiments other than those discussed with reference to the other figures, and the embodiments discussed with reference to these other figures can perform operations different from those discussed with reference to the flowchart.

[0040] At operation 210, an observability data producer sends an observability capabilities report to an observability configuration component. The observability capabilities report can indicate what types of observability data the observability data producer is capable of generating (e.g., logs, metrics, and / or traces), the granularity of the observability data that the observability data producer is capable of generating (e.g., for each type, indicating the specific observability data items that can be generated (e.g., a list of performance-related metrics and / or a list of traces supported by microservices)), and / or the frequency at which the observability data producer is capable of generating observability data (e.g., indicating that the most frequent interval at which a certain observability data item can be generated is once per minute). In an embodiment, the observability capabilities report indicates whether the generation of certain observability data types and / or certain data items can be dynamically activated, and the possible mechanisms for activating it (e.g., commands that can be used to activate the generation of observability data). If the observability data producer only supports generating static observability data (e.g., the observability data producer cannot be dynamically configured to generate different types of observability data), the observability data producer can send an observability capabilities report to the observability management system indicating this situation. The observability management system can consider this information when generating the observability configuration and / or observability data aggregation model.

[0041] As an example, in the context of a fifth-generation (5G) mobile network, network functions (which can act as observability data producers) may be able to generate performance measurements defined by the Third Generation Partnership Project (3GPP) standards (e.g., 3GPP Technical Specification (TS) 28.552). When these network functions are deployed, they can send an observability capabilities report to the observability configuration component indicating which performance measurements they are capable of generating.

[0042] At operation 220, an observability data consumer sends an observability data request to an observability data subscription component. The observability data request may indicate the specific observability data being requested and related information. The observability data request may specify the observability data requested at a fine-grained level (e.g., specifying specific logs, metrics, or traces on a target component) or at a coarse-grained level (e.g., specifying general performance metrics of a subsystem). In an embodiment, the observability data request indicates the target component (which components to monitor / observe), the type of observability data to be collected (e.g., logs, metrics, and / or traces), the observability data collection frequency, the observability data collection duration, and / or the observability data reporting instruction (e.g., how the generated observability data should be reported to the observability data consumer (e.g., the observability data request may indicate the service endpoint to which the observability data should be sent, or indicate that the observability management system should provide an endpoint / location from which the observability data consumer can extract the observability data)). It should be noted that different observability data consumers may request the same or different sets of observability data on the same target component. As will be described in further detail herein, the observability management system may be responsible for consolidating and collating various observability data requests related to the same target component, and configuring the observability data producers such that they generate the observability data required to meet the requirements of different observability data requests.

[0043] As an example, the observability data request may include the following parameters: target, measurement, interval, duration, and report type. The target parameter may indicate which components to observe / monitor. The measurement parameter may indicate the measurement being requested (e.g., it may be a standardized or predefined measurement, such as 5G network-related measurements defined by 3GPP standards). The interval parameter may indicate the desired measurement frequency. The duration parameter may indicate the duration for which the measurement is generated. The report type parameter may indicate how / where the generated measurement results are reported (e.g., sending the measurement results as a stream or storing the measurement results in a file).

[0044] At operation 230, the observability configuration component generates an observability configuration and an observability data aggregation model based on collating the observability data requests (and possibly considering the capabilities of the observability data producers reported by the observability data producers). The observability configuration component may send the observability data aggregation model to the observability data aggregation component.

[0045] As an example, the observability configuration component can translate any coarser-grained requirements of the observability data request (specified at a higher level of generality) into more specific requirements. In the example of a mobile network use case, the observability data request can indicate the target component as a set of network functions, radio access network (RAN), or core network deployed at a specific site, or as two endpoints (e.g., if the observability data is related to latency). The observability configuration component can communicate with the OAM (Operation, Administration, and Maintenance) system of the mobile network to export more specific network elements accordingly. As another example, the observability data request can indicate the measurement results as packet delay, packet discard rate, and / or other types of performance metrics (e.g., KPIs (Key Performance Indicators) defined by 3GPP standards). The observability configuration component can map these measurement results to specific observability data items using the capabilities reported by the observability data producers. The observability configuration component can maintain information about this mapping. For example, the observability data request can use the measurement result names or KPIs defined by 3GPP standards to indicate the measurement results. The observability configuration component can map these measurement results to specific measurement results that can be generated by network functions (which can act as observability data producers) according to 3GPP standards. In some cases, the measurement results indicated by the observability data request may imply some aggregation on the original measurement results (e.g., KPIs). In such cases, the observability configuration component can generate an observability data aggregation model indicating how to aggregate the collected measurement results to generate the final measurement results.

[0046] In an embodiment, for each observability data item related to the same target component, the observability configuration component can maintain a list of observability data consumers that have requested it, along with the corresponding frequencies and / or durations. For each data item, the observability configuration component can determine the appropriate frequency for generating the observability data item and the duration for generating the observability data item. That is, the observability configuration component can collate the frequencies and durations indicated by different observability data requests to determine the appropriate frequency for generating the observability data item and the appropriate duration for generating the observability data item in order to meet the requirements of different observability data requests.

[0047] As an example, three different observability data consumers (U1, U2, and U3) can request the same packet counter representing the number of packets received by a specific user plane function (UPF), but U1 can request the packet counter at a frequency of once per minute and for a duration of 25 minutes, U2 can request the packet counter at a frequency of once every 30 seconds and for a duration of 5 minutes, and U3 can request the packet counter at a frequency of once every 15 seconds and for a duration of 10 minutes. In response, the observability configuration component can generate an observability configuration that indicates that the observability data producer will count the number of packets received by the UPF every 15 seconds for the first 10 minutes (from minute #0 to #10), and then adjust to count the number of packets received by the UPF every minute for the next 15 minutes (from minute #11 to #25).

[0048] The observability configuration component can also generate an observability data aggregation model that indicates how to sample the packet counter values to meet the requirements of the corresponding observability data requests. For example, the observability data aggregation model can indicate the following: (1) from minute #0 to #10, send the raw packet counter values to U3; (2) sample every other raw packet counter value and send it to U2 from minute #0 to #5; (3) sample every fourth raw packet counter value and send it to U1 from minute #0 to #10; and (4) from minute #11 to #25, send the raw packet counter values to U1 (since starting at minute #11, the observability data producer is configured to generate packet counter values every minute).

[0049] In the case where the observability data producer has already been configured to generate specific observability data items at a specific frequency and / or for a specific duration, the observability configuration component can consider this existing configuration to determine an appropriate configuration that will meet both the existing requirements and the requirements of the new observability data requests. If a new observability data consumer submits an observability data request that adds new requirements, the observability configuration component can dynamically update / adjust the observability configuration and / or the observability data aggregation model to accommodate the new requirements. For example, if a new observability data consumer U4 requests the packet counter at a frequency of once per minute and for a duration of 50 minutes, the observability configuration component can update the observability configuration to indicate that the observability data producer will count the number of packets received by the UPF every minute from minute #26 to #50, and update the observability data aggregation model to indicate that the raw packet counter values will be sent to U4 during that time period.

[0050] In a case where the observability configuration component determines that it is impossible to meet the requirements of the observability requests submitted by the observability data consumers (e.g., after considering the capabilities of the reported observability data producers), the observability configuration component may send a message to certain observability data consumers indicating which of the requested observability data can be provided and which cannot be provided.

[0051] At operation 240, the observability configuration component configures one or more observability data producers according to the observability configuration to cause the one or more observability data producers to generate observability data. To support efficient and low-overhead generation and transmission of observability data, the observability data producers can be designed in a way that enables dynamic configuration and / or activation of observability-related functions. Thus, when the observability configuration component generates a new observability configuration, it can reconfigure the relevant observability data producers accordingly.

[0052] As an example, the observability configuration component can configure network functions (which can be observability data producers) in a mobile network by creating measurement jobs in a performance management (PM) service (e.g., as specified by 3GPP standards) to cause the relevant network functions to generate measurement results required to meet the requirements of the observability data requests.

[0053] At operation 250, the observability data aggregation component collects the observability data generated by one or more observability data producers.

[0054] At operation 260, the observability data aggregation component aggregates the collected observability data according to the observability data aggregation model to generate an aggregation result for the corresponding observability data request. As described above, the observability configuration component generates the observability configuration and the observability data aggregation model based on collating the observability data requests submitted by multiple different observability data consumers, and configures the observability data producers according to the observability configuration. The observability data aggregation component can perform operations that can be regarded as reverse / inferential operations. For example, in a case where the frequency of the observability data requested by the observability data consumers is lower than the frequency of the observability data generated by the observability data producers, the observability data aggregation component can sample the observability data generated by the observability data producers to meet the requirements of the observability data requests. As another example, in a case where multiple observability data items are required to meet the (coarse-grained) requirements of the observability data request, the observability data aggregation component can aggregate the relevant observability data items by combining / processing the relevant observability data items and then reporting the aggregation result to the observability data consumers.

[0055] As an example, if the observability data being requested is the packet delay of a data path within a mobile network, the delays between each hop of the data path can be collected and combined (e.g., summed) to generate the total packet delay. For each relevant observability data request submitted by an observability data consumer, the observability data aggregation component can aggregate the collected observability data according to the observability data aggregation model to meet the requirements of the corresponding observability data request. The advantages of this method are that the aggregation operations that are usually required for the same observability data can be performed once (avoiding duplicate processing), and the aggregation results can be shared and reported to each relevant observability data consumer.

[0056] As an example, the observability data aggregation component can act as a Management Service (MnS) consumer, which collects the measurement results generated by network functions and aggregates the collected measurement results according to the observability data aggregation model, and then reports the aggregation results to the corresponding observability data consumers.

[0057] At operation 270, the observability data aggregation component reports the aggregation results for the observability data requests submitted by each observability data consumer to that observability data consumer. The observability data aggregation component can send the aggregation results directly to the observability data consumers, and / or store the aggregation results at a storage location accessible by the observability data consumers (e.g., as a file).

[0058] Figure 3 is a component interaction diagram showing how to generate and collect observability data in a coordinated manner according to some embodiments.

[0059] As shown in the figure, the network function 320 (which is a gNB in this example) can report the measurement results it supports (e.g., report a list of measurement results that the gNB is capable of generating) to the observability configuration component 140.

[0060] An observability data consumer 110 (which can be a network automation system, for example) can send an observability data request to the observability data subscription component 130 ("ODS") to request the downlink delay of 5G connections experienced by a specific group of UEs (user equipment) in a given area (e.g., to determine whether a 5G cell needs to be enabled / disabled, whether the resource allocation in the gNB should be adjusted, and / or whether the throughput configuration of the relevant UPF should be adjusted - for quality assurance).

[0061] The observability data request can include the following information:

[0062] {

[0063] Target: 5G connections with network slice ns# in area a#

[0064] Measurement: Latency

[0065] Interval: 5 minutes

[0066] Duration: 1 hour

[0067] Report type: Streaming

[0068] }

[0069] In the above information, "ns#" refers to the single network slice selection assistance information (S-NSSAI) associated with a specific network slice that serves the UE, and "a#" is the identifier of a specific tracking area (TA) known to the mobile network management system.

[0070] The observability data subscription component 130 can notify the observability configuration component 140 ("OC") of this request. The observability configuration component 140 can consult the mobile network OAM 310 to determine the network slice instances associated with a specified network slice ("ns1") that serves a specified tracking area ("a1"), and determine the network functions associated with those network slice instances (e.g., RAN network functions (e.g., gNB) and 5G core (5GC) network functions (e.g., AMF, SMF, UPF)).

[0071] The observability configuration component 140 can determine the measurement results required to meet the requirements of the observability data request and the corresponding observability data producers that can generate those measurement results. In this example, the observability configuration component 140 determines that the required measurement results are (1) M1: the downlink packet latency between the gNB 320 and the UE (e.g., DR.DelayDlNgranUeDist.SNSSAI as defined in the 3GPP standard); and (2) M2: the downlink packet latency between the gNB 320 and the UPF (e.g., GTP.DelayDlPsaUpfNgranMean.SNSSAI as defined in the 3GPP standard). Additionally, in this example, the observability configuration component 140 determines that the gNB 320 (e.g., which has the ID of "gNB-1" and is part of the network slice "ns1") can generate both of these measurement results (the gNB 320 is the source). The gNB 320 may have previously reported that it is capable of generating such latency measurements.

[0072] The observability configuration component 140 can determine whether there are any existing measurement jobs being performed by the gNB 320 (e.g., by checking a configuration list maintained by the observability management system 120 or performing the operation listMeasurementJobs (as defined in the 3GPP standard) to query for existing measurement jobs being performed by the gNB 320). In this example, it is assumed that there are no existing measurement jobs being performed by the gNB 320.

[0073] The observability configuration component 140 can generate a measurement configuration and send it to the mobile network OAM 310 to cause the gNB 320 to generate latency measurements. For example, the measurement configuration can indicate that the target network function is the gNB 320, the measurements to be generated are the downlink packet latency between the gNB 320 and the UE and the downlink packet latency between the gNB 320 and the UPF, the measurement interval is five minutes, and the measurement duration is one hour. The measurement configuration can be represented as shown in Table I below (two "createMeasurementJob" for two latency measurements M1 and M2 respectively). The mobile network OAM 310 can activate the generation of measurement data in the gNB 320 according to this measurement configuration. Various examples provided herein (e.g., Table I and Table II) use terms adopted from the 3GPP standard. However, it should be understood that the embodiments are not limited to being used with the 3GPP standard.

[0074]

[0075] Table I

[0076] The observability configuration component 140 can generate an observability data aggregation model and send it to the observability data aggregation component 150 ("ODA"). In this example, the observability data aggregation model indicates that the total latency is the sum of the downlink packet latency between the gNB 320 and the UE and the downlink packet latency between the gNB 320 and the UPF (i.e., M1 + M2).

[0077] The gNB 320 can generate these two latency measurements (i.e., M1 and M2) according to its configuration and send them to the observability data aggregation component 150. The observability data aggregation component 150 can collect the latency measurements from the gNB 320 (or an OAM component capable of providing these latency measurements generated by the gNB 320). The observability data aggregation component 150 can aggregate these two latency measurements by calculating the sum of the two latency measurements (i.e., calculating M1 + M2 as indicated by the observability data aggregation model). The observability data aggregation component 150 can send the aggregation result (which is the total downlink latency) to the observability data consumer 110 (and do so every five minutes).

[0078] Figure 4 shows a component interaction diagram for updating an observability configuration and an observability data aggregation model according to some embodiments. The example shown in this figure is Figure 3 a continuation of the example shown and described above.

[0079] In the example shown in this figure, another observability data consumer 110B (which can be, for example, a network performance KPI reporting system) can send an observability data request to the observability data subscription component 130 to request downlink delay measurements between a UE and a gNB 320 (with the ID of "gNB-1") regarding the same network slice ("ns1") as the network slice in the Figure 3 example shown. The requested measurement interval is one minute, and the requested measurement duration is one hour. The observability data consumer 110A shown in this figure can correspond to the observability data consumer 110 mentioned above regarding the Figure 3 example shown.

[0080] The observability data request may include the following information:

[0081] {

[0082] Target: Network slice "ns1" served by gNB "gNB-1"

[0083] Measurement: Downlink delay between UE and gNB

[0084] Interval: 1 minute

[0085] Duration: 1 hour

[0086] Report type: Streaming

[0087] }

[0088] The observability data subscription component 130 can notify the observability configuration component 140 of this request. The observability configuration component 140 can determine the measurement results required to meet the requirements of the observability data request and the corresponding observability data producers capable of generating these measurement results. In this example, the observability configuration component 140 determines that the required measurement result is the downlink packet delay between the gNB 320 and the UE, and the gNB 320 can generate this measurement result (the gNB 320 is the source).

[0089] The observability configuration component 140 can determine whether there are any existing measurement jobs being performed by the gNB 320. In this example, the observability configuration component 140 determines that a measurement job for measuring the downlink packet delay between the gNB 320 and the UE is already being performed by the gNB 320, but with an interval of 5 minutes.

[0090] The observability configuration component 140 can cause the existing measurement job that the gNB 320 is exiting to be terminated (e.g., using "stopMeasurementJob" as defined in the 3GPP standard), and generate a new measurement configuration and send it to the mobile network OAM 310 to cause the gNB 320 to generate appropriate delay measurements (e.g., using "createMeasurementJob" as defined in the 3GPP standard). For example, this measurement configuration can indicate that: the target network function is the gNB 320, the measurement result to be generated is the downlink packet delay between the gNB 320 and the UE, the measurement interval is one minute, and the measurement duration is one hour. This measurement configuration can be represented as shown in Table II below (the "createMeasurementJob" representing the delay measurement M1). The mobile network OAM 310 can activate the measurement data generation in the gNB 320 according to this measurement configuration.

[0091]

[0092] Table II

[0093] It should be noted that if there are multiple requests for measurements related to the same network function, the interval can be determined based on collating these multiple measurement requests.

[0094] The observability configuration component 140 can update the existing observability data aggregation model and send it to the observability data aggregation component 150. In this example, the updated observability data aggregation model indicates that: (1) for the observability data consumer 110A, sample every fifth measurement result of the downlink packet delay between the gNB 320 and the UE, and sum it with the downlink packet delay between the gNB 320 and the UPF (e.g., sum "DR.DelayDlNgranUeDist.SNSSAI" and "GTP.DelayDlPsaUpfNgranMean.SNSSAI"); and (2) for the observability data consumer 110B, provide the downlink packet delay between the gNB 320 and the UE (e.g., stream the raw data of "DR.DelayDlNgranUeDist.SNSSAI").

[0095] The gNB 320 may generate latency measurements (updated versions of M1 and M2) according to its updated configuration and send them to the observability data aggregation component 150. The observability data aggregation component 150 may collect latency measurements from the gNB 320 (or an OAM component capable of providing the latency measurements generated by the gNB 320) (e.g., collect measurement data of the updated M1 (generated every 1 minute) and M2 (generated every 5 minutes) from the gNB 320). The observability data aggregation component 150 may aggregate the measurement data of the observability data consumers 110A and 110B according to the observability data aggregation model. For example, for the observability data consumer 110A, the observability data aggregation component 150 may sample every fifth measurement result of the downlink latency between the gNB 320 and the UE and sum it with the downlink latency between the gNB 320 and the UPF (e.g., sample every 5 measurement results of "DR.DelayDlNgranUeDist.SNSSAI" and sum it with "GTP.DelayDlPsaUpfNgranMean.SNSSAI") to aggregate these two latency measurements. For the observability data consumer 110B, the observability data aggregation component 150 may only extract the downlink latency between the gNB 320 and the UE without having to sample it or combine it with other measurements (e.g., perform each measurement of "DR.DelayDlNgranUeDist.SNSSAI"). The observability data aggregation component 150 may send the aggregation result to the corresponding observability data consumer 110 (e.g., do so every five minutes for the observability data consumer 110A and every minute for the observability data consumer 110B).

[0096] Figure 5 It is a flowchart of a method for collecting observability data about a distributed system according to some embodiments. In an embodiment, the method is performed by an observability management system implemented by one or more computing devices. In an embodiment, the distributed system is a mobile network.

[0097] At operation 510, the observability management system obtains observability data requests submitted by multiple observability data consumers. In an embodiment, the observability data request indicates one or more of the following: one or more targets to be observed / monitored, one or more types of observability data to be collected, the observability data collection frequency, the observability data collection duration, and the observability data reporting instruction.

[0098] At operation 520, the observability management system generates an observability configuration and an observability data aggregation model based on collating observability data requests. In an embodiment, the collation includes: converting the observability data requirements of the observability data requests submitted by multiple observability data consumers into lower-level observability data requirements, and determining whether there are overlaps in the lower-level observability data requirements.

[0099] In an embodiment, the observability management system receives an observability capability report from at least one of one or more observability data producers, where the observability capability report indicates which types of observability data the at least one of the one or more observability data producers is capable of generating. The observability management system may consider the information included in the observability capability report when generating the observability configuration.

[0100] At operation 530, the observability management system configures one or more observability data producers of the distributed system according to the observability configuration to cause the one or more observability data producers to generate observability data. In an embodiment, the configuration includes configuring a particular observability data producer among the one or more observability data producers to perform a single job that generates observability data capable of meeting the requirements of more than one of the observability data requests. In an embodiment, a particular observability data producer among the observability data producers is configured to generate observability data with a frequency and duration that are determined based on collating the frequencies and durations indicated by more than one of the observability data requests.

[0101] At operation 540, the observability management system collects the observability data generated by one or more observability data producers. In an embodiment, the collected observability data includes one or more of the following: logs, metrics, and traces.

[0102] At operation 550, the observability management system aggregates the collected observability data according to the observability data aggregation model to generate an aggregation result for each observability data request. In an embodiment, the aggregation includes one or more of the following: sampling observability data items (e.g., specific measurement results) from the collected observability data, and combining observability data items from the collected observability data.

[0103] At operation 560, the observability management system reports the aggregation result for the observability data request submitted by each observability data consumer to that observability data consumer.

[0104] Embodiments can be used in the context of a mobile network to meet the requirements for coordinated management of data production defined by 3GPP standards. For example, components of the Open Network Automation Platform (ONAP), a popular open-source project for network management systems, such as DMaaP (Data Movement as a Platform) and / or DCAE (Data Collection, Analysis, and Events), can be extended by implementing the techniques described herein. More specifically, an observability data consumer (e.g., a VES consumer and / or a file consumer) can submit an observability data request for network measurement data from DMaaP to an observability data subscription component. The observability configuration component can merge and collate different observability data requests, generate appropriate measurement jobs (e.g., a "PerfMetricJob" with appropriate parameters such as an interval), and configure the relevant network functions and 3GPP PM mappers accordingly. The observability configuration component can also generate an observability data aggregation model and send it to the observability data aggregation component. The observability data aggregation component can collect the measurement data generated by the network functions, aggregate the collected measurement data, and send the aggregation result to the observability data consumers (e.g., a VES consumer and / or a file consumer). Embodiments can be implemented as a set of microservices deployed / distributed in a cloud environment.

[0105] Figure 6A Shows the connectivity between network devices (NDs) within an exemplary network and three exemplary implementations of NDs according to some embodiments of the present invention. Figure 6A Shows NDs 600A to 600H and their connectivity is shown by lines between 600A to 600B, 600B to 600C, 600C to 600D, 600D to 600E, 600E to 600F, 600F to 600G, and between 600A to 600G and between each of 600H and 600A, 600C, 600D, and 600G. These NDs are physical devices, and the connectivity between these NDs can be wireless or wired (commonly referred to as a link). Additional lines extending from NDs 600A, 600E, and 600F show that these NDs act as entry and exit points of the network (and thus these NDs are sometimes referred to as edge NDs; while other NDs can be referred to as core NDs).

[0106] Figure 6A Two exemplary ND implementations are: 1) a dedicated network device 602 using a custom application-specific integrated circuit (ASIC) and a dedicated operating system (OS); and 2) a general-purpose network device 604 using a commercial off-the-shelf processor (COTS) and a standard OS.

[0107] The dedicated network device 602 includes network hardware 610, which includes a collection 612 of one or more processors, forwarding resources 614 (which typically include one or more ASICs and / or network processors), and physical network interfaces (NIs) 616 (through which network connections are made, such as the network connection shown by the connectivity between ND 600A and ND 600H), and a non-transitory machine-readable storage medium 618 in which network software 620 is stored. During operation, the network software 620 can be executed by the network hardware 610 to instantiate a collection 622 of one or more network software instances. Each network software instance 622 and the portion of the network hardware 610 that executes that network software instance (if it is hardware dedicated to that network software instance and / or a time slice of hardware time-shared by that network software instance with other network software instances 622) form separate virtual network elements 630A through 630R. Each virtual network element (VNE) 630A through 630R includes control communication and configuration modules 632A through 632R (sometimes referred to as local control modules or control communication modules) and forwarding tables 634A through 634R, such that a given virtual network element (e.g., 630A) includes a control communication and configuration module (e.g., 632A), a collection of one or more forwarding tables (e.g., 634A), and the portion of the network hardware 610 that executes the virtual network element (e.g., 630A).

[0108] The dedicated network device 602 is generally considered physically and / or logically to include: 1) an ND control plane 624 (sometimes referred to as the control plane), including processors 612 that execute the control communication and configuration modules 632A through 632R; and 2) an ND forwarding plane 626 (sometimes referred to as the forwarding plane, data plane, or media plane), including the forwarding resources 614 that utilize the forwarding tables 634A through 634R and the physical NIs 616. As an example where the ND is a router (or implements routing functionality), the ND control plane 624 (the processors 612 that execute the control communication and configuration modules 632A through 632R) is generally responsible for participating in controlling how to route (e.g., the next hop of data and the output physical NI for that data) data (e.g., packets) and for storing that routing information in the forwarding tables 634A through 634R, and the ND forwarding plane 626 is responsible for receiving the data on the physical NIs 616 and forwarding the data out the appropriate physical NI in the physical NIs 616 based on the forwarding tables 634A through 634R.

[0109] In an embodiment, the software 620 includes code such as an observability management component 623 that, when executed by the network hardware 610, causes the dedicated network device 602 to perform the operations of one or more embodiments disclosed herein (e.g., for generating and collecting observability data related to a distributed system).

[0110] Figure 6B Shows an example way for implementing a dedicated network device 602 according to some embodiments of the present invention. Figure 6B Shows a dedicated network device including a card 638 (usually hot-pluggable). Although in some embodiments, the card 638 has two types (one or more (sometimes referred to as line cards) operating as an ND forwarding plane 626, and one or more (sometimes referred to as control cards) for implementing an ND control plane 624), alternative embodiments may combine functions onto a single card and / or include additional card types (e.g., an additional type of card is referred to as a service card, a resource card, or a multi-application card). The service card may provide special processing (e.g., layer 4 to layer 7 services (e.g., firewall, Internet Protocol Security (IPsec), Secure Sockets Layer (SSL) / Transport Layer Security (TLS), Intrusion Detection System (IDS), peer-to-peer (P2P), Voice over IP (VoIP) session border controller, mobile radio gateway (Gateway General Packet Radio Service (GPRS) Support Node (GGSN), Evolved Packet Core (EPC) gateway))). As an example, the service card may be used to terminate an IPsec tunnel and perform the accompanying authentication and encryption algorithms. These cards are coupled together by one or more interconnection mechanisms shown as a backplane 636 (e.g., a first full mesh coupling the line cards and a second full mesh coupling all the cards).

[0111] Return to Figure 6A, the general network device 604 includes hardware 640, which includes a set 642 of one or more processors (usually COTS processors) and a physical NI 646, as well as a non-transitory machine-readable storage medium 648 in which software 650 is stored. During operation, the processor 642 executes the software 650 to instantiate one or more sets of one or more applications 664A through 664R. Although one embodiment does not implement virtualization, alternative embodiments may use different forms of virtualization. For example, in one such alternative embodiment, the virtualization layer 654 represents the kernel of the operating system (or a simplified version executed on top of the underlying operating system), which allows the establishment of multiple instances, referred to as software containers 662A through 662R, each of which can be used to execute one (or more) applications from the set of applications 664A through 664R; where the multiple software containers (also referred to as virtualization engines, virtual private servers, or jails) are user spaces (usually virtual memory spaces), which are separated from each other and from the kernel space in which the operating system runs; and where, unless explicitly permitted, the set of applications running in a given user space cannot access the memory of other processes. In another such alternative embodiment, the virtualization layer 654 represents a hypervisor (sometimes referred to as a virtual machine monitor (VMM)) or a hypervisor executed on top of the host operating system, and each of the sets of applications 664A through 664R runs on top of a guest operating system within instances 662A through 662R, which are referred to as virtual machines running on top of the hypervisor (in some cases, they can be considered a form of tightly isolated software containers), and the guest operating system and applications may not be aware that they are running on a virtual machine rather than on a "bare metal" host electronic device, or through paravirtualization, the operating system and / or applications can be aware of the existence of virtualization for optimization purposes. In other alternative embodiments, one, some, or all of the applications are implemented as unikernels, which can be generated by directly leveraging the application to compile only a limited set of libraries that provide the specific OS services required by the application (e.g., from a library operating system (LibOS), including drivers / libraries for OS services). Since unikernels can be implemented to run directly on the hardware 640, directly on the hypervisor (in which case the unikernel is sometimes described as running within a LibOS virtual machine), or within a software container, embodiments can be implemented entirely by unikernels as follows: unikernels running directly on the hypervisor represented by the virtualization layer 654, unikernels running within the software containers represented by instances 662A through 662R, or a combination of unikernels and the above techniques (e.g., unikernels and virtual machines both running directly on top of the hypervisor, unikernels and sets of applications running within different software containers).

[0112] The instantiation and virtualization (if implemented) of one or more collections of one or more of applications 664A through 664R are collectively referred to as software instance 652. Each collection of applications 664A through 664R, the corresponding virtualization constructs (e.g., instances 662A through 662R) (if implemented), and the portion of hardware 640 that executes them (which is dedicated hardware for that execution and / or a time slice of hardware shared temporarily) form separate virtual network elements 660A through 660R.

[0113] Virtual network elements 660A through 660R perform functions similar to those of virtual network elements 630A through 630R - e.g., similar to control communication and configuration module 632A and forwarding table 634A (this virtualization of hardware 640 is sometimes referred to as network function virtualization (NFV)). Thus, NFV can be used to unify many network device types into industry-standard high-volume server hardware, physical switches, and physical storage devices, which can be located in data centers, ND, and customer premise equipment (CPE). While embodiments of the present invention are shown with each of instances 662A through 662R corresponding to one of VNE 660A through 660R, alternative embodiments can implement this correspondence at a finer granularity level (e.g., line card virtual machines virtualize line cards, control card virtual machines virtualize control cards, etc.); it should be understood that the techniques described herein with reference to the correspondence of instances 662A through 662R with VNE also apply to embodiments using such finer granularity levels and / or single kernels.

[0114] In some embodiments, virtualization layer 654 includes a virtual switch that provides forwarding services similar to those of a physical Ethernet switch. Specifically, the virtual switch forwards traffic between instances 662A through 662R and physical NI 646 and optionally between instances 662A through 662R; in addition, the virtual switch can enforce network isolation between VNE 660A through 660R, which, according to a policy, are not allowed to communicate with each other (e.g., by implementing virtual local area network (VLAN)).

[0115] In an embodiment, software 650 includes code such as observability management component 653 that, when executed by hardware 640, causes general network device 604 to perform the operations of one or more embodiments disclosed herein (e.g., for generating and collecting observability data related to a distributed system).

[0116] Figure 6AA third exemplary ND implementation is the hybrid network device 606, which includes a custom ASIC / proprietary OS and a COTS processor / standard OS in a single ND or on a single card within an ND. In some embodiments of such a hybrid network device, a platform VM (i.e., a VM that implements the functionality of the dedicated network device 602) may provide paravirtualization to the network hardware present in the hybrid network device 606.

[0117] Some portions of the foregoing detailed description have been presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithms and representations are the means by which those skilled in the data processing arts most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived as a self-consistent sequence of operations leading to a desired result. The operations are those requiring physical manipulation of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to represent these signals by quantities such as bits, values, elements, symbols, characters, terms, numbers, etc.

[0118] However, it should be borne in mind that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless specifically stated otherwise as apparent from the above discussion, it should be appreciated that throughout the specification, discussions using terms such as "processing" or "computing" or "calculating" or "determining" or "displaying" etc. refer to the actions and processes of a computer system or similar electronic computing device that manipulates and transforms data represented as physical (electronic) quantities within the registers and memories of the computer system into other data similarly represented as physical quantities in the memories or registers or other such information storage, transmission, or display devices of the computer system.

[0119] The algorithms and displays presented herein are not inherently related to any particular computer or other apparatus. Various general purpose systems may be used with programs in accordance with the teachings herein, or it may prove convenient to construct more specialized apparatus to perform the required method operations. The structure of various such systems will be apparent from the description above. In addition, no specific programming language has been referenced to describe embodiments. It should be understood that a variety of programming languages may be used to implement the teachings of the embodiments as described herein.

[0120] An embodiment can be an article in which instructions (e.g., computer code) for programming one or more data processing components (collectively referred to herein as "processors") to perform the above operations are stored on a non-transitory machine-readable storage medium (e.g., a microelectronic memory). In other embodiments, some of these operations may be performed by specific hardware components that include hardwired logic (e.g., dedicated digital filter blocks and state machines). Alternatively, these operations may be performed by any combination of programmed data processing components and fixed hardwired circuit components.

[0121] Throughout the specification, embodiments have been presented via flowcharts. It should be understood that the transactions and the order of transactions described in these flowcharts are for illustrative purposes only and are not intended to be limiting. Those of ordinary skill in the art will recognize that changes may be made to the flowcharts.

[0122] In the foregoing specification, embodiments have been described with reference to their specific exemplary embodiments. Obviously, various modifications can be made thereto without departing from the broader spirit and scope of the disclosure provided herein. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive.

Claims

1. A method for collecting observability data about a distributed system, performed by an observability management system implemented by one or more computing devices, the method comprises: obtaining (510) observability data requests submitted by a plurality of observability data consumers; generating (520) an observability configuration and an observability data aggregation model based on collating the observability data requests; configuring (530) one or more observability data producers of the distributed system according to the observability configuration, so that the one or more observability data producers generate observability data; collecting (540) the observability data generated by the one or more observability data producers; aggregating (550) the collected observability data according to the observability data aggregation model to generate an aggregation result for each observability data request; and reporting (560) the aggregation result for the observability data request submitted by the observability data consumer to each of the plurality of observability data consumers.

2. The method according to claim 1, wherein, the collected observability data includes one or more of the following items: logs, metrics, and traces.

3. The method according to claim 1, further comprises: receiving an observability capability report from at least one of the one or more observability data producers, wherein the observability capability report indicates which types of observability data the at least one of the one or more observability data producers is capable of generating.

4. The method according to claim 1, wherein, the observability data request indicates one or more of the following items: one or more targets to be observed, one or more types of observability data to be collected, observability data collection frequency, observability data collection duration, and observability data reporting instructions.

5. The method according to claim 1, wherein, the collating includes: converting the observability data requirements of the observability data requests submitted by the plurality of observability data consumers into lower-level observability data requirements, and determining whether there are overlaps in the lower-level observability data requirements.

6. The method according to claim 1, wherein, the configuring includes configuring a specific observability data producer among the one or more observability data producers to execute a single job, and the single job generates observability data that can be used to meet the requirements of more than one observability data request in the observability data requests.

7. The method according to claim 6, wherein, the specific observability data producer among the observability data producers is configured to generate observability data at the following frequency and duration: the frequency and duration are determined based on collating the frequency and duration indicated by the more than one observability data request in the observability data requests.

8. The method according to claim 1, wherein, The aggregation includes one or more of the following: sampling observability data items from the collected observability data, and combining observability data items from the collected observability data.

9. The method according to claim 1, wherein, the distributed system is a mobile network.

10. A non-transitory machine-readable medium including computer program code that, when executed by one or more computing devices, causes the one or more computing devices to perform operations for collecting observability data about a distributed system, the operations including: obtaining (510) observability data requests submitted by a plurality of observability data consumers; generating (520) an observability configuration and an observability data aggregation model based on collating the observability data requests; configuring (530) one or more observability data producers of the distributed system according to the observability configuration so that the one or more observability data producers generate observability data; collecting (540) the observability data generated by the one or more observability data producers; aggregating (550) the collected observability data according to the observability data aggregation model to generate an aggregation result for each observability data request; and reporting (560) the aggregation result for the observability data request submitted by the observability data consumer to each of the plurality of observability data consumers.

11. The non-transitory machine-readable medium according to claim 10, wherein, the collected observability data includes one or more of the following: logs, metrics, and traces.

12. The non-transitory machine-readable medium according to claim 10, wherein, the operations further include: receiving an observability capability report from at least one of the one or more observability data producers, wherein the observability capability report indicates which types of observability data the at least one of the one or more observability data producers is capable of generating.

13. The non-transitory machine-readable medium according to claim 10, wherein, the observability data request indicates one or more of the following: one or more targets to be observed, one or more types of observability data to be collected, observability data collection frequency, observability data collection duration, and an observability data reporting instruction.

14. The non-transitory machine-readable medium according to claim 10, wherein, the collating includes: converting the observability data requirements of the observability data requests submitted by the plurality of observability data consumers into lower-level observability data requirements, and determining whether there are overlaps in the lower-level observability data requirements.

15. The non-transitory machine-readable medium according to claim 10, wherein, The configuration includes: configuring a particular observability data producer among the one or more observability data producers to perform a single job that generates observability data capable of meeting the requirements of more than one of the observability data requests.

16. The non-transitory machine-readable medium according to claim 15, wherein, the particular observability data producer among the observability data producers is configured to generate observability data at a frequency and duration that are determined based on collating the frequencies and durations indicated by the more than one of the observability data requests in the observability data request.

17. The non-transitory machine-readable medium according to claim 10, wherein, the aggregation includes one or more of the following: sampling observability data items from the collected observability data and combining observability data items from the collected observability data.

18. The non-transitory machine-readable medium according to claim 10, wherein, the distributed system is a mobile network.

19. A computing device for implementing an observability management system, the computing device comprises: one or more processors; and a non-transitory machine-readable storage medium storing computer instructions that, if executed by the one or more processors, cause the observability management system to perform the method steps according to any one of claims 1-9.