Operating twinning causal inference system (CAROT) for development and operation of 5G CNF
By operating a twin causal reasoning system (CAROT) to analyze observation data of 5G cloud native network functions (CNF) in a digital twin environment, the problem of difficult to optimize CNF configuration and design under non-ideal cloud infrastructure conditions in the existing technology is solved, and more efficient performance and robust optimization is achieved.
Patent Information
- Application Number
- CN202280101207.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-20
- Publication Date
- 2025-05-30
AI Technical Summary
The existing technology is difficult to effectively identify and solve the configuration and design problems of 5G cloud native network functions (CNF) in the software development cycle under non-ideal cloud infrastructure conditions, resulting in difficult performance and robustness optimization.
The operation twin causal reasoning system (CAROT) is used to optimize the configuration and design of CNF by creating a digital twin environment, copying the operating conditions in the production environment, generating observation data, and applying causal reasoning functions.
It realizes identification and resolution of CNF configuration and design issues early in the software development cycle, improves performance and robustness, reduces cost and time, and enhances insights into the configuration settings and design of 5G CNF.
Smart Images

Figure CN120077366A_ABST
Abstract
Description
Technical Field
[0001] Examples and non - limiting example embodiments generally relate to chaos engineering and, more particularly, to an operational twin causal reasoning system (CAROT) for the development and operation of 5G CNFs. Background Art
[0002] It is well - known to develop fault - tolerant systems within the software development life cycle. Summary of the Invention
[0003] According to one aspect, a device includes: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the device to at least: operate copies of one or more cloud - native network functions; generate observation data of the copies of the one or more cloud - native network functions, the observation data being generated based on multiple operating conditions of the one or more cloud - native network functions; and use the observation data to apply a causal reasoning function to analyze a causal relationship between at least one observed cause of at least one observed effect of the one or more cloud - native network functions and at least one observed effect of the one or more cloud - native network functions.
[0004] According to one aspect, a device includes: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the device to at least: operate copies of one or more target applications; generate observation data of the copies of the one or more target applications, the observation data being generated based on multiple operating conditions of the one or more target applications; and use the observation data to apply a causal reasoning function to analyze a causal relationship between at least one observed cause of at least one observed effect of the one or more target applications and at least one observed effect of the one or more target applications.
[0005] According to one aspect, a device includes: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the device to at least: select one or more cloud-native network functions; select at least one characteristic of the one or more cloud-native network functions; select at least one environmental condition of the one or more cloud-native network functions; select a load for the one or more cloud-native network functions, the load including: the intensity and duration of processing of the one or more cloud-native network functions under at least one environmental condition; utilize the one or more cloud-native network functions to perform at least one experiment based on at least one characteristic, the load, and at least one environmental condition of the one or more cloud-native network functions; collect observation data from the at least one experiment; and use the observation data to apply a causal inference function to analyze the causal relationship between at least one observed cause of at least one observed effect of the one or more cloud-native network functions and the at least one observed effect of the one or more cloud-native network functions.
[0006] According to one aspect, a method includes: operating replicas of one or more cloud-native network functions; generating observation data for the replicas of the one or more cloud-native network functions, the observation data being generated based on multiple operating conditions of the one or more cloud-native network functions; and using the observation data to apply a causal inference function to analyze the causal relationship between at least one observed cause of at least one observed effect of the one or more cloud-native network functions and the at least one observed effect of the one or more cloud-native network functions.
[0007] According to one aspect, a method includes: operating replicas of one or more target applications; generating observation data for the replicas of the one or more target applications, the observation data being generated based on multiple operating conditions of the one or more target applications; and using the observation data to apply a causal inference function to analyze the causal relationship between at least one observed cause of at least one observed effect of the one or more target applications and the at least one observed effect of the one or more target applications.
[0008] According to one aspect, a method includes: selecting one or more cloud-native network functions; selecting at least one characteristic of the one or more cloud-native network functions; selecting at least one environmental condition of the one or more cloud-native network functions; selecting a load for the one or more cloud-native network functions, the load including: the intensity and duration of processing of the one or more cloud-native network functions under at least one environmental condition; using the one or more cloud-native network functions to perform at least one experiment based on at least one characteristic, the load, and at least one environmental condition of the one or more cloud-native network functions; collecting observational data from the at least one experiment; and using the observational data to apply a causal inference function to analyze the causal relationship between at least one observed cause of at least one observed effect of the one or more cloud-native network functions and at least one observed effect of the one or more cloud-native network functions.
[0009] According to one aspect, a device includes: means for operating replicas of one or more cloud-native network functions; means for generating observational data for the replicas of the one or more cloud-native network functions, the observational data being generated based on multiple operating conditions of the one or more cloud-native network functions; and means for using the observational data to apply a causal inference function to analyze the causal relationship between at least one observed cause of at least one observed effect of the one or more cloud-native network functions and at least one observed effect of the one or more cloud-native network functions.
[0010] According to one aspect, a device includes: means for operating replicas of one or more target applications; means for generating observational data for the replicas of the one or more target applications, the observational data being generated based on multiple operating conditions of the one or more target applications; and means for using the observational data to apply a causal inference function to analyze the causal relationship between at least one observed cause of at least one observed effect of the one or more target applications and at least one observed effect of the one or more target applications.
[0011] According to one aspect, a device includes: components for selecting one or more cloud-native network functions; components for selecting at least one characteristic of one or more cloud-native network functions; components for selecting at least one environmental condition of one or more cloud-native network functions; components for selecting a load of one or more cloud-native network functions, the load including: the intensity and duration of the processing of one or more cloud-native network functions under at least one environmental condition; components for performing at least one experiment using one or more cloud-native network functions, the experiment being based on at least one characteristic, the load, and at least one environmental condition of one or more cloud-native network functions; components for collecting observation data from at least one experiment; and components for using the observation data to apply a causal inference function to analyze the causal relationship between at least one observed cause of at least one observed effect of one or more cloud-native network functions and at least one observed effect of one or more cloud-native network functions.
[0012] According to one aspect, there is provided a machine-readable non-transitory program storage device tangibly embodying a machine-executable instruction program for performing operations, the operations including: operating copies of one or more cloud-native network functions; generating observation data of the copies of one or more cloud-native network functions, the observation data being generated based on multiple operating conditions of one or more cloud-native network functions; and using the observation data to apply a causal inference function to analyze the causal relationship between at least one observed cause of at least one observed effect of one or more cloud-native network functions and at least one observed effect of one or more cloud-native network functions.
[0013] According to one aspect, a machine-readable non-transitory program storage device tangibly embodies a machine-executable instruction program for performing operations, the operations including: operating copies of one or more target applications; generating observation data of the copies of one or more target applications, the observation data being generated based on multiple operating conditions of one or more target applications; and using the observation data to apply a causal inference function to analyze the causal relationship between at least one observed cause of at least one observed effect of one or more target applications and at least one observed effect of one or more target applications.
[0014] According to one aspect, there is provided a machine-readable non-transitory program storage device tangibly embodying a machine-executable instruction program for performing operations, the operations including: selecting one or more cloud-native network functions; selecting at least one characteristic of one or more cloud-native network functions; selecting at least one environmental condition of one or more cloud-native network functions; selecting a load for one or more cloud-native network functions, the load including: the intensity and duration of processing of one or more cloud-native network functions under at least one environmental condition; performing at least one experiment using one or more cloud-native network functions, the experiment being based on at least one characteristic, load, and at least one environmental condition of one or more cloud-native network functions; collecting observation data from at least one experiment; and using the observation data to apply a causal inference function to analyze a causal relationship between at least one observed cause of at least one observed effect of one or more cloud-native network functions and at least one observed effect of one or more cloud-native network functions. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The above aspects and other features are explained in the following description taken in conjunction with the accompanying drawings.
[0016] Figure 1 Illustrates a CI / CD pipeline.
[0017] Figure 2 Depicts the left-shift paradigm framework in a CI / CD pipeline.
[0018] Figure 3 Support for a CI / CD left-shift architecture pipeline is illustrated by examples described herein.
[0019] Figure 4 Depicts the layers of a causal inference system for an operational twin advanced architecture.
[0020] Figure 5 Depicts the lower build box (chaos framework) for the operational twin causal inference system described herein.
[0021] Figure 6 Depicts the upper build box (causal inference) for the operational twin causal inference system described herein.
[0022] Figure 7 Depicts the general problem statement for the operational twin causal inference system described herein.
[0023] Figure 8 Depicts the high-level task description of the operational twin causal inference system described herein.
[0024] Figure 9 Depicts the digital twin component of the operational twin causal inference system described herein.
[0025] Figure 10 Depicts the digital twin component workflow of the operational twin causal inference system described herein.
[0026] Figure 11 Depicts the causal inference component.
[0027] Figure 12 Depicts the implementation of a randomized controlled trial (RCT).
[0028] Figure 13 Depicts the causal inference workflow.
[0029] Figure 14 Is a block diagram of one possible non - limiting system in which example embodiments may be practiced.
[0030] Figure 15 Is an example apparatus configured to implement the examples described herein.
[0031] Figure 16 Illustrates a representation of an example of a non - volatile storage medium.
[0032] Figure 17 Is an example method for implementing the examples described herein.
[0033] Figure 18 Is an example method for implementing the examples described herein.
[0034] Figure 19 Is an example method for implementing the examples described herein. Detailed Description
[0035] Modern software, including 5G cloud - native network functions (e.g., CNF) devices, is typically cloud - based, highly dynamic, and service - oriented. These software devices are often important for configuration but difficult to optimize, especially in an environment where the infrastructure exhibits non - ideal conditions such as a production setting.
[0036] Industry practice for software development methods dictates that a continuous integration / continuous delivery (CI / CD) pipeline should generally resemble Figure 1 The pipeline 101 shown.
[0037] In accordance with these terms, the 5G CNF development process flows from left to right. For example, source code is checked into a repository (version control 202), then built 204, tested 206, released 208, deployed 210, operated 212, and monitored 214. Refer to Figure 2, certain tasks that typically occur later during development can be brought forward in the pipeline - the left-shift principle includes left-shifting 216 - to help improve productivity by identifying and correcting issues early in the development cycle, thus reducing costs and improving compliance with requirements. This principle is central to the CAROT disclosure. Figure 2 The CAROT (e.g., Operational Twin Causal Reasoning) framework 201 is depicted in Figure 2 , where the operational environmental conditions (including during operations 212 and monitoring 214 in the pipeline) are brought into the test 206 phase, thus effectively reconstructing the digital twin.
[0038] The causal reasoning of the operational twin system replicates the operational environmental conditions in a safe, controllable, and repeatable manner. Its novelty includes: analyzing the observations collected from the digital twin setup to infer possible cause / effect relationships, thus providing insights into the configuration settings, design assertions, and hypothetical scenarios of 5G CNFs, and supporting the root cause analysis (RCA) basis for zero-touch management automation.
[0039] The ideas described herein generally apply to any containerized software development production cycle, however the focus of the examples described herein is on associating the system with CNFs (Cloud Native Network Functions) (specifically CNFs related to 5G radio and core infrastructure).
[0040] Throughout the submission, all references to software, applications, devices, software appliances / applications can be understood as 5G-related CNFs, which are microservices-based software applications that support functional communication networks (e.g., 5G SA or NSA). One or more of the following technical effects can be selected: verifying 5G CNFs in the operational environment replica and digital twin, robust operation of the communication network based on 5G CNFs, optimizing the configuration of the communication network based on 5G CNFs, discovering insights into unforeseen / unplanned 5G CNF behavior, and / or continuously improving the design of the communication network based on 5G CNFs.
[0041] The examples described herein include the following features:
[0042] 1. Generally, as Figure 3 shown, the overall positioning of the CAROT framework supports the "left-shift" CI / CD paradigm.
[0043] The causal inference system 201 of the operational twin system spans the engineering layer and includes results 302, ML 304, data 306, and system 308. Results 302 include optimal code 310, RCA and faults 324, and RCA and fixes 326. ML 304 includes code review 312, causal analysis 328, and anomaly detection 330. Data 306 includes code configuration 314 and production metrics 332. System 308 includes IDE 316 and production 334. Development 318 includes optimal code 310, code review 312, code configuration 314, and IDE 316. CAROT 201 is a component of the development operation pipeline 320 and implements simulation and digital twins 322. Operations 336 include RCA and faults 324, RCA and fixes 326, causal analysis 328, anomaly detection 330, production metrics 332, and production 334.
[0044] As Figure 3 shown, CAROT 201 receives inputs from IDE 316 and production metrics 332 and provides outputs to optimal code 310, causal analysis 328, and production 334.
[0045] 2. As Figure 4 depicted in the high-level architecture shown, CAROT 201 is built on top of multiple specialized components. The specialized components include chaos framework 402, chaos metrics 404, causal inference 406, causal discovery 408, and design optimization 410. Causal discovery provides results to causal analysis 328 through interface 412, and causal discovery 408 receives information from production metrics 332 through interface 412. The results of design optimization 410 are provided through interface 410 to develop optimal code 310. CAROT 201 (including chaos framework 402) receives inputs from IDE 316 through interface 414. The term "chaos" refers to studying fault generation and fault observations by changing the configuration, parameters, or components of a system.
[0046] 2a. CAROT builds on top of it a digital twin engine that includes an automated planned fault injection (including chaos fault injection) and load stress (e.g., REST and network traffic benchmark) framework. This subjects the target application (5G CNF) to configurable workloads during which planned infrastructure non-optimal conditions can be injected and observed through the captured metrics. As Figure 5 emphasized in the system component 502 shown, this capability is provided by the lower two boxes (chaos framework 402 and chaos metrics 404) in the CAROT framework high-level architecture. Enabling the observations to help strengthen or rule out measurable quantitative hypotheses (e.g., one or more hypotheses) about application performance and / or robustness.
[0047] 2b. As described in (2a), the automation supported by the lower two boxes (Chaos Framework 402 and Chaos Metrics 404) provides repeatable stimuli for load stress and fault sensing to generate a wide range of applicable (depending on the application type and observation requirements, which are also configurable) observation datasets. These are processed by causal inference (machine learning) algorithms that provide inferences or insights into the relationship between causes and their reasonable effects on the application itself, represented by the upper boxes of Causal Inference 406 and Causal Discovery 408 in the advanced architecture, including Figure 6 the system components 602 shown. These insights include: measuring and understanding treatment effects (e.g., the magnitude of the effect change due to a change in the cause), discerning which application configuration options are optimal for a specific network setup, and identifying which cloud infrastructure choices can have the greatest impact on overall performance and robustness.
[0048] 2c. The inferences from (2b) are intended to provide feedback to the application design phase, to prove or disprove early theoretical assumptions, and potentially to identify unplanned behavior. This is represented by Figure 6 the Design Optimization box (410) in. These insights can also be used to supplement root cause analysis (RCA) assumptions and models, including RCA and Fault 324 and RCA and Fix 326, as depicted by the data flow shown in Figure 4 which includes the Causal Relationship Analysis box (328).
[0049] Before an application is released to the production environment, the application is compiled and tested to verify compliance with functional and non-functional requirements. In modern software settings, these tests are typically integrated as part of a continuous integration / continuous deployment (CI / CD) pipeline immediately after the application is built (compiled) in order to detect and fix any potential unplanned issues early in the development cycle, thereby reducing costs and increasing productivity.
[0050] However, in many cases, these test tasks are typically performed in a cloud infrastructure that shows ideal conditions, making it difficult to identify configuration and design issues that may arise after the application is deployed to a production environment where resources may be limited and the application typically shows signs of stress.
[0051] Additionally, this traditional approach fails to explore hypothetical scenarios that can help identify configuration setting or design principle issues of an application that would otherwise not be detected during the normal execution of CI / CD tasks and may have a negative impact on its expected performance and behavior (SLO) after deployment to a production operating environment.
[0052] Examples include determining or refining infrastructure thresholds (such as increasing or decreasing specific resource availability), and how these thresholds affect the application to help save costs (e.g., suppressing resources with negligible impact), refining environmental requirements, validating scaling and performance capabilities, more effectively identifying root causes, etc.
[0053] Motivations for developing the systems described herein include: 99% of misconfiguration events in the cloud go unnoticed (Help Net Security, September 25, 2019), cloud misconfigurations cost companies nearly $5 trillion (TechRepublic, February 20, 2020), and it is well known that non-functional failures are difficult. Hardware performance failures in large production systems fail slowly at scale.
[0054] Therefore, it is important to emphasize the relevance of the left-shift software development paradigm approach described herein, as this provides an opportunity to identify and address problems that typically arise much later in the process, to help reduce costs, increase productivity, and meet requirements or expectations. Figure 7 An overall overview of the problem statement is provided.
[0055] As Figure 7 shown in item 702 of, during the test phase 206 of the CI / CD pipeline, due to the environmental conditions not replicating production or generating a digital twin, configuration, performance, or design problems are not identified. As Figure 7 further shown in item 702 of, during the operation phase 212 of the pipeline, design, configuration, or performance problems are identified, but at this stage of the pipeline, design, configuration, or performance problems are costly and tend to require more time to fix, which results in an extended deployment timeline, reduced quality, and a general loss of confidence from stakeholders. As Figure 7 shown in item 704 of, CAROT 201 identifies design, configuration, or performance problems early in the CI / CD pipeline, and earlier than the process shown in item 702, soon after the test phase 206, and before the release phase 208 (and much earlier than the operation phase 212), when corrections are faster and less costly.
[0056] Chaos engineering methods and tools are generally available. This also applies to causal inference machine learning techniques. However, the examples described herein combine these tools and techniques in a unique and novel original solution that, for example, captures observations for a digital twin of a 5G CNF cloud infrastructure (Kubernetes) operating environment, which are aggregated and processed early in the software production cycle for cause / effect insights (causal inference) to achieve the left-shift principle. These components specifically include the left-shift paradigm, digital twin, and causal inference.
[0057] The left shift paradigm includes conditions of a similar operating environment inserted early in the CNF CI / CD testing task.
[0058] Digital twins include deploying and configuring CNFs to expose them to controlled variable rates (benchmarks for REST and network services), replicating optimal and suboptimal cloud infrastructure conditions of the operating environment (chaos engineering), and observing infrastructure and interactions across CNF microservices and components, including monitoring and tracking.
[0059] Causal reasoning includes: determining possible cause / effect relationships, analyzing the magnitude of the impact between cause and effect, leveraging or constructing a directed acyclic graph (DAG) by performing causal discovery, determining root cause analysis (RCA), and developing and identifying hypothetical scenarios to generate inferential insights.
[0060] The automation enabled by the example (CAROT) described in this article allows experiments to be performed based on the scale of the type of target application, develop and iterate permutations of different cloud infrastructure conditions (e.g., digital twins) where possible. The observations and metrics collected become a dataset that is analyzed by causal reasoning algorithms, such as using machine learning to validate existing design assumptions and / or discover new assumptions to analyze causal impacts. The system described in this article can also provide supplementary information for root cause analysis (RCA) of operations in the operating environment.
[0061] The system described in this article, including CAROT, provides a framework that is intended to be integrated into the CI / CD 5G CNF pipeline to explore behavior in typical operating environments, create digital twins, infer possible causes and effects, help explore hypothetical scenarios, assert or ignore CNF design assumptions, identify optimal CNF configuration settings, validate SLO / SLA under non-ideal cloud infrastructure conditions, and support root cause analysis (e.g., zero-touch management automation).
[0062] Figure 8 Components of CAROT are referenced in. The system is intended for 5G CNF CI / CD pipeline integration via an open API, as shown in (802). As Figure 8 shown, the system can be integrated into the test section 206 of the CI / CD pipeline. The system places the target application 816 in a configurable infrastructure (e.g., digital twin) that simulates a production environment: this includes load 814 - REST or network services (804), and cloud infrastructure corruption (806). Injecting chaotic infrastructure corruption into the software application 816 includes complete crashes 818, resource stress 820, network stress 822, and I / O corruption 824.
[0063] The system automatically captures infrastructure and application observations (e.g., tracing and monitoring) for posterior analysis (808). Causal reasoning 406 and discovery 408 algorithms (ML) are applied to the observations 826 along with one or more causal directed acyclic graphs (DAGs) (828) to provide insights (810). The output of the framework (CAROT) involves hypothetical scenarios 830, decision insights 836, important features 832, and a model 834 for RCA that leads to zero-touch management automation (812).
[0064] The system (CAROT) described herein is supported by design principles and two components, which will be described in detail below. These individual methods, techniques, and tools that support the framework have been assembled in this unique configuration for this specific purpose.
[0065] 1. Design principle: Integrated into the CI / CD pipeline - the left-shift paradigm
[0066] Research shows that the cost of fixing problems near or very near the end of the process is much higher than at the beginning when designing and creating the codebase. This approach helps improve the delivery timeline as well as end-consumer satisfaction and confidence. By adopting the left-shift design principle, the CAROT framework can replicate what typically happens in a production environment (e.g., resource stress, hardware failures, oversubscription, etc.) in a planned, repeatable, and controlled setting (digital twin). Since the observations collected are then processed by the causal reasoning system, SLO assumptions and optimal configuration insights can be asserted or negated before the CNF reaches the operational deployment phase. Figure 7 This feature is depicted.
[0067] 2. Digital twin component - Reconstruct the operational environment
[0068] Figure 9 A functional representation of the digital twin component 900 of CAROT is shown. The following description provides additional insights.
[0069] 2a. 5G CNF target - At the center of the image is the 5G CNF 902, which must be packaged in a helm template if deployment and management are handled internally by CAROT. Alternatively, the CNF 902 can reside externally, in which case the target system details, such as the URL endpoints for load localization and monitoring (such as for K8s label chaos feature localization and monitoring), must be provided.
[0070] 2b. 5G CNF Load - Above the application 816 is the load module 814. The load module 814 can target 5G CNF 902 with REST or network traffic such as IMIX traffic. The load level is configurable, and based on user selection, CAROT automatically generates the necessary workers to support it appropriately. The higher the load, the more workers are generated.
[0071] 2c. The overall experiment length (e.g., the amount of time the 5G CNF 902 is subjected to the controlled environment characteristics while being monitored) is set by the load duration, which is a configurable parameter ranging from a few seconds to several days.
[0072] 2d. Environmental Conditions - The infrastructure damage injection module 904 supports a configurable range of cloud infrastructure conditions that can be applied to the environment in which the target application provided by the chaos engineering tool is deployed. The damage characteristics can be combined or completely excluded to represent the ideal state.
[0073] 2e. Observations 826 are collected from the application 816, e.g., using tracing, and the infrastructure is encapsulated and tagged with a unique id maintained in the API persistence layer of the framework for future reference. The infrastructure can include node exporters and cAdvisor via Prometheus.
[0074] Figure 10 Shows the general workflow of the digital twin component. Figure 10 The workflow shown can be related to the description provided above. Three modules manage and operate the workflow / component 1000, namely the API module 1001, the operator module 1003, and the engine module 1005. The API module 1001 is an open API endpoint for all tasks related to requests (e.g., submit, update, cancel, etc.). The API module 1001 maintains the persistence layer of the component. The operator module 1003 manages the auto-scaling of the K8s worker cluster. The engine module 1005 processes the sequence stored in the persistence layer, coordinates the deployment, termination for application and environmental characteristics. The engine module 1005 processes the collection and packaging of observations.
[0075] The component 1000 is designed and capable of vertical and horizontal scaling to allow for the simultaneous execution of a large number of experiments when computing resources are available. For reference, the component 1000 has processed tens of thousands of short-duration experiments.
[0076] At 1002, the system determines whether the target application 816 (e.g., 5G CNF 902) is internally managed. If the system determines at 1002 that the 5G CNF is not internally managed but externally managed, then at 1004, the system provides endpoint details such as a URL or a label. If the system determines at 1002 that the 5G CNF is internally managed, then at 1006, the system selects the target 5G CNF that includes replicas. At 1008, the system selects 5G CNF features that include computing characteristics and instances. From items 1004 and 1008, the method transitions to 1010. At 1010, the system selects 5G CNF loads that include intensity and duration. At 1012, the system selects environmental conditions such as type, intensity, and periodicity. At 1014, the system determines whether the 5G CNF is internally managed.
[0077] If at 1014, the system determines that the 5G CNF is internally managed, then the method transitions to 1016. If at 1014, the system determines that the 5G CNF is not internally managed but externally managed, then the method transitions to 1018. At 1016, the 5G CNF is deployed, for example, using a helm chart or a template. At 1018, the environmental conditions are deployed, for example, using a helm chart or a template. At 1020, the load is deployed, for example, using a helm chart or a template. At 1022, the system waits for the experiment to complete. At 1024, the system collects, packages, and labels the observations 826 such as metrics. At 1026, one or more causal inference components (such as causal inference 406 or causal discovery 408) process the observations 826. At 1028, the 5G CNF, the load, and the chaos deployment are terminated, for example, using a helm chart or a template. At 1030, the method ends.
[0078] 3. Causal Inference - Intelligent Inference of the Performance of the Target Application (e.g., 5G CNF)
[0079] This component includes applying causal inference techniques to the observations captured by the digital twin module to generate application configuration optimization and performance insight discovery.
[0080] Causal inference breaks traditional machine learning methods because it attempts to answer the reasons behind the decision-making process, effectively computationally solving the counterfactual hypothesis problem: for example, if vCPU resource type X is used for the application, how good will the latency KPI be? Causal inference is a step towards general artificial intelligence (AGI), and it is actually applied to software configuration and performance optimization problems in the examples described in this article.
[0081] Reference Figure 11The causal inference component 1100 shown is constructed based on four modules, which include generating and collecting observational data (1104, 1116, 402), formulating inputs (1103, 1122, 1128), performing causal inference (1102, 1110, 1112, 1114, 1118), and causal discovery (1124, 1126), as well as design optimization 410.
[0082] Generating and collecting observational data (1104, 1116, 402) - The first step in this process is to collect observational data 1108 using at least part of module 1103. In the context of the example described herein, the observational data 1108 is automatically collected from the digital twin component. The data 1108 is generated and stored in tabular form (1122) such that each record corresponds to an experiment performed (1116) and each record is associated with a set of attributes. The set of these attributes corresponds to nodes in a directed acyclic graph (DAG) in which causal inference calculations are performed. The attributes can be related to configurations (e.g., cloud computing settings), controls (e.g., chaos characteristics), or observable metrics and combinations thereof. As Figure 11 shown, the chaos framework 402 includes Domain Name System (DNS) corruption 801, kernel crash 803, complete crash 818, network stress 822, computational resource stress 820, HTTP stress 805 (e.g., internet stress), time corruption 807, web service (WS) chaos or crash 809, and input / output (I / O) corruption 824. The chaos framework 402 can also include network load and application programming interface load. The run experiment box (1116) sets chaos conditions (1107), runs experiments (402), makes observations by collecting data (1105), and then repeats the process multiple times. The collective observational data is stored (1115) in a data set (1122). Module 1104 includes chaos metrics 404 and the chaos framework 402.
[0083] Formalized Input (1103, 1122, 1128) - To perform causal reasoning, the process requires at least in part a set 1122 of observational data generated 1115 using experiment 1116, and a causal graph (one of the causal graphs 1128), such as a DAG that can be formulated using domain expert knowledge. The causal graph is generated (1127) using at least ambiguity removal 1126. The person designing the DAG carefully defines the edges connecting the nodes. The edges represent a possible causal relationship from a parent node to a child node. For example, the parent node defines the cause 1130, and the child node defines the effect 1132, where the effect 1132 can be one or more KPIs. In a DAG, the edges are essentially unidirectional, and the acyclic name means that there cannot be any cycles in the graph structure. In causal reasoning analysis, cycles are considered useless because it is almost impossible to decipher causality between them. The strength of the causal influence between two nodes does not need to be defined during the design phase. During the design phase, all that is simply defined is that a causal relationship can exist. When in doubt, it is recommended to add an edge because ignoring an edge is considered a very strong assumption. Essentially, the causal reasoning function is to quantitatively analyze and calculate the strength between a change in the cause and its causal impact on the effect attributes. This will be described in more detail below.
[0084] Perform causal reasoning (1102, 1110, 1112, 1114, 1118) - This module is responsible for processing the raw observation data 1122 using the input from the DAG (e.g., using one or more causal graphs 1128) to perform causal reasoning 406. The metrics are reflected by the causality 1110, which can include ATE (Average Treatment Effect), ATT (Average Treatment Effect on the Treated), conditional ATE, ITE (Individual Treatment Effect), mediation analysis such as NDE (Natural Direct Effect) and NIE (Natural Indirect Effect). Any configuration or control node can be a treatment candidate, and any value within that node can be used as a control or baseline, and some other values will be considered treatment values. For example, by checking the throughput of a service (e.g., the DAG result node) average treatment effect (e.g., ATE) when modifying the K8s cluster worker node type (e.g., the treatment node), where the application 816 (5G CNF 902) is deployed from a low-power CPU to a medium-power CPU. In this case, it can be considered that the worker node type is being treated from a low-power CPU to a medium-power CPU, and the aim is to understand the impact on the service throughput. This part of the design generally involves using doCalculus (see 1106) or some other causal reasoning technique as a first step to perform the formulation of the estimate (see identification 1114) using the input 1111 from the causal graph 1128 and the input 1113 from the column-row data 1122. This allows the formulation of counterfactual scenarios - the reasoning aspect - which then allows the next step of invoking the statistical analysis technique 1112 to calculate the causal reasoning result metrics described herein. The output 1119 of the ML and statistical analysis 1112 is provided to the robustness check 1118, and another output 1121 of the ML and statistical analysis 1112 is provided to the causality module 1110. The output 1123 of the identification 1114 is provided to the ML and statistical analysis 1112. Additionally, there are algorithm-based refutation techniques (1118) that can be used to confirm the robustness and credibility of the CI results. The output 1117 of the robustness check 1118 is provided to the causality module 1110. For details, see the DoWhy library. Item 1102 includes both causal reasoning 406 and causal discovery 408.
[0085] The DAG 1128 (one or more DAGs 1128) can 1) be derived using a causal discovery process (1124, 1125, 1126, 1127), or 2) the DAG 1128 (one or more DAGs 1128) can be proposed (e.g., developed) by a human expert.
[0086] Causal discovery (1124, 1126) - This module can be considered optional. A DAG can be semi-automatically generated using a causal discovery (CD) process 1124 and input 1125 from observational data 1122. Many of these algorithms can be used to infer a graph by using observational data 1122 as input. However, considering the state-of-the-art techniques in this field, the resulting graph may not necessarily be a DAG inherently, as edges may lack direction or show other types of ambiguity. Expert knowledge is usually still applied to help remove the ambiguity in the graph 1126 to formulate a DAG. The causal discovery process may generate many candidate graphs, among which expert knowledge is applied to select the closest one. The resulting causally discovered DAG can be used as part of a regular causal reasoning process. Additionally, this DAG can be used for root cause analysis. In some use cases, the causal discovery function can be used independently, rather than just as an auxiliary function for the causal reasoning process.
[0087] Design optimization - This final step involves evaluating the resulting causal reasoning to search for the validation or exclusion of the hypotheses established during the design phase. Design tasks can be done manually in CAROT. Alternatively, with sufficient automation, design optimization can be tightly integrated into the CI / CD pipeline to streamline the entire process.
[0088] Another feature of the system described in this paper is the ability to perform RCTs (randomized controlled trials) 1202 in a digital twin environment, as Figure 12 shown and referenced in the advanced architecture of. RCTs are generally touted as the gold standard for causal relationship analysis and treatment effect studies, and at least the baseline standard.
[0089] However, due to cost, ethics, and many other practical limitations, RCTs are not always feasible. Another limitation of RCTs is that they usually can only study one causal relationship pair at a time.
[0090] In the case of CAROT, RCTs are performed by testing one treatment effect at a time using digital twin components. When considering the possible permutations of treatment effects, the described causal reasoning process remains the most advantageous. If the goal is to understand many specific treatment combinations, the causal reasoning process is an effective process. In this case, the RCT setup is used as a validation function to spot-check the causal reasoning results, and specifically study treatment result pairs and compare the results with those generated by the causal reasoning analysis. Figure 12 An interface 415 used to provide information from the chaos framework 402 to the experiment 1116 is also shown in.
[0091] Figure 13FIG. 1300 shows the workflow of the individual modules of the cross-causal reasoning component 1100. Observation data is generated and collected (1108), including generating digital twin observations (1104) and formatting the observations into rows in a table (1116), where the columns represent experimental features and KPIs. At 1302, a determination is made as to whether to apply causal discovery. If it is determined at 1302 to apply causal discovery, the method transitions to causal discovery 1304, and if it is determined at 1302 not to apply causal discovery, the method transitions to formulate input 1122. Optional causal discovery 1304 includes automatic DAG generation 1124 and ambiguity removal 1126, and ambiguity removal 1126 can be a human-assisted process. Formulating input 1122 includes generating a causal graph using domain knowledge (1128).
[0092] After generating and collecting observation data (1108) and formulating the input (1122), the method transitions to causal reasoning (1102). Causal reasoning 1102 includes identification 1114, ML and statistical analysis 1112, a robustness check 1118 that can include refutation, and generating causals 1110, such as ATE, CATE, and causal tree generation. Design optimization 410 is performed after causal reasoning 1102, and then the process ends at 1306.
[0093] The operational twin causal reasoning system described herein provides several advantages and technical effects. The unique design of the system allows for inferring causal relationships from an environment that is very similar to production conditions by creating digital twins. This provides several advantages to help improve the overall application development cycle and quality, particularly for validating or excluding applications, such as 5G CNF design assumptions regarding ideal and impaired environments, including performance, robustness, and application configuration settings. The system also provides functionality for validating or excluding applications, such as 5G CNF design assumptions regarding platform requirements such as computing and network requirements. The system also includes functionality for validating root cause assumptions in zero-touch and troubleshooting, and for identifying potential applications (e.g., 5G CNF issues) long before they are encountered in a production environment.
[0094] Thus, the examples described herein relate to software development, agile methods, application architectures for deployment (e.g., CI / CD), and testing of applications in consumer production environments (e.g., 5G CNF). The examples described herein focus on the application of causal reasoning (ML) techniques in 5G CNF software testing in a CI / CD pipeline, as well as the ability to analyze hypothetical scenarios, the ability to solve counterfactual problems, and the ability to provide equivalent treatment effect estimation for performing causal reasoning.
[0095] Go to Figure 14, which shows a block diagram of one possible and non - limiting example in which the example can be practiced. A user equipment (UE) 110, a radio access network (RAN) node 170, and one or more network elements 190 are shown. In Figure 14 the example, the user equipment (UE) 110 communicates wirelessly with a wireless network 100. The UE is a wireless device that can access the wireless network 100. The UE 110 includes one or more processors 120, one or more memories 125, and one or more transceivers 130 interconnected by one or more buses 127. Each of the one or more transceivers 130 includes a receiver Rx132 and a transmitter Tx 133. The one or more buses 127 can be address, data, or control buses and can include any interconnection mechanism such as a series of lines on a motherboard or integrated circuit, optical fibers, or other optical communication devices, etc. The one or more transceivers 130 are connected to one or more antennas 128. The one or more memories 125 include computer program code 123. The UE110 includes a module 140, which includes a part or both of 140 - 1 and / or 140 - 2, and the module can be implemented in a variety of ways. The module 140 can be implemented in hardware as module 140 - 1, such as being implemented as part of one or more processors 120. The module 140 - 1 can also be implemented as an integrated circuit or by other hardware such as a programmable gate array. In another example, the module 140 can be implemented as module 140 - 2, which is implemented as computer program code 123 and executed by one or more processors 120. For example, the one or more memories 125 and the computer program code 123 can be configured to, together with one or more processors 120, cause the user equipment 110 to perform one or more of the operations described herein. The UE 110 communicates with the RAN node 170 via a wireless link 111.
[0096] In this example, the RAN node 170 is a base station that provides access to the wireless network 100 for wireless devices such as the UE 110. The RAN node 170 can be, for example, a base station for 5G, also known as New Radio (NR). In 5G, the RAN node 170 can be an NG-RAN node, which is defined as a gNB or an ng-eNB. A gNB is a node that provides NR user plane and control plane protocol termination towards the UE and is connected to the 5GC (e.g., (one or more) network elements 190) via an NG interface (such as connection 131). An ng-eNB is a node that provides E-UTRA user plane and control plane protocol termination towards the UE and is connected to the 5GC via an NG interface (such as connection 131). The NG-RAN node can include multiple gNBs, which can also include a Central Unit (CU) (gNB-CU) 196 and one or more Distributed Units (DU) (gNB-DU), where the DU 195 is shown. Note that the DU 195 can include or be coupled to and control a Radio Unit (RU). The gNB-CU 196 is a logical node that hosts the Radio Resource Control (RRC), SDAP, and PDCP protocols of the gNB or controls the operation of one or more gNB-DUs, or the RRC and PDCP protocols of the en-gNB. The gNB-CU 196 terminates at the F1 interface connected to the gNB-DU 195. The F1 interface is shown as reference numeral 198, although reference numeral 198 also shows the link between the remote element of the RAN node 170 and the centralized element of the RAN node 170, such as the link between the gNB-CU 196 and the gNB-DU 195. The gNB-DU 195 is a logical node that hosts the RLC, MAC, and PHY layers of the gNB or en-gNB, and its operation is partially controlled by the gNB-CU 196. One gNB-CU 196 supports one or more cells. One cell can be supported by one gNB-DU 195, or one cell can be supported / shared by multiple DUs under RAN sharing. The gNB-DU 195 terminates at the F1 interface 198 connected to the gNB-CU 196. Note that the DU 195 is considered to include the transceiver 160, for example, as part of the RU, but some examples in this regard can have the transceiver 160 as part of a separate RU, for example, under the control of the DU 195 and connected to the DU 195. The RAN node 170 can also be an eNB (evolved NodeB) base station for LTE (Long-Term Evolution), or any other suitable base station or node.
[0097] The RAN node 170 includes one or more processors 152, one or more memories 155, one or more network interfaces ((multiple) N / W I / F) 161, and one or more transceivers 160 interconnected by one or more buses 157. Each of the one or more transceivers 160 includes a receiver Rx 162 and a transmitter Tx 163. The one or more transceivers 160 are connected to one or more antennas 158. The one or more memories 155 include computer program code 153. The CU 196 may include (multiple) processors 152, (multiple) memories 155, and a network interface 161. Note that the DU 195 may also include its own one / more memories and (multiple) processors, and / or other hardware, but these are not shown.
[0098] The RAN node 170 includes a module 150, which includes part or both of 150-1 and / or 150-2, and the module can be implemented in various ways. The module 150 can be implemented in hardware as module 150-1, such as being implemented as part of one or more processors 152. The module 150-1 can also be implemented as an integrated circuit or by other hardware such as a programmable gate array. In another example, the module 150 can be implemented as module 150-2, which is implemented as computer program code 153 and executed by one or more processors 152. For example, the one or more memories 155 and the computer program code 153 are configured to work with one or more processors 152 such that the RAN node 170 performs one or more of the operations described herein. Note that the functions of the module 150 can be distributed, such as being distributed between the DU 195 and the CU 196, or implemented separately in the DU 195.
[0099] One or more network interfaces 161 communicate via a network (such as via links 176 and 131). Two or more gNBs 170 can communicate using, for example, link 176. The link 176 can be wired or wireless or both, and can implement, for example, the Xn interface for 5G, the X2 interface for LTE, or other suitable interfaces for other standards.
[0100] One or more buses 157 can be address, data, or control buses and can include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optic or other optical communication devices, wireless channels, etc. For example, one or more transceivers 160 can be implemented as a remote radio head (RRH) 195 for LTE, or a distributed unit (DU) 195 implemented for a gNB for 5G, where other elements of the RAN node 170 may be physically located at a different location from the RRH / DU 195, and one or more buses 157 can be partially implemented as, for example, a fiber optic cable, or other suitable network connection that connects other elements of the RAN node 170 (e.g., a central unit (CU), gNB-CU 196) to the RRH / DU 195. The reference numeral 198 also indicates these suitable (multiple) network links.
[0101] Note that the description herein indicates that a "cell" performs a function, but it should be clear that the device forming the cell can perform the function. A cell forms part of a base station. That is, each base station can have multiple cells. For example, there can be three cells for a single carrier frequency and associated bandwidth, each cell covering one-third of a 360-degree area, so that the coverage area of a single base station covers an approximately elliptical or circular shape. In addition, each cell can correspond to a single carrier, and a base station can use multiple carriers. So if there are 3 cells of 120 degrees for each carrier and there are 2 carriers, the base station has a total of 6 cells.
[0102] The wireless network 100 may include one or more network functions 190, which may include core network functions and provide connectivity to other networks such as a telephone network and / or a data communication network (e.g., the Internet) via one or more links 181. Such core network functions for 5G may include location management function(s) ((multiple) LMF) and / or (multiple) access and mobility management function(s) ((multiple) AMF) and / or user plane function(s) ((multiple) UPF) and / or (multiple) session management function(s) ((multiple) SMF). Such core network functions for LTE may include MME (Mobility Management Entity) / SGW (Serving Gateway) function. Such core network functions may include SON (Self-Organizing / Optimizing Network) function. These are merely example functions that may be supported by (multiple) network functions 190, and note that both 5G functions and LTE functions may be supported. The RAN node 170 is coupled to the network function 190 via the link 131. The link 131 may be implemented as, for example, the NG interface for 5G, or the S1 interface for LTE, or other suitable interfaces for other standards. The network function 190 includes one or more processors 175, one or more memories 171, and one or more network interfaces ((multiple) N / W I / F) 180 interconnected by one or more buses 185. One or more memories 171 include computer program code 173.
[0103] The wireless network 100 may implement network virtualization, which is the process of combining hardware and software network resources and network functions into a single software-based management entity (virtual network). Network virtualization involves platform virtualization, which is typically combined with resource virtualization. Network virtualization is classified as external network virtualization or internal network virtualization. External network virtualization combines many networks or parts of networks into virtual units, and internal network virtualization provides network-like functions to software containers on a single system. Note that the virtualized entities resulting from network virtualization are still implemented using hardware such as processors 152 or 175 and memories 155 and 171 to some extent, and such virtualized entities also create technical effects.
[0104] The computer-readable memories 125, 155, and 171 can be of any type suitable for the local technical environment and can be implemented using any suitable data storage technology, such as semiconductor-based memory devices, flash memory, magnetic memory devices and systems, optical memory devices and systems, non-transitory memory, transitory memory, fixed memory, and removable memory. The computer-readable memories 125, 155, and 171 can be components for performing storage functions. The processors 120, 152, and 175 can be of any type suitable for the local technical environment, and by way of non-limiting example, can include one or more of a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), and a processor based on a multi-core processor architecture. The processors 120, 152, and 175 can be components for performing functions such as controlling the UE 110, the RAN node 170, the network function(s) 190, and other functions described herein.
[0105] Generally, various embodiments of the user equipment 110 can include, but are not limited to, a cellular phone with wireless communication capabilities (such as a smart phone, a tablet computer, a personal digital assistant (PDA)), a portable computer with wireless communication capabilities, an image capture device with wireless communication capabilities (such as a digital camera), a game device with wireless communication capabilities, a music storage and playback device with wireless communication capabilities, an Internet device that allows wireless Internet access and browsing, a tablet computer with wireless communication capabilities, a head-mounted display (such as a display implementing virtual / augmented / mixed reality), and a portable unit or terminal that combines a combination of such functions.
[0106] The UE 110, the RAN node 170, and / or the network function(s) 190 (and associated memories, computer program code, and modules) can be configured to implement (e.g., in part) the methods described herein, including the operational twin causal reasoning system (CAROT) for the development and operation of 5G CNFs. Thus, Figure 14 The computer program code 123, the modules 140-1, the modules 140-2, and other elements / features of the UE 110 shown can implement the user equipment-related aspects of the methods described herein. Similarly, Figure 14 The computer program code 153, the modules 150-1, the modules 150-2, and other elements / features of the RAN node 170 shown can implement the gNB / TRP-related aspects (such as for a target gNB or a source gNB) of the methods described herein. Figure 14 The computer program code 173 and other elements / features of the network function(s) 190 shown can be configured to implement the network function / element-related aspects (such as for an OAM node) of the methods described herein.
[0107] Figure 15 It is the example device 1500, which can be implemented in hardware and is configured to implement the examples described herein. The device 1500 includes at least one processor 1502 (e.g., FPGA and / or CPU), and at least one memory 1504 including computer program code 1505, wherein the at least one memory 1504 and the computer program code 1505 are configured to, together with the at least one processor 1502, cause the device 1500 to implement circuitry, processes, components, modules or functions (collectively referred to as controls 1506) to implement the examples described herein, including an operational twin causal reasoning system (CAROT) for the development and operation of 5G CNF. The memory 1504 can be non-transitory memory, transitory memory, volatile memory (e.g., RAM) or non-volatile memory (e.g., ROM).
[0108] The device 1500 optionally includes a display and / or an I / O interface 1508, which can be used to display aspects or states of the methods described herein (e.g., when one of these methods is being executed or will be executed at a later time), or to receive input from a user, such as using a keypad, a camera, a touch screen, a touch area, a microphone, biometrics, one or more sensors, etc. The device 1500 includes one or more communication interfaces ((plural) I / F) 1510, such as one or more network (N / W) interfaces. The (plural) communication I / F 1510 can be wired and / or wireless and communicate via any communication technology over the Internet / (plural) other networks. The (plural) communication I / F 1510 can include one or more transmitters and one or more receivers. The (plural) communication I / F 1510 can include standard well-known components, such as amplifiers, filters, frequency converters, modulators (demodulators) and encoder / decoder circuitry and one or more antennas.
[0109] The apparatus 1500 that implements the functions of the control 1506 can be the UE 110, the RAN node 170 (e.g., gNB), the (multiple) network elements 190, or any other example described herein. Thus, the processor 1502 can correspond to the (multiple) processors 120, the (multiple) processors 152, and / or the (multiple) processors 175, the memory 1504 can correspond to the (multiple) memories 125, the (multiple) memories 155, and / or the (multiple) memories 171, the computer program code 1505 can correspond to the computer program code 123, the modules 140-1, the modules 140-2, and / or the computer program code 153, the modules 150-1, the modules 150-2, and / or the computing program code 173, and the (multiple) communication I / Fs 1510 can correspond to the transceiver 130, the (multiple) antennas 128, the transceiver 160, the (multiple) antennas 158, the (multiple) N / W I / Fs 161, and / or the (multiple) N / W I / Fs 180. Alternatively, the apparatus 1500 may not correspond to any of the UE 110, the RAN node 170, or the (multiple) network functions 190, because the apparatus 1500 can be part of a self-organizing / optimizing network (SON) node, such as in the cloud.
[0110] The apparatus 1500 can also be distributed throughout the network (e.g., 100), including inside and between the apparatus 1500 and any network elements, such as the network control function (NCE) 190 and / or the RAN node 170 and / or the UE 110.
[0111] As Figure 15 shown, the interface 1512 enables data communication between the various items of the apparatus 1500. For example, the interface 1512 can be one or more buses, such as an address, data, or control bus, and can include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, optical fiber, or other optical communication device, etc. The computer program code 1505 that includes the control 1506 can include object-oriented software that is configured to transfer data and messages between objects within the computer program code 1505. The apparatus 1500 does not need to include each of the above features, or may also include other features.
[0112] Figure 16 A schematic diagram of non-volatile storage media 1600a (e.g., a computer optical disc (CD) or digital versatile disc (DVD)) and 1600b (e.g., a universal serial bus (USB) memory stick) storing instructions and / or parameters 1602 is shown, which, when executed by a processor, allows the processor to perform one or more steps of the methods described herein.
[0113] Figure 17is an example method 1700 for implementing the example embodiments described herein. At 1710, the method includes operating replicas of one or more cloud-native network functions. At 1720, the method includes generating observational data of the replicas of one or more cloud-native network functions, the observational data being generated based on multiple operating conditions of the one or more cloud-native network functions. At 1730, the method includes using the observational data to apply a causal inference function to analyze the causal relationship between at least one observed cause of at least one observed effect of the one or more cloud-native network functions and at least one observed effect of the one or more cloud-native network functions. Method 1700 can be executed using CAROT 201, apparatus 1100, or apparatus 1500.
[0114] Figure 18 is an example method 1800 for implementing the example embodiments described herein. At 1810, the method includes operating replicas of one or more target applications. At 1820, the method includes generating observational data of the replicas of one or more target applications, the observational data being generated based on multiple operating conditions of the one or more target applications. At 1830, the method includes using the observational data to apply a causal inference function to analyze the causal relationship between at least one observed cause of at least one observed effect of the one or more target applications and at least one observed effect of the one or more target applications. Method 1800 can be executed using CAROT 201, apparatus 1100, or apparatus 1500.
[0115] Figure 19 is an example method 1900 for implementing the example embodiments described herein. At 1910, the method includes selecting one or more cloud-native network functions. At 1920, the method includes selecting at least one characteristic of the one or more cloud-native network functions. At 1930, the method includes selecting at least one environmental condition of the one or more cloud-native network functions. At 1940, the method includes selecting a load for the one or more cloud-native network functions, the load including: the intensity and duration of processing of the one or more cloud-native network functions under at least one environmental condition. At 1950, the method includes using the one or more cloud-native network functions to perform at least one experiment based on at least one characteristic, load, and at least one environmental condition of the one or more cloud-native network functions. At 1960, the method includes collecting observational data from the at least one experiment. At 1970, the method includes using the observational data to apply a causal inference function to analyze the causal relationship between at least one observed cause of at least one observed effect of the one or more cloud-native network functions and at least one observed effect of the one or more cloud-native network functions. Method 1900 can be executed using CAROT 201, apparatus 1100, or apparatus 1500.
[0116] The following examples are provided and described herein.
[0117] Example 1. An apparatus, comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to at least: operate copies of one or more cloud-native network functions; generate observation data of the copies of the one or more cloud-native network functions, the observation data being generated based on multiple operating conditions of the one or more cloud-native network functions; and use the observation data to apply a causal inference function to analyze a causal relationship between at least one observed cause of at least one observed effect of the one or more cloud-native network functions and the at least one observed effect of the one or more cloud-native network functions.
[0118] Example 2. The apparatus according to Example 1, wherein the apparatus is caused to: design one or more cloud-native network functions based on the causal relationship between at least one observed cause and at least one observed effect of the one or more cloud-native network functions analyzed.
[0119] Example 3. The apparatus according to Example 2, wherein the one or more cloud-native network functions are designed during or after a test phase of a continuous integration and continuous deployment pipeline and before releasing the one or more cloud-native network functions in a production environment, the release being a release of the continuous integration and continuous deployment pipeline.
[0120] Example 4. The apparatus according to any one of Examples 1 to 3, wherein the apparatus is caused to: operate or configure the one or more cloud-native network functions during or after a test phase of a continuous integration and continuous deployment pipeline and before releasing the one or more cloud-native network functions in a production environment, the release being a release of the continuous integration and continuous deployment pipeline.
[0121] Example 5. The apparatus according to any one of Examples 1 to 4, wherein the one or more cloud-native network functions include: software applications for a 5G network.
[0122] Example 6. The apparatus according to any one of Examples 1 to 5, wherein the copy includes: a digital twin of the one or more cloud-native network functions.
[0123] Example 7. The apparatus according to any one of Examples 1 to 6, wherein the at least one observed cause includes: operating attributes of the one or more cloud-native network functions, and the configuration of the one or more cloud-native network functions includes: types of the operating attributes.
[0124] Example 8. The apparatus according to any one of Examples 1 to 7, wherein the at least one observed effect includes: at least one performance metric of the one or more cloud-native network functions.
[0125] Example 9. The apparatus according to any one of Examples 1 to 8, wherein at least one observed effect includes at least one characteristic, load, or at least one environmental condition of one or more cloud-native network functions.
[0126] Example 10. The apparatus according to any one of Examples 1 to 9, wherein analyzing a causal relationship between at least one observed cause and at least one observed effect of one or more cloud-native network functions includes at least one of the following: determining the existence of a relationship between at least one observed cause and at least one observed effect; determining the magnitude of the relationship between at least one observed cause and at least one observed effect; or determining the manner in which at least one observed cause and at least one observed effect are related.
[0127] Example 11. The apparatus according to any one of Examples 1 to 10, wherein the plurality of operating conditions includes at least one operating condition of one or more cloud-native network functions, and at least one operating condition includes at least one of the following: network load; application programming interface load; domain name system corruption; kernel crash; complete crash; network stress; computing resource stress; time corruption; network service crash; or input / output corruption.
[0128] Example 12. The apparatus according to any one of Examples 1 to 11, wherein the apparatus is further configured to: form at least one graph representing a causal relationship between at least one observed cause and at least one observed effect of one or more cloud-native network functions.
[0129] Example 13. The apparatus according to Example 12, wherein the at least one graph includes: a directed acyclic graph.
[0130] Example 14. The apparatus according to any one of Examples 12 to 13, wherein the apparatus is further configured to: form at least one graph using domain expert knowledge to remove at least one ambiguity related to a causal relationship between at least one observed cause and at least one observed effect of one or more cloud-native network functions.
[0131] Example 15. The apparatus according to any one of Examples 1 to 14, wherein the apparatus is further configured to: generate observed data as a table, wherein the rows of the table include observed results, and the columns of the table include: cause attributes, effect attributes, or attributes including both cause and effect.
[0132] Example 16. The apparatus according to any one of Examples 1 to 15, wherein the apparatus is further configured to: generate observed data as a table, wherein the first column of the table includes at least a part of the operating configuration of one or more cloud-native network functions, and the second column of the table includes at least one observed effect of one or more cloud-native network functions.
[0133] Example 17. The apparatus according to any one of Examples 1 to 16, wherein the apparatus is further configured to: generate observational data as a table, where the table includes: an experimental group and a control group, where the experimental group includes: a set of experiments in which at least one operating attribute is set to a specific value to be studied, and where the control group includes: a set of experiments in which the operating attribute is randomized.
[0134] Example 18. The apparatus according to Example 17, wherein the apparatus is further configured to: determine an average treatment effect of one or more cloud-native network functions when at least one operating attribute is set to a specific value.
[0135] Example 19. The apparatus according to any one of Examples 1 to 18, wherein the apparatus is further configured to: determine the magnitude of the relationship between at least one observed cause and at least one observed effect; where the magnitude of the relationship includes at least one of the following: average treatment effect, conditional average treatment effect, individual treatment effect, natural direct effect, or natural indirect effect.
[0136] Example 20. The apparatus according to any one of Examples 1 to 19, wherein applying a causal inference function includes: performing do-calculus to replace the do-operator with at least one conditional probability of at least one observed effect of one or more cloud-native network functions given at least one observed cause, the at least one conditional probability being used to infer the causal relationship between at least one observed cause and at least one observed effect.
[0137] Example 21. The apparatus according to any one of Examples 1 to 20, wherein applying a causal inference function includes: performing a statistical analysis of at least one observed cause and at least one observed effect of one or more cloud-native network functions.
[0138] Example 22. The apparatus according to any one of Examples 1 to 21, wherein applying a causal inference function includes: applying machine learning to analyze the causal relationship between at least one observed cause and at least one observed effect of one or more cloud-native network functions.
[0139] Example 23. The apparatus according to any one of Examples 1 to 22, wherein the apparatus is further configured to: based on the application of the causal inference function, verify or exclude at least one design or configuration assumption related to at least one observed cause or at least one observed effect of one or more cloud-native network functions.
[0140] Example 24. The apparatus according to any one of Examples 1 to 23, wherein the apparatus is further configured to: based on the application of the causal inference function, determine that at least one operating condition of one or more cloud-native network functions is sub-optimal.
[0141] Example 25. The apparatus according to any one of Examples 1 to 24, wherein the apparatus is further caused to: perform a root cause analysis to determine at least one cause of an operational failure of one or more cloud-native network functions.
[0142] Example 26. The apparatus according to Example 25, wherein the apparatus is further caused to: operate or configure one or more cloud-native network functions using at least one result of the root cause analysis to supplement the application of the causal reasoning function.
[0143] Example 27. The apparatus according to any one of Examples 1 to 26, wherein the apparatus is further caused to: determine whether one or more cloud-native network functions are external applications or internal applications.
[0144] Example 28. The apparatus according to any one of Examples 1 to 27, wherein one or more cloud-native network functions are in the form of containerized software.
[0145] Example 29. An apparatus, comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to at least: operate copies of one or more target applications; generate observation data of the copies of the one or more target applications, the observation data being generated based on multiple operating conditions of the one or more target applications; and use the observation data to apply a causal reasoning function to analyze a causal relationship between at least one observed cause of at least one observed effect of the one or more target applications and the at least one observed effect of the one or more target applications.
[0146] Example 30. The apparatus according to Example 29, wherein the apparatus is caused to: design one or more target applications based on the causal relationship between at least one observed cause and at least one observed effect of the one or more target applications analyzed.
[0147] Example 31. The apparatus according to Example 30, wherein the one or more target applications are designed during or after a test phase of a continuous integration and continuous deployment pipeline and before releasing the one or more target applications in a production environment, the release being a release of the continuous integration and continuous deployment pipeline.
[0148] Example 32. The apparatus according to any one of Examples 29 to 31, wherein the apparatus is caused to: operate or configure one or more target applications during or after a test phase of a continuous integration and continuous deployment pipeline and before releasing the one or more target applications in a production environment, the release being a release of the continuous integration and continuous deployment pipeline.
[0149] Example 33. The apparatus according to any one of Examples 29 to 32, wherein the one or more target applications include: software applications for a 5G network.
[0150] Example 34. The apparatus according to any one of Examples 29 to 33, wherein the copy includes: digital twins of one or more target applications.
[0151] Example 35. The apparatus according to any one of Examples 29 to 34, wherein at least one observed cause includes: operational attributes of one or more target applications, and wherein the configuration of the one or more target applications includes: the type of operational attributes.
[0152] Example 36. The apparatus according to any one of Examples 29 to 35, wherein at least one observed effect includes: at least one performance metric of one or more target applications.
[0153] Example 37. The apparatus according to any one of Examples 29 to 36, wherein at least one observed effect includes: at least one feature, load, or at least one environmental condition of one or more target applications.
[0154] Example 38. The apparatus according to any one of Examples 29 to 37, wherein analyzing the causal relationship between at least one observed cause and at least one observed effect of one or more target applications includes at least one of the following: determining the existence of the relationship between at least one observed cause and at least one observed effect; determining the magnitude of the relationship between at least one observed cause and at least one observed effect; or determining the manner in which at least one observed cause and at least one observed effect are related.
[0155] Example 39. The apparatus according to any one of Examples 29 to 38, wherein the plurality of operating conditions includes: at least one operating condition of one or more target applications, and the at least one operating condition includes at least one of the following: network load; application programming interface load; domain name system corruption; kernel crash; complete crash; network stress; computing resource stress; time corruption; network service crash; or input / output corruption.
[0156] Example 40. The apparatus according to any one of Examples 29 to 39, wherein the apparatus is further caused to: form at least one graph representing the causal relationship between at least one observed cause and at least one observed effect of one or more target applications.
[0157] Example 41. The apparatus according to Example 40, wherein the at least one graph includes: a directed acyclic graph.
[0158] Example 42. The apparatus according to any one of Examples 40 to 41, wherein the apparatus is further caused to: form at least one graph using domain expert knowledge to remove at least one ambiguity related to the causal relationship between at least one observed cause and at least one observed effect of one or more target applications.
[0159] Example 43. The apparatus according to any one of Examples 29 to 42, wherein the apparatus is further configured to: generate the observed data as a table, wherein the rows of the table include the observation results, and the columns of the table include: a cause attribute, an effect attribute, or an attribute including both cause and effect.
[0160] Example 44. The apparatus according to any one of Examples 29 to 43, wherein the apparatus is further configured to: generate the observed data as a table, wherein the first column of the table includes: at least a part of the operation configurations of one or more target applications, and the second column of the table includes: at least one observed effect of one or more target applications.
[0161] Example 45. The apparatus according to any one of Examples 29 to 44, wherein the apparatus is further configured to: generate the observed data as a table, wherein the table includes: an experimental group and a control group, wherein the experimental group includes: a set of experiments in which at least one operation attribute is set to a specific value to be studied, and wherein the control group includes: a set of experiments in which the operation attribute is randomized.
[0162] Example 46. The apparatus according to Example 45, wherein the apparatus is further configured to: determine the average treatment effect of one or more target applications when at least one operation attribute is set to a specific value.
[0163] Example 47. The apparatus according to any one of Examples 29 to 46, wherein the apparatus is further configured to: determine the magnitude of the relationship between at least one observed cause and at least one observed effect; and wherein the magnitude of the relationship includes at least one of the following terms: average treatment effect, conditional average treatment effect, individual treatment effect, natural direct effect, or natural indirect effect.
[0164] Example 48. The apparatus according to any one of Examples 29 to 47, wherein applying the causal reasoning function includes: performing do-calculus to replace the do-operator with at least one conditional probability of at least one observed effect of one or more target applications given at least one observed cause, the at least one conditional probability being used to infer the causal relationship between at least one observed cause and at least one observed effect.
[0165] Example 49. The apparatus according to any one of Examples 29 to 48, wherein applying the causal reasoning function includes: performing a statistical analysis of at least one observed cause and at least one observed effect of one or more target applications.
[0166] Example 50. The apparatus according to any one of Examples 29 to 49, wherein applying the causal reasoning function includes: applying machine learning to analyze the causal relationship between at least one observed cause and at least one observed effect of one or more target applications.
[0167] Example 51. The apparatus according to any one of Examples 29 to 50, wherein the apparatus is further configured to: based on the application of the causal reasoning function, verify or exclude at least one design or configuration assumption related to at least one observed cause or at least one observed effect of one or more target applications.
[0168] Example 52. The apparatus according to any one of Examples 29 to 51, wherein the apparatus is further configured to: based on the application of the causal reasoning function, determine that at least one operating condition of one or more target applications is sub-optimal.
[0169] Example 53. The apparatus according to any one of Examples 29 to 52, wherein the apparatus is further configured to: perform a root cause analysis to determine at least one cause of an operating failure of one or more target applications.
[0170] Example 54. The apparatus according to Example 53, wherein the apparatus is further configured to: use at least one result of the root cause analysis to operate or configure one or more target applications to supplement the application of the causal reasoning function.
[0171] Example 55. The apparatus according to any one of Examples 29 to 54, wherein the apparatus is further configured to: determine whether one or more target applications are external applications or internal applications.
[0172] Example 56. The apparatus according to any one of Examples 29 to 55, wherein one or more target applications are in the form of containerized software.
[0173] Example 57. The apparatus according to any one of Examples 29 to 56, wherein one or more target applications include: cloud-native network functions.
[0174] Example 58. The apparatus according to Example 57, wherein the cloud-native network function includes: 5G cloud-native network functions.
[0175] Example 59. The apparatus according to any one of Examples 57 to 58, wherein the cloud-native network function is in the form of containerized software.
[0176] Example 60. An apparatus, comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to at least: select one or more cloud-native network functions; select at least one characteristic of the one or more cloud-native network functions; select at least one environmental condition of the one or more cloud-native network functions; select a load of the one or more cloud-native network functions, the load including: the intensity and duration of processing of the one or more cloud-native network functions under at least one environmental condition; utilize the one or more cloud-native network functions to perform at least one experiment based on at least one characteristic, load, and at least one environmental condition of the one or more cloud-native network functions; collect observational data from the at least one experiment; and use the observational data to apply a causal inference function to analyze the causal relationship between at least one observed cause of at least one observed effect of the one or more cloud-native network functions and the at least one observed effect of the one or more cloud-native network functions.
[0177] Example 61. The apparatus according to Example 60, wherein at least one observed cause includes: operational attributes of the one or more cloud-native network functions, and wherein the configuration of the one or more cloud-native network functions includes: the type of operational attributes.
[0178] Example 62. The apparatus according to any one of Examples 60 to 61, wherein at least one observed effect includes: at least one performance metric of the one or more cloud-native network functions.
[0179] Example 63. The apparatus according to any one of Examples 60 to 62, wherein at least one observed effect includes: at least one characteristic, load, or at least one environmental condition of the one or more cloud-native network functions.
[0180] Example 64. The apparatus according to any one of Examples 60 to 63, wherein the apparatus is further caused to: design the one or more cloud-native network functions based on the analyzed causal relationship between at least one observed cause and at least one observed effect of the one or more cloud-native network functions.
[0181] Example 65. The apparatus according to any one of Examples 60 to 64, wherein the apparatus is further caused to: determine whether the one or more cloud-native network functions are external applications or internal applications.
[0182] Example 66. The apparatus according to Example 65, wherein the apparatus is further caused to: provide an endpoint of the one or more cloud-native network functions in response to determining that the one or more cloud-native network functions are external applications.
[0183] Example 67. The apparatus according to Example 66, wherein the endpoint includes: a uniform resource locator.
[0184] Example 68. The apparatus according to any one of Examples 65 to 67, wherein the apparatus is further configured to: deploy one or more cloud-native network functions in response to determining that one or more cloud-native network functions are internal applications; wherein at least one experiment is performed using one or more cloud-native network functions.
[0185] Example 69. The apparatus according to Example 68, wherein the apparatus is further configured to: wrap one or more cloud-native network functions within a helm template during application of a causal inference function.
[0186] Example 70. The apparatus according to any one of Examples 60 to 69, wherein environmental conditions of one or more cloud-native network functions include at least one of the following: network load; application programming interface load; domain name system corruption; kernel crash; complete crash; network stress; computing resource stress; time corruption; network service crash; or input / output corruption.
[0187] Example 71. The apparatus according to any one of Examples 60 to 70, wherein the apparatus is further configured to: form at least one graph representing a causal relationship between at least one observed cause and at least one observed effect of one or more cloud-native network functions.
[0188] Example 72. The apparatus according to any one of Examples 60 to 71, wherein the apparatus is further configured to: determine a magnitude of a relationship between at least one observed cause and at least one observed effect; and wherein the magnitude of the relationship includes at least one of the following: average treatment effect, conditional average treatment effect, individual treatment effect, natural direct effect, or natural indirect effect.
[0189] Example 73. The apparatus according to any one of Examples 60 to 72, wherein one or more cloud-native network functions are in the form of containerized software.
[0190] Example 74. The apparatus according to any one of Examples 60 to 74, wherein one or more cloud-native network functions are fifth-generation (5G) cloud-native network functions.
[0191] Example 75. A method, comprising: operating replicas of one or more cloud-native network functions; generating observed data of the replicas of one or more cloud-native network functions, the observed data being generated based on multiple operating conditions of one or more cloud-native network functions; and using the observed data to apply a causal inference function to analyze a causal relationship between at least one observed cause and at least one observed effect of one or more cloud-native network functions.
[0192] Example 76. A method includes: operating a copy of one or more target applications; generating observation data of the copy of one or more target applications, the observation data being generated based on multiple operating conditions of one or more target applications; and using the observation data to apply a causal reasoning function to analyze a causal relationship between at least one observed cause of at least one observed effect of one or more target applications and at least one observed effect of one or more target applications.
[0193] Example 77. A method includes: selecting one or more cloud-native network functions; selecting at least one characteristic of one or more cloud-native network functions; selecting at least one environmental condition of one or more cloud-native network functions; selecting a load of one or more cloud-native network functions, the load including: the intensity and duration of processing of one or more cloud-native network functions under at least one environmental condition; using one or more cloud-native network functions to perform at least one experiment, the experiment being based on at least one characteristic, load, and at least one environmental condition of one or more cloud-native network functions; collecting observation data from at least one experiment; and using the observation data to apply a causal reasoning function to analyze a causal relationship between at least one observed cause of at least one observed effect of one or more cloud-native network functions and at least one observed effect of one or more cloud-native network functions.
[0194] Example 78. An apparatus includes: components for operating a copy of one or more cloud-native network functions; components for generating observation data of the copy of one or more cloud-native network functions, the observation data being generated based on multiple operating conditions of one or more cloud-native network functions; and components for using the observation data to apply a causal reasoning function to analyze a causal relationship between at least one observed cause of at least one observed effect of one or more cloud-native network functions and at least one observed effect of one or more cloud-native network functions.
[0195] Example 79. An apparatus includes: components for operating a copy of one or more target applications; components for generating observation data of the copy of one or more target applications, the observation data being generated based on multiple operating conditions of one or more target applications; and components for using the observation data to apply a causal reasoning function to analyze a causal relationship between at least one observed cause of at least one observed effect of one or more target applications and at least one observed effect of one or more target applications.
[0196] Example 80. An apparatus includes: components for selecting one or more cloud-native network functions; components for selecting at least one characteristic of one or more cloud-native network functions; components for selecting at least one environmental condition of one or more cloud-native network functions; components for selecting a load of one or more cloud-native network functions, the load including: the intensity and duration of processing of one or more cloud-native network functions under at least one environmental condition; components for performing at least one experiment using one or more cloud-native network functions, the experiment being based on at least one characteristic, the load, and at least one environmental condition of one or more cloud-native network functions; components for collecting observation data from at least one experiment; and components for applying a causal inference function using the observation data to analyze the causal relationship between at least one observed cause of at least one observed effect of one or more cloud-native network functions and at least one observed effect of one or more cloud-native network functions.
[0197] Example 81. A machine-readable non-transitory program storage device tangibly embodies a machine-executable instruction program for performing operations, the operations including: operating copies of one or more cloud-native network functions; generating observation data of the copies of one or more cloud-native network functions, the observation data being generated based on multiple operating conditions of one or more cloud-native network functions; and using the observation data to apply a causal inference function to analyze the causal relationship between at least one observed cause of at least one observed effect of one or more cloud-native network functions and at least one observed effect of one or more cloud-native network functions.
[0198] Example 82. A machine-readable non-transitory program storage device tangibly embodies a machine-executable instruction program for performing operations, the operations including: operating copies of one or more target applications; generating observation data of the copies of one or more target applications, the observation data being generated based on multiple operating conditions of one or more target applications; and using the observation data to apply a causal inference function to analyze the causal relationship between at least one observed cause of at least one observed effect of one or more target applications and at least one observed effect of one or more target applications.
[0199] Example 83. A machine-readable non-transitory program storage device tangibly embodies a machine-executable instruction program for performing operations, the operations including: selecting one or more cloud-native network functions; selecting at least one characteristic of one or more cloud-native network functions; selecting at least one environmental condition of one or more cloud-native network functions; selecting a load for one or more cloud-native network functions, the load including: the intensity and duration of processing of one or more cloud-native network functions under at least one environmental condition; utilizing one or more cloud-native network functions to perform at least one experiment based on at least one characteristic, load, and at least one environmental condition of one or more cloud-native network functions; collecting observational data from at least one experiment; and using the observational data to apply a causal reasoning function to analyze a causal relationship between at least one observed cause of at least one observed effect of one or more cloud-native network functions and at least one observed effect of one or more cloud-native network functions.
[0200] References to "computers", "processors", etc. should be understood to include not only computers having different architectures (such as single / multi-processor architectures and sequential or parallel architectures), but also dedicated circuits such as field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), signal processing devices, and other processing circuitry. References to computer programs, instructions, code, etc. should be understood to include software or firmware for programmable processors, such as, for example, the programmable content of a hardware device, whether instructions for a processor or configuration settings for a fixed-function device, gate array, or programmable logic device, etc.
[0201] The (multiple) memories described herein can be implemented using any suitable data storage technology (such as semiconductor-based memory devices, flash memory, magnetic memory devices and systems, optical memory devices and systems, non-transitory memory, transitory memory, fixed memory, and removable memory). The (multiple) memories can include a database for storing data.
[0202] As used herein, the term "circuitry" can refer to the following: (a) a hardware circuit implementation, such as an implementation in analog and / or digital circuitry, and (b) a combination of circuitry and software (and / or firmware), such as, where applicable: (i) a combination of (one or more) processors, or (ii) a portion of (one or more) processors / software (including (one or more) digital signal processors, software, and (one or more) memories that work together to enable a device to perform various functions), and (c) circuitry that requires software or firmware to operate (even if the software or firmware is not physically present), such as (one or more) microprocessors or a portion of (one or more) microprocessors. As a further example, as used herein, the term "circuitry" will also cover an implementation of only a processor (or processors) or a portion of a processor and its accompanying software and / or firmware. For example, if applicable to a particular element, the term "circuitry" will also cover a baseband integrated circuit or an application processor integrated circuit for a mobile phone, or a similar integrated circuit in a server, a cellular network device, or another network device.
[0203] In the figures, the arrows between individual boxes represent the operative couplings between them and the direction of the data flow on these couplings.
[0204] It should be understood that the above description is illustrative only. Those skilled in the art can design various alternatives and modifications. For example, the features recited in the respective dependent claims can be combined with each other in any suitable combination. Additionally, features from the different example embodiments above can be selectively combined into new example embodiments. Accordingly, this specification is intended to cover all such alternatives, modifications, and variations that fall within the scope of the appended claims.
[0205] The following abbreviations and acronyms that may appear in the specification and / or the drawings are defined as follows (the abbreviations and acronyms may be appended to each other or to other characters using, for example, dashes, hyphens, or numbers):
[0206] 3GPP Third Generation Partnership Project
[0207] 4G Fourth Generation
[0208] 5G Fifth Generation
[0209] 5GC 5G Core Network
[0210] AGI Artificial General Intelligence
[0211] AMF Access and Mobility Management Function
[0212] API Application Programming Interface
[0213] ASIC Application Specific Integrated Circuit
[0214] ATE Average Treatment Effect
[0215] ATT Average Treatment Effect on the Treated
[0216] CATE Conditional Average Treatment Effect
[0217] CAROT Operational Twin Causal Reasoning
[0218] CD Causal Discovery
[0219] CI Causal Inference
[0220] CI / CD Continuous Integration / Continuous Deployment (or Delivery)
[0221] CNF Cloud Native Network Function
[0222] Col Column
[0223] CPU Central Processing Unit
[0224] CSV Comma-Separated Values
[0225] CU Central Unit or Centralized Unit
[0226] DAG Directed Acyclic Graph
[0227] DNS Domain Name System / Service
[0228] DSP Digital Signal Processor
[0229] eNB Evolved Node B (e.g., an LTE base station)
[0230] EN-DC E-UTRAN New Radio - Dual Connectivity
[0231] en-gNB A node that provides NR user plane and control plane protocol termination towards the UE and acts as a secondary node in EN-DC
[0232] E-UTRA Evolved Universal Terrestrial Radio Access, i.e., LTE radio access technology
[0233] E-UTRAN E-UTRA network
[0234] evn. Environment
[0235] F1 Interface between the CU and the DU
[0236] FPGA Field Programmable Gate Array
[0237] gNB A base station for 5G / NR, i.e., a node that provides NR user plane and control plane protocol termination towards the UE and is connected to the 5GC via the NG interface
[0238] Helm or helm is a packaging manager for Kubernetes
[0239] HTTP Hypertext Transfer Protocol
[0240] id identifier
[0241] IDE Integrated Development Environment
[0242] I / F Interface
[0243] IMIX Internet Mix
[0244] incl. including
[0245] I / O Input / Output
[0246] ITE Individual Treatment Effect
[0247] K8 or K8s Kubernetes
[0248] KPI Key Performance Indicator
[0249] LMF Location Management Function
[0250] LTE Long-Term Evolution (4G)
[0251] MAC Media Access Control
[0252] ML Machine Learning
[0253] MME Mobility Management Entity
[0254] NCE Network Control Element
[0255] NDE Natural Direct Effect
[0256] NIE Natural Indirect Effect
[0257] ng or NG Next Generation
[0258] ng-eNB Next Generation eNB
[0259] NG-RAN Next Generation Radio Access Network
[0260] NR New Radio (5G)
[0261] NSA Non-Standalone, usually in the context of 5G core
[0262] N / W Network
[0263] Obs. Observation
[0264] PDA Personal Digital Assistant
[0265] PDCP Packet Data Convergence Protocol
[0266] PHY Physical Layer
[0267] RAM Random Access Memory
[0268] RAN Radio Access Network
[0269] RCA Root Cause Analysis
[0270] RCT Randomized Controlled Trial
[0271] REST Representational State Transfer
[0272] RLC Radio Link Control
[0273] ROM Read Only Memory
[0274] RRC Radio Resource Control (Protocol)
[0275] RU Radio Unit
[0276] Rx Receiver or Receive
[0277] SA Standalone, typically in the context of 5G Core
[0278] SGW Serving Gateway
[0279] SLA Service Level Agreement
[0280] SLO Service Level Objective
[0281] SMF Session Management Function
[0282] SON Self-Organizing / Optimizing Network
[0283] SOTA State of the Art
[0284] TRP Transmission and Reception Point
[0285] Tx Transmitter or Transmission
[0286] UE User Equipment (e.g., wireless device, typically a mobile device)
[0287] UPF User Plane Function
[0288] URL Uniform Resource Locator
[0289] vCPU Virtual Central Processing Unit
[0290] WS (Multiple) Web Services
[0291] Network interfaces between X2 RAN nodes and between RAN and the core network
[0292] Network interface between Xn NG-RAN nodes
Claims
1. A device, comprising: at least one processor; and at least one memory storing instructions which, when executed by the at least one processor, cause the device to at least: operate copies of one or more cloud-native network functions; generate observation data of the copies of the one or more cloud-native network functions, the observation data being generated based on multiple operating conditions of the one or more cloud-native network functions; and use the observation data to apply a causal inference function to analyze the causal relationship between at least one observed cause of at least one observed effect of the one or more cloud-native network functions and the at least one observed effect of the one or more cloud-native network functions.
2. The device according to claim 1, wherein the device is caused to: design the one or more cloud-native network functions based on the causal relationship between the at least one observed cause and the at least one observed effect of the one or more cloud-native network functions analyzed.
3. The device according to claim 2, wherein the one or more cloud-native network functions are designed during or after a test phase of a continuous integration and continuous deployment pipeline and before releasing the one or more cloud-native network functions in a production environment, the release being the release of the continuous integration and continuous deployment pipeline.
4. The device according to any one of claims 1 to 3, wherein the device is caused to: operate or configure the one or more cloud-native network functions during or after a test phase of a continuous integration and continuous deployment pipeline and before releasing the one or more cloud-native network functions in a production environment, the release being the release of the continuous integration and continuous deployment pipeline.
5. The device according to any one of claims 1 to 4, wherein the one or more cloud-native network functions comprise: software applications for 5G networks.
6. The device according to any one of claims 1 to 5, wherein the copy comprises: a digital twin of the one or more cloud-native network functions.
7. The device according to any one of claims 1 to 6, wherein the at least one observed cause comprises: operating attributes of the one or more cloud-native network functions, wherein the configuration of the one or more cloud-native network functions comprises: the type of operating attributes.
8. The device according to any one of claims 1 to 7, wherein the at least one observed effect comprises: at least one performance metric of the one or more cloud-native network functions.
9. The device according to any one of claims 1 to 8, wherein the at least one observed effect comprises: at least one characteristic, load, or at least one environmental condition of the one or more cloud-native network functions.
10. The device according to any one of claims 1 to 9, wherein analyzing the causal relationship between the at least one observed cause and the at least one observed effect of the one or more cloud-native network functions comprises at least one of the following: determining the existence of the relationship between the at least one observed cause and the at least one observed effect; Determine the magnitude of the relationship between the at least one observed cause and the at least one observed effect; or Determine the manner in which the at least one observed cause and the at least one observed effect are related.
11. The apparatus according to any one of claims 1 to 10, wherein the plurality of operating conditions comprise: at least one operating condition of the one or more cloud-native network functions, the at least one operating condition comprising at least one of the following items: Network load; Application programming interface load; Domain Name System corruption; Kernel crash; Complete crash; Network stress; Computing resource stress; Time corruption; Network service crash; or Input / output corruption.
12. The apparatus according to any one of claims 1 to 11, wherein the apparatus is further configured to: Form at least one graph representing the causal relationship between the at least one observed cause and the at least one observed effect of the one or more cloud-native network functions.
13. The apparatus according to claim 12, wherein the at least one graph comprises: A directed acyclic graph.
14. The apparatus according to any one of claims 12 to 13, wherein the apparatus is further configured to: Use domain expert knowledge to form the at least one graph to remove at least one ambiguity related to the causal relationship between the at least one observed cause and the at least one observed effect of the one or more cloud-native network functions.
15. The apparatus according to any one of claims 1 to 14, wherein the apparatus is further configured to: Generate the observed data as a table, wherein the rows of the table include the observation results, and the columns of the table comprise: Cause attributes, effect attributes, or attributes including both cause and effect.
16. The apparatus according to any one of claims 1 to 15, wherein the apparatus is further configured to: Generate the observed data as a table, wherein the first column of the table comprises: At least a part of the operating configuration of the one or more cloud-native network functions, and the second column of the table comprises: the at least one observed effect of the one or more cloud-native network functions.
17. The apparatus according to any one of claims 1 to 16, wherein the apparatus is further configured to: Generate the observed data as a table, wherein the table comprises: An experimental group and a control group, wherein the experimental group comprises: a set of experiments in which at least one operating attribute is set to a specific value to be studied, and wherein the control group comprises: a set of experiments in which the operating attributes are randomized.
18. The apparatus according to claim 17, wherein the apparatus is further configured to: Determine the average treatment effect of the one or more cloud-native network functions when the at least one operating attribute is set to the specific value.
19. The apparatus according to any one of claims 1 to 18, wherein the apparatus is further configured to: Determine the magnitude of the relationship between the at least one observed cause and the at least one observed effect; The magnitude of the relationship includes at least one of the following: average treatment effect, conditional average treatment effect, individual treatment effect, natural direct effect, or natural indirect effect.
20. The apparatus according to any one of claims 1 to 19, wherein applying the causal inference function comprises: Performing do-calculus to replace the do-operator with at least one conditional probability of at least one observed effect of the one or more cloud-native network functions given the at least one observed cause, the at least one conditional probability being used to infer the causal relationship between the at least one observed cause and the at least one observed effect.
21. The apparatus according to any one of claims 1 to 20, wherein applying the causal inference function comprises: Performing a statistical analysis of the at least one observed cause and the at least one observed effect of the one or more cloud-native network functions.
22. The apparatus according to any one of claims 1 to 21, wherein applying the causal inference function comprises: Applying machine learning to analyze the causal relationship between the at least one observed cause and the at least one observed effect of the one or more cloud-native network functions.
23. The apparatus according to any one of claims 1 to 22, wherein the apparatus is further configured to: Based on the application of the causal inference function, verify or exclude at least one design or configuration assumption related to the at least one observed cause or the at least one observed effect of the one or more cloud-native network functions.
24. The apparatus according to any one of claims 1 to 23, wherein the apparatus is further configured to: Based on the application of the causal inference function, determine that at least one operating condition of the one or more cloud-native network functions is sub-optimal.
25. The apparatus according to any one of claims 1 to 24, wherein the apparatus is further configured to: Perform root cause analysis to determine at least one cause of an operational failure of the one or more cloud-native network functions.
26. The apparatus according to claim 25, wherein the apparatus is further configured to: Use at least one result of the root cause analysis to operate or configure the one or more cloud-native network functions to supplement the application of the causal inference function.
27. The apparatus according to any one of claims 1 to 26, wherein the apparatus is further configured to: Determine whether the one or more cloud-native network functions are external applications or internal applications.
28. The apparatus according to any one of claims 1 to 27, wherein the one or more cloud-native network functions are in the form of containerized software.
29. An apparatus, comprising: At least one processor; and At least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to at least: Operate copies of one or more target applications; Generate observed data of the copies of the one or more target applications, the observed data being generated based on multiple operating conditions of the one or more target applications; and Use the observation data to apply a causal inference function to analyze the causal relationship between at least one observed cause of at least one observed effect of the one or more target applications and the at least one observed effect of the one or more target applications.
30. The apparatus according to claim 29, wherein the apparatus is configured to: Design the one or more target applications based on the causal relationship between the at least one observed cause and the at least one observed effect of the one or more target applications analyzed.
31. The apparatus according to claim 30, wherein the one or more target applications are designed during or after a test phase of a continuous integration and continuous deployment pipeline and before releasing the one or more target applications in a production environment, and the release is a release of the continuous integration and continuous deployment pipeline.
32. The apparatus according to any one of claims 29 to 31, wherein the apparatus is configured to: Operate or configure the one or more target applications during or after a test phase of a continuous integration and continuous deployment pipeline and before releasing the one or more target applications in a production environment, and the release is a release of the continuous integration and continuous deployment pipeline.
33. The apparatus according to any one of claims 29 to 32, wherein the one or more target applications comprise: Software applications for 5G networks.
34. The apparatus according to any one of claims 29 to 33, wherein the replica comprises: A digital twin of the one or more target applications.
35. The apparatus according to any one of claims 29 to 34, wherein the at least one observed cause comprises: Operating attributes of the one or more target applications, wherein the configuration of the one or more target applications includes: the type of operating attributes.
36. The apparatus according to any one of claims 29 to 35, wherein the at least one observed effect comprises: At least one performance metric of the one or more target applications.
37. The apparatus according to any one of claims 29 to 36, wherein the at least one observed effect comprises: At least one feature, load, or at least one environmental condition of the one or more target applications.
38. The apparatus according to any one of claims 29 to 37, wherein analyzing the causal relationship between the at least one observed cause and the at least one observed effect of the one or more target applications includes at least one of the following: Determining the existence of the relationship between the at least one observed cause and the at least one observed effect; Determining the magnitude of the relationship between the at least one observed cause and the at least one observed effect; or Determining the manner in which the at least one observed cause and the at least one observed effect are related.
39. The apparatus according to any one of claims 29 to 38, wherein the plurality of operating conditions comprise: At least one operating condition of the one or more target applications, and the at least one operating condition includes at least one of the following: Network load; Application programming interface load; Domain Name System corruption; Kernel crash; Complete crash; Network pressure; Computing resource pressure; Time corruption; Network service crash; or Input / output corruption.
40. The apparatus according to any one of claims 29 to 39, wherein the apparatus is further caused to: Form at least one graph, the at least one graph representing: the causal relationship between the at least one observed cause and the at least one observed effect of the one or more target applications.
41. The apparatus according to claim 40, wherein the at least one graph comprises: A directed acyclic graph.
42. The apparatus according to any one of claims 40 to 41, wherein the apparatus is further caused to: Use domain expert knowledge to form the at least one graph to remove at least one ambiguity related to the causal relationship between the at least one observed cause and the at least one observed effect of the one or more target applications.
43. The apparatus according to any one of claims 29 to 42, wherein the apparatus is further caused to: Generate the observed data as a table, wherein the rows of the table include observation results, and the columns of the table comprise: Cause attributes, effect attributes, or attributes including both cause and effect.
44. The apparatus according to any one of claims 29 to 43, wherein the apparatus is further caused to: Generate the observed data as a table, wherein the first column of the table comprises: At least a part of the operation configuration of the one or more target applications, and the second column of the table comprises: the at least one observed effect of the one or more target applications.
45. The apparatus according to any one of claims 29 to 44, wherein the apparatus is further caused to: Generate the observed data as a table, wherein the table comprises: An experimental group and a control group, wherein the experimental group comprises: a set of experiments in which at least one operation attribute is set to a specific value to be studied, and wherein the control group comprises: a set of experiments in which the operation attributes are randomized.
46. The apparatus according to claim 45, wherein the apparatus is further caused to: Determine the average treatment effect of the one or more target applications when the at least one operation attribute is set to the specific value.
47. The apparatus according to any one of claims 29 to 46, wherein the apparatus is further caused to: Determine the magnitude of the relationship between the at least one observed cause and the at least one observed effect; and wherein the magnitude of the relationship comprises at least one of the following terms: average treatment effect, conditional average treatment effect, individual treatment effect, natural direct effect, or natural indirect effect.
48. The apparatus according to any one of claims 29 to 47, wherein applying the causal reasoning function comprises: Performing do-calculus to replace the do operator with at least one conditional probability of the at least one observed effect of the one or more target applications given the at least one observed cause, the at least one conditional probability being used to infer the causal relationship between the at least one observed cause and the at least one observed effect.
49. The apparatus according to any one of claims 29 to 48, wherein applying the causal reasoning function comprises: performing a statistical analysis of the at least one observed cause and the at least one observed effect of the one or more target applications.
50. The apparatus according to any one of claims 29 to 49, wherein applying the causal reasoning function comprises: applying machine learning to analyze the causal relationship between the at least one observed cause and the at least one observed effect of the one or more target applications.
51. The apparatus according to any one of claims 29 to 50, wherein the apparatus is further configured to: verify or eliminate at least one design or configuration assumption related to the at least one observed cause or the at least one observed effect of the one or more target applications based on the application of the causal reasoning function.
52. The apparatus according to any one of claims 29 to 51, wherein the apparatus is further configured to: determine that at least one operating condition of the one or more target applications is sub-optimal based on the application of the causal reasoning function.
53. The apparatus according to any one of claims 29 to 52, wherein the apparatus is further configured to: perform a root cause analysis to determine at least one cause of an operational failure of the one or more target applications.
54. The apparatus according to claim 53, wherein the apparatus is further configured to: operate or configure the one or more target applications using at least one result of the root cause analysis to complement the application of the causal reasoning function.
55. The apparatus according to any one of claims 29 to 54, wherein the apparatus is further configured to: determine whether the one or more target applications are external applications or internal applications.
56. The apparatus according to any one of claims 29 to 55, wherein the one or more target applications are in the form of containerized software.
57. The apparatus according to any one of claims 29 to 56, wherein the one or more target applications comprise: cloud-native network functions.
58. The apparatus according to claim 57, wherein the cloud-native network functions comprise: 5G cloud-native network functions.
59. The apparatus according to any one of claims 57 to 58, wherein the cloud-native network functions are in the form of containerized software.
60. An apparatus, comprising: at least one processor; and at least one memory storing instructions which, when executed by the at least one processor, cause the apparatus to at least: select one or more cloud-native network functions; select at least one characteristic of the one or more cloud-native network functions; select at least one environmental condition of the one or more cloud-native network functions; select a load of the one or more cloud-native network functions, the load comprising: the intensity and duration of processing of the one or more cloud-native network functions under the at least one environmental condition; Perform at least one experiment using the one or more cloud-native network functions, the experiment being based on the at least one characteristic of the one or more cloud-native network functions, the load, and the at least one environmental condition; Collect observational data from the at least one experiment; and Use the observational data to apply a causal reasoning function to analyze the causal relationship between at least one observed cause of at least one observed effect of the one or more cloud-native network functions and the at least one observed effect of the one or more cloud-native network functions.
61. The apparatus according to claim 60, wherein the at least one observed cause comprises: The operational attributes of the one or more cloud-native network functions, wherein the configuration of the one or more cloud-native network functions comprises: the type of operational attributes.
62. The apparatus according to any one of claims 60 to 61, wherein the at least one observed effect comprises: At least one performance metric of the one or more cloud-native network functions.
63. The apparatus according to any one of claims 60 to 62, wherein the at least one observed effect comprises: The at least one characteristic, the load, or the at least one environmental condition of the one or more cloud-native network functions.
64. The apparatus according to any one of claims 60 to 63, wherein the apparatus is further caused to: Design the one or more cloud-native network functions based on the causal relationship between the at least one observed cause and the at least one observed effect of the one or more cloud-native network functions analyzed.
65. The apparatus according to any one of claims 60 to 64, wherein the apparatus is further caused to: Determine whether the one or more cloud-native network functions are external applications or internal applications.
66. The apparatus according to claim 65, wherein the apparatus is further caused to: In response to determining that the one or more cloud-native network functions are external applications, provide endpoints of the one or more cloud-native network functions.
67. The apparatus according to claim 66, wherein the endpoint comprises: Uniform Resource Locator.
68. The apparatus according to any one of claims 65 to 67, wherein the apparatus is further caused to: In response to determining that the one or more cloud-native network functions are internal applications, deploy the one or more cloud-native network functions; wherein the at least one experiment is performed using the one or more cloud-native network functions.
69. The apparatus according to claim 68, wherein the apparatus is further caused to: During the application of the causal reasoning function, wrap the one or more cloud-native network functions in a helm template.
70. The apparatus according to any one of claims 60 to 69, wherein the environmental condition of the one or more cloud-native network functions comprises at least one of the following: Network load; Application Programming Interface load; Domain Name System corruption; Kernel crash; Total crash; Network stress; Computing resource stress; Time corruption; Network service crash; or Input / Output corruption.
71. The apparatus according to any one of claims 60 to 70, wherein the apparatus is further caused to: Form at least one graph, the at least one graph representing: the causal relationship between the at least one observed cause and the at least one observed effect of the one or more cloud-native network functions.
72. The apparatus according to any one of claims 60 to 71, wherein the apparatus is further caused to: Determine the magnitude of the relationship between the at least one observed cause and the at least one observed effect; and Wherein the magnitude of the relationship includes at least one of the following: average treatment effect, conditional average treatment effect, individual treatment effect, natural direct effect, or natural indirect effect.
73. The apparatus according to any one of claims 60 to 72, wherein the one or more cloud-native network functions are in the form of containerized software.
74. The apparatus according to any one of claims 60 to 74, wherein the one or more cloud-native network functions are fifth-generation (5G) cloud-native network functions.
75. A method, Comprising: Operating replicas of one or more cloud-native network functions; Generating observed data of the replicas of the one or more cloud-native network functions, the observed data being generated based on multiple operating conditions of the one or more cloud-native network functions; And Using the observed data to apply a causal inference function to analyze the causal relationship between at least one observed cause and at least one observed effect of the one or more cloud-native network functions.
76. A method, Comprising: Operating replicas of one or more target applications; Generating observed data of the replicas of the one or more target applications, the observed data being generated based on multiple operating conditions of the one or more target applications; And Using the observed data to apply a causal inference function to analyze the causal relationship between at least one observed cause and at least one observed effect of the one or more target applications.
77. A method, Comprising: Selecting one or more cloud-native network functions; Selecting at least one feature of the one or more cloud-native network functions; Selecting at least one environmental condition of the one or more cloud-native network functions; Selecting the load of the one or more cloud-native network functions, the load including: the intensity and duration of processing of the one or more cloud-native network functions under the at least one environmental condition; Performing at least one experiment using the one or more cloud-native network functions, the experiment being based on the at least one feature, the load, and the at least one environmental condition of the one or more cloud-native network functions; Collecting observed data from the at least one experiment; and Using the observed data to apply a causal inference function to analyze the causal relationship between at least one observed cause and at least one observed effect of the one or more cloud-native network functions.
78. A device comprising: means for operating replicas of one or more cloud-native network functions; means for generating observation data of the replicas of the one or more cloud-native network functions, the observation data being generated based on multiple operating conditions of the one or more cloud-native network functions; and means for using the observation data to apply a causal reasoning function to analyze a causal relationship between at least one observed cause of at least one observed effect of the one or more cloud-native network functions and the at least one observed effect of the one or more cloud-native network functions.
79. A device comprising: means for operating replicas of one or more target applications; means for generating observation data of the replicas of the one or more target applications, the observation data being generated based on multiple operating conditions of the one or more target applications; and means for using the observation data to apply a causal reasoning function to analyze a causal relationship between at least one observed cause of at least one observed effect of the one or more target applications and the at least one observed effect of the one or more target applications.
80. A device comprising: means for selecting one or more cloud-native network functions; means for selecting at least one characteristic of the one or more cloud-native network functions; means for selecting at least one environmental condition of the one or more cloud-native network functions; means for selecting a load of the one or more cloud-native network functions, the load including: the intensity and duration of processing of the one or more cloud-native network functions under the at least one environmental condition; means for performing at least one experiment using the one or more cloud-native network functions, the experiment being based on the at least one characteristic, the load, and the at least one environmental condition of the one or more cloud-native network functions; means for collecting observation data from the at least one experiment; and means for using the observation data to apply a causal reasoning function to analyze a causal relationship between at least one observed cause of at least one observed effect of the one or more cloud-native network functions and the at least one observed effect of the one or more cloud-native network functions.
81. A machine-readable non-transitory program storage device tangibly embodying a program of machine-executable instructions for performing operations that comprise: operating replicas of one or more cloud-native network functions; generating observation data of the replicas of the one or more cloud-native network functions, the observation data being generated based on multiple operating conditions of the one or more cloud-native network functions; and using the observation data to apply a causal reasoning function to analyze a causal relationship between at least one observed cause of at least one observed effect of the one or more cloud-native network functions and the at least one observed effect of the one or more cloud-native network functions.
82. A machine-readable non-transitory program storage device tangibly embodying a program of machine-executable instructions for performing operations that Comprising: Operating copies of one or more target applications; Generating observation data of the copies of the one or more target applications, the observation data being generated based on multiple operating conditions of the one or more target applications; And Using the observation data to apply a causal reasoning function to analyze the causal relationship between at least one observed cause of at least one observed effect of the one or more target applications and the at least one observed effect of the one or more target applications.
83. A machine-readable non-transitory program storage device tangibly implementing a machine-executable instruction program for performing operations, the operations Comprising: Selecting one or more cloud-native network functions; Selecting at least one characteristic of the one or more cloud-native network functions; Selecting at least one environmental condition of the one or more cloud-native network functions; Selecting the load of the one or more cloud-native network functions, the load including: the intensity and duration of processing of the one or more cloud-native network functions under the at least one environmental condition; Utilizing the one or more cloud-native network functions to perform at least one experiment based on the at least one characteristic, the load, and the at least one environmental condition of the one or more cloud-native network functions; Collecting observation data from the at least one experiment; and Using the observation data to apply a causal reasoning function to analyze the causal relationship between at least one observed cause of at least one observed effect of the one or more cloud-native network functions and the at least one observed effect of the one or more cloud-native network functions.