Auto-healing control in consideration of silent failure
The network management device addresses silent failures in virtualized environments by deploying alternative resources and scaling to minimize performance degradation and service disruptions.
Patent Information
- Application Number
- JP2024039947
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-14
- Publication Date
- 2025-09-29
AI Technical Summary
Silent failures in network virtualization environments, which are not detected by normal operational monitoring mechanisms, can lead to performance degradation and potential large-scale failures, causing long-term disadvantages and reputational risks for users and service providers.
A network management device and method that identifies performance degradation in VNFs or CNFs, triggering auto-healing by deploying alternative hardware resources and, if necessary, auto-scaling to minimize adverse effects on services.
Automatically minimizes the impact of silent failures by quickly identifying and addressing performance degradation through auto-healing and auto-scaling, preventing further service disruptions.
Smart Images

Figure 2025140508000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to auto-healing control that takes silent faults into account. [Background technology]
[0002] With the improvement in performance of general-purpose servers and the expansion of network infrastructure, cloud computing (hereafter referred to as "cloud"), which uses virtualized computing resources on physical resources such as servers on demand, has become widespread. NFV (Network Function Virtualization), which virtualizes network functions and provides them on the cloud, is also well known. NFV is a technology that uses virtualization and cloud technologies to separate the hardware and software of various network services that previously ran on dedicated hardware, and runs the software on a virtualized platform. This is expected to lead to more advanced operations and cost reductions. In recent years, virtualization has also been progressing in mobile networks. The European Telecommunications Standards Institute (ETSI) NFV defines the architecture of NFV (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] International Publication No. 2016 / 121830 Summary of the Invention [Problem to be solved by the invention]
[0004] Auto-healing, as defined in ETSI NFV and other standards, enables rapid service recovery by automatically evacuating the faulty VNF / CNF to healthy hardware resources when a failure is detected. Auto-healing removes the faulty VNF or CNF, deploys other hardware resources for that VNF or CNF (creates a VNF or CNF), and then configures the VNF or CNF to optimize its operation.
[0005] Generally, failures that trigger auto-healing can be broadly classified into two types: one is a failure of the hardware resources on which the VNF / CNF runs (such as a physical breakdown), and the other is a failure of the VNF / CNF itself (such as a software bug in the application).
[0006] On the other hand, there are also silent failures that do not fall into the above categories. In a network virtualization environment, a performance degradation that is not detected by normal operational monitoring mechanisms is called a "silent failure." If a silent failure is left unattended, it can eventually affect the entire system, causing an inability to connect to the network, and can lead to an even larger-scale failure. The cause of a silent failure is presumed to be a physical malfunction of the hardware or a malfunction of software such as firmware that runs on the hardware. However, identifying the location and cause of silent failures often requires a lot of time and effort, which means users tend to suffer long-term disadvantages due to performance degradation, and service providers may suffer long-term reputational risks.
[0007] As mentioned above, silent failures are not detected by normal operational monitoring mechanisms, so auto-healing will not be performed unless the silent failure causes a failure that triggers auto-healing.
[0008] On the other hand, in NFV-based mobile networks, if there is no alert report indicating an abnormality (which could trigger auto-healing) and a phenomenon such as performance degradation of a component occurs, auto-scaling or manual performance expansion may be implemented. Autoscaling is a process in a network virtualization environment that distributes the processing load by automatically adding new components (VNFs or CNFs) when the processing load of a component exceeds its tolerance. Human performance expansion is a process in which the processing load is distributed by manually adding new components (VNFs or CNFs). Whether through autoscaling or manual scaling, the added components may perform well, but the degraded components continue to be used, which can lead to silent failures that continue to negatively impact service to users of those components and, in some cases, can lead to larger-scale outages.
[0009] Therefore, the present disclosure provides an automation technology that can minimize adverse effects on services as quickly as possible when a degradation in performance of a component that could result in a silent failure occurs. [Means for solving the problem]
[0010] One aspect of the present disclosure provides a network management device, including at least one processor, that performs the following operations: receiving a signal that is provided when a degradation in performance of a VNF or CNF exceeding a threshold is detected in a virtualized environment of a network; identifying a hardware resource associated with the degradation in performance upon receiving the signal; and transmitting a first command transmission command to perform autohealing, such as deploying a hardware resource other than the identified hardware resource for the VNF or CNF whose performance has degraded.
[0011] One aspect of the present disclosure provides a network management method, including receiving a signal in a virtualized environment of a network when a degradation in performance of a VNF or CNF exceeding a threshold is detected, and upon receiving the signal, identifying a hardware resource associated with the degradation in performance and transmitting an instruction to perform autohealing such that a hardware resource other than the identified hardware resource is deployed for the VNF or CNF experiencing the degradation. [Effects of the Invention]
[0012] In an aspect of the present disclosure, when a degradation in the performance of a component that may result in a silent failure occurs in a network implemented in a virtualized environment, the adverse effect on services can be automatically minimized as quickly as possible. [Brief explanation of the drawings]
[0013] [Figure 1] Figure 1 is a block diagram illustrating components in a network virtualization environment defined by ETSI NFV. [Figure 2] FIG. 2 is a diagram illustrating an example of installation of a determination device according to the present disclosure. [Figure 3] FIG. 3 is a diagram showing another example of installation of a determination device according to the present disclosure. [Figure 4] FIG. 4 is a diagram showing another example of installation of a determination device according to the present disclosure. [Figure 5] FIG. 5 is a diagram showing another example of installation of a determination device according to the present disclosure. [Figure 6] FIG. 6 is a block diagram illustrating an example of a hardware configuration of the determination device. [Figure 7] FIG. 7 is a table showing examples of monitoring items for the performance of a VNF or CNF according to the present disclosure. [Figure 8] FIG. 8 is a sequence diagram illustrating an example of processing in a virtualized network environment according to the present disclosure. [Figure 9] FIG. 9 is a flowchart showing the determination operation of the determination device. [Figure 10] FIG. 10 is a diagram for explaining the details of the determination operation of the determination device. DETAILED DESCRIPTION OF THE INVENTION
[0014] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings.
[0015] Figure 1 is a block diagram illustrating components in a network virtualization environment defined by ETSI NFV. The solid lines in the diagram represent logical connections between components.
[0016] A VNF (Virtual Network Function) corresponds to applications that run on a virtual machine (VM) on a server, and realizes network functions such as directory services, routers, firewalls, and load balancers in software. VNFs may also be implemented as software (virtual machines) to implement elements of the EPC (Evolved Packet Core), which is the core network of a mobile network, or elements of the IMS (IP Multimedia Subsystem). CNF (Containerized Network Function, or Cloud-native Network Function) is an evolution of VNF, and supports applications running in containers on servers, realizing network functions in software. As a CNF, elements of EPC or IMS may be realized in software (containers). Hereinafter, "VNF / CNF" means "VNF" or "CNF".
[0017] An EMS (Element Management System) is a management function attached to each VNF / CNF. Each EMS is connected to the corresponding VNF / CNF and monitors that VNF / CNF.
[0018] NFVI (Network Function Virtualization Infrastructure) is the execution platform for VNF / CNF. NFVI is a platform that enables the flexible handling of hardware resources of physical machines (servers), such as computing, storage, and network functions, as virtualized hardware resources such as virtualized computing, virtualized storage, and virtualized networks, which are virtualized using a virtualization layer such as a hypervisor. In reality, multiple NFVIs are provided, and each NFVI is connected to multiple VNFs / CNFs and monitors those VNFs / CNFs.
[0019] The VIM (Virtualized Infrastructure Manager) plays the role of a cloud controller. In other words, the VIM controls the NFVI via the virtualization layer (managing computing, storage, and network resources, monitoring NFVI faults, and monitoring resource information, which is the execution platform for NFV). In practice, multiple VIMs are provided, and each VIM is connected to and monitors multiple NFVIs.
[0020] The VNFM (VNF Manager) manages VNFs / CNFs. Specifically, the VNFM controls NFVIs via the virtualization layer (managing computing, storage, and network resources, monitoring NFVI faults, monitoring resource information, etc.). In practice, multiple VNFMs are provided, and each VNFM is connected to multiple VNFs / CNFs and multiple VIMs and monitors those VNFs / CNFs and VIMs.
[0021] The NFVO (NFV Orchestrator) orchestrates NFVI resources, manages network resources, and manages network services. The NFVO also has a repository of NFV instances and a repository of NFVI resources. The NFVO is connected to multiple VNFMs and multiple VIMs and monitors those VNFMs and VIMs. In addition, when the NFVO receives a report indicating that a component in the virtualized environment has failed, it has the function of sending an auto-healing command to a component related to auto-healing in order to restore the failed VNF / CNF. Furthermore, when the NFVO receives a report indicating that the processing load of a component in the virtualized environment exceeds the allowable value, the NFVO has the function of sending an auto-scaling command to a component related to auto-scaling in order to automatically add a new VNF / CNF.
[0022] The NFVO, VNFM, and VIM constitute the Management and Orchestration (MANO), which has the management and orchestration functions for the virtualized environment.
[0023] OSS (Operations Support System) is a system (equipment, software, mechanisms, etc.) required by telecommunications carriers to build and operate services. BSS (Business Support System) is an information system (equipment, software, mechanisms, etc.) used by telecommunications carriers for charging, billing, customer support, etc. The OSS and BSS function in tandem and may be established as an inseparable unit. Hereinafter, "OSS / BSS" refers to a system that includes both the OSS and BSS, but the OSS and BSS may also be established separately. The OSS / BSS is connected to the NFVO and also to multiple EMSs.
[0024] In a virtualized environment, some of the functions of the NFVO (including the control of auto-healing and auto-scaling) may be performed by the OSS. When the OSS performs the functions related to receiving reports of auto-healing and auto-scaling failures and issuing commands for auto-healing and auto-scaling (control of auto-healing and auto-scaling), the connection corresponding to the solid line between the OSS / BSS and the NFVO in the figure is used to receive the reports and send the commands. In other words, auto-healing and auto-scaling are controlled by the OSS or NFVO. Hereinafter, "OSS / NFVO" refers to either the OSS or the NFVO.
[0025] As mentioned above, the VNF / CNF may function as an element of the EPC or an element of the IMS for mobile services. Therefore, there exists a network for applications (VNF / CNF) to provide services.,The thick solid line in Fig. 7 indicates the mobile service,network used for mobile services in the virtualized environment. The mobile service network connects the mobile network Mo and multiple VNFs / CNFs. The VNFs / CNFs communicate with the mobile network Mo via the mobile service network and communicate with each other via the mobile service network.
[0026] To implement auto-healing and auto-scaling, multiple logical networks are used in virtualized environments. First, a platform monitoring network between the VIM and NFVI is used. Each VIM is connected to multiple NFVIs, and the NFVIs report failures or performance degradation of the hardware resources on which the VNFs / CNFs run to the VIM. In other words, each VIM monitors the NFVIs under its control using the platform monitoring network between the VIM and NFVI. In addition, a platform monitoring network between the OSS / NFVO-VIM is used. The NFVO is connected to multiple VIMs, and each VIM reports hardware resource failures or performance degradation reported by the NFVI to the NFVO. When the OSS controls auto-healing and auto-scaling, the NFVO forwards the failure or performance degradation report to the OSS via the connection between the OSS and NFVO. In this way, the platform monitoring network is used to report failures or degradation of the hardware resources on which the VNFs / CNFs run.
[0027] Auto-healing and auto-scaling commands for hardware resources are transmitted via the platform control network between the OSS / NFVO and VIM. When the OSS controls auto-healing and auto-scaling, it sends commands to the NFVO via the OSS-NFVO connection. The NFVO then sends commands to the VIM corresponding to the VNF / CNF that has a failure or performance degradation. The autohealing command for hardware resources is a command to remove a faulty VNF / CNF and deploy other hardware resources for that VNF / CNF (create a VNF / CNF).The autoscaling command for hardware resources is a command to add a new VNF / CNF and deploy hardware resources for that VNF / CNF.
[0028] On the other hand, networks for application monitoring between EMS / VNFM and VNF / CNF (network between VNFM and VNF / CNF and connection between EMS and VNF / CNF) are used. Each VNFM is connected to multiple VNF / CNF and monitors the applications of the VNF / CNFs under it. Each EMS is connected to the corresponding VNF / CNF and monitors the applications of that VNF / CNF. In addition, networks for application monitoring between OSS / NFVO-EMS / VNFM (a network between NFVO-VNFM, a connection between OSS-EMS, and a connection between OSS and NFVO) are used. The NFVO is connected to multiple VNFMs, and each VNFM reports failures or performance degradation of VNF / CNF applications to the NFVO. When the OSS controls auto-healing and auto-scaling, the NFVO forwards the failure or performance degradation report to the OSS via the connection between OSS and NFVO. The OSS / BSS is connected to multiple EMSs, and each EMS may report failures of VNF / CNF applications to the OSS / BSS. When the OSS does not control auto-healing and auto-scaling, the OSS forwards the failure or performance degradation report to the NFVO via the connection between OSS and NFVO. In this way, the application monitoring network is used to report failures or performance degradation of applications in the VNF / CNF.
[0029] Auto-healing commands and auto-scaling commands for applications are transmitted via the application control network between the OSS / NFVO-EMS / VNFM (the network between the NFVO and VNFM, the connection between the OSS and EMS, and the connection between the OSS and NFVO). When the OSS controls auto-healing and auto-scaling, the OSS provides commands to the NFVO, and the NFVO forwards the commands to the VNFM corresponding to the created or added VNF / CNF, or provides the commands to the EMS corresponding to the created or added VNF / CNF. When the OSS does not control auto-healing and auto-scaling, the NFVO provides commands to the OSS, and the OSS forwards the commands to the EMS corresponding to the created or added VNF / CNF, or provides the commands to the VNFM corresponding to the created or added VNF / CNF. Auto-healing instructions for applications are configuration instructions to optimize the behavior of created VNFs / CNFs (i.e., to incorporate VNFs / CNFs into operation). Auto-scaling instructions for applications are configuration instructions to optimize the behavior of added VNFs / CNFs (i.e., to incorporate VNFs / CNFs into operation).
[0030] As described above, in a virtualized environment, auto-healing can be performed when there is a failure in the VNF / CNF hardware resources or applications. Also, auto-scaling can be performed when the performance of the VNF / CNF hardware resources or applications deteriorates. Degraded performance of VNFs / CNFs that may cause silent failures usually does not result in auto-healing, but may result in auto-scaling. VNFs / CNFs added through auto-scaling may perform well. However, degraded VNFs / CNFs may continue to be used, potentially causing silent failures to persist.
[0031] Therefore, in this embodiment, when a degradation in performance of a VNF / CNF occurs that could result in a silent failure, the adverse impact on services is automatically minimized as quickly as possible. Specifically, in this embodiment, in a network virtualization environment, a reception process is performed to receive a signal that is provided when a degradation in performance of a VNF / CNF that exceeds a threshold is detected. Upon receiving the signal, a determination process is performed to identify hardware resources associated with the performance degradation. A first command transmission process is performed to transmit a command to perform autohealing so as to deploy hardware resources other than the identified hardware resources for the VNF or CNF whose performance has degraded. By excluding the hardware resources associated with the performance degradation and performing autohealing, the adverse impact on services can be automatically minimized quickly. However, when a signal is received in the reception process, if a predetermined condition is met, a second command transmission process is executed to transmit a command to implement auto-scaling.
[0032] In this embodiment, a determination device is provided that determines whether to perform auto-healing or auto-scaling. Figures 2 to 5 show examples of the installation of the determination device. In the example of Figure 2, the decision device is provided in the OSS / BSS, which is preferable when the OSS controls auto-healing and auto-scaling. In the example of Figure 3, the decision device is located in the NVFO, which is preferable when the OSS does not control autohealing and autoscaling. In the example of Figure 4, the decision device is connected to the OSS / BSS, which is preferable when the OSS controls auto-healing and auto-scaling. In the example of Figure 5, the decision device is connected to the NVFO, which is preferable if the OSS does not control autohealing and autoscaling.
[0033] As shown in FIG. 6, the determination device includes a CPU (Central Processing Unit), ie, a processor, a ROM (Read Only Memory), a RAM (Random Access Memory), a HDD (Hard Disk Drive), and a UI (User Interface). The ROM or HDD stores computer programs required for the operation of the determination device, and also stores data such as parameters required for the operation of the determination device. The CPU uses data stored in the ROM or HDD to execute a computer program stored in the ROM or HDD, and operates in accordance with the computer program.
[0034] RAM is used as a work area for the CPU. The UI may be a combination of a display device and a pointing device (e.g., a mouse or touchpad), or a touch panel that functions as both a display device and a pointing device. Using the UI, the user of the determination device can give instructions to the CPU. In the installation examples of Figures 4 and 5, the decision device has a communication interface (not shown) for communicating with the OSS or NVFO.
[0035] 7 is a table showing examples of monitoring items for the performance of VNF / CNF according to this embodiment. As described above, VNF / CNF can function as an element of the EPC or an element of the IMS for a mobile service, and the monitoring items mainly relate to performance related to the mobile service.
[0036] Performance degradation of the platform layer monitoring items in Figure 7 is reported by the platform monitoring network described above. Therefore, data corresponding to the platform layer monitoring items (CPU usage rate to job or task completion rate) and thresholds (T1 to T7) in Figure 7 is stored in each VIM. The platform layer monitoring items may be monitored for each VNF / CNF, or may be monitored for each process running on the NFVI, which is the platform for VNFs / CNFs. The degradation of performance of the application layer monitoring items in Figure 7 is reported by the application monitoring network described above. Therefore, data corresponding to the application layer monitoring items (CPU usage to CPS) and thresholds (T8 to T14) in Figure 7 is stored in each VNFM and each EMS. The application layer monitoring items may be monitored for each VNF / CNF, or may be monitored for each process running in the VNF / CNF application.
[0037] Disk IO / s is the number of reads and writes to the disk per unit time. Network IO / s is the number of transmissions and receptions to the mobile network per unit time. Number of accommodated users is the number of UEs (User Equipment) to which the VNF / CNF provides mobile services. CPS is the CPS (Calls Per Second) of the mobile services provided by the VNF / CNF.
[0038] Each VIM compares each monitoring item in the platform layer in Figure 7 with the corresponding threshold (one of T1 to T7). If CPU usage exceeds threshold T1, it is estimated that the VNF / CNF's performance has degraded. If memory usage exceeds threshold T2, it is estimated that the VNF / CNF's performance has degraded. If disk IO / s falls below threshold T3, it is estimated that the VNF / CNF's performance has degraded. If disk read / write latency exceeds threshold T4, it is estimated that the VNF / CNF's performance has degraded. If network IO / s falls below threshold T5, it is estimated that the VNF / CNF's performance has degraded. If packet loss rate exceeds threshold T6, it is estimated that the VNF / CNF's performance has degraded. If job or task completion rate falls below threshold T7, it is estimated that the VNF / CNF's performance has degraded. Therefore, if any of the above conditions are met, the VIM sends a signal to the OSS / NFVO indicating a degradation in performance.
[0039] Each VNFM and each EMS compares each monitoring item in the application layer in Figure 7 with the corresponding threshold (one of T8 to T14). If CPU usage exceeds threshold T8, it is estimated that the performance of the VNF / CNF has degraded. If memory usage exceeds threshold T9, it is estimated that the performance of the VNF / CNF has degraded. If disk IO / s falls below threshold T10, it is estimated that the performance of the VNF / CNF has degraded. If disk read / write latency exceeds threshold T11, it is estimated that the performance of the VNF / CNF has degraded. If the job or task completion rate falls below threshold T12, it is estimated that the performance of the VNF / CNF has degraded. If the number of accommodated users falls below threshold T13, it is estimated that the performance of the VNF / CNF has degraded. If CPS exceeds threshold T14, it is estimated that the performance of the VNF / CNF has degraded. Therefore, if any of the above conditions are met, the VNFM or EMS sends a signal to the OSS / NFVO indicating a degradation in performance.
[0040] In this embodiment, the CPU usage, memory usage, disk IO / sec, disk read / write latency, and job or task completion rate of the VNF / CNF are monitored at both the platform layer and the application layer. Threshold T1 may be the same as threshold T8. Threshold T2 may be the same as threshold T9. Threshold T3 may be the same as threshold T10. Threshold T4 may be the same as threshold T11. Threshold T7 may be the same as threshold T11. These monitoring items may be monitored at only either the platform layer or the application layer.
[0041] The sequence diagram of FIG. 8 shows an example of processing in a virtualized environment according to this embodiment. The operations enclosed in the dashed rectangle in FIG. 8 are signal exchanges between components in the platform monitoring network and application monitoring network described above. When the VIM, VNFM, or EMS reports a signal indicating a performance degradation to the OSS / NFVO, the decision block in the figure, "Performance degradation detected?", becomes positive. In this case, the OSS / NFVO records the performance degradation report and sends a decision request signal to the decision device, requesting a decision on whether to perform auto-healing or auto-scaling. The decision request signal includes information identifying the VNF / CNF whose performance has degraded. When the determination request signal is received, the CPU of the determination device determines whether to perform auto-healing or auto-scaling.
[0042] The CPU of the decision device returns the decision result to the OSS / NFVO. Based on the decision result, the OSS / NFVO performs auto-healing or auto-scaling. If it is determined that auto-healing should be performed, the CPU of the determination device executes a process to identify hardware resources related to the degradation of the VNF / CNF performance. The determination result instructing the OSS / NFVO to perform auto-healing, which is sent as a reply from the determination device, includes information identifying the hardware resources identified in the process. Upon receiving a decision result instructing the implementation of auto-healing, the OSS / NFVO sends a request to the VIM corresponding to the degraded VNF / CNF not to deploy the VNF / CNF to the hardware resources related to the identified performance degradation. In response to this request, the VIM and NFVI corresponding to the degraded VNF / CNF perform a procedure to exclude the hardware resources related to the identified performance degradation from the hardware resources for the VNF / CNF to be recovered. After completing the procedure, the VIM returns a completion acknowledgment signal (ACK) to the OSS / NFVO. In response to the completion confirmation signal, the OSS / NFVO sends an auto-healing start command to the VIM. Upon receiving the auto-healing start command, the VIM deletes the VNF / CNF with degraded performance and deploys other hardware resources for that VNF / CNF (creates a VNF / CNF). Although not shown, after the VNF / CNF creation is complete, the OSS / NFVO commands the VNFM or EMS corresponding to the created VNF / CNF to configure the VNF / CNF to incorporate the created VNF / CNF into operation. In accordance with the command, the VNFM or EMS configures the created VNF / CNF to incorporate it into operation.
[0043] Although not shown in the figure, upon receiving a judgment result instructing the implementation of auto-scaling, the OSS / NFVO sends a command to implement auto-healing to the VIM corresponding to the VNF / CNF whose performance has deteriorated. Upon receiving this command, the VIM adds a new VNF / CNF and deploys hardware resources for that VNF / CNF. After the VNF / CNF has been added, the OSS / NFVO commands the VNFM or EMS corresponding to the added VNF / CNF to configure it to incorporate the added VNF / CNF into operation. In accordance with the command, the VNFM or EMS configures the added VNF / CNF to incorporate it into operation.
[0044] Next, the determination operation of the CPU of the determination device will be described in more detail with reference to the flowchart of FIG. In step S1, the CPU determines whether or not it has received a determination request signal from the OSS / NFVO. If the determination in step S1 is affirmative, in step S2, the CPU determines whether or not it has already received a determination request signal regarding a performance degradation of another VNF / CNF belonging to a group consisting of multiple load-balanced VNFs / CNFs. That is, if the VNF / CNF with degraded performance related to the determination request signal in step S1 belongs to a group consisting of multiple load-balanced VNFs or CNFs, the CPU of the determination device executes a first determination process to determine whether or not a degradation exceeding a threshold in the performance of another VNF / CNF belonging to that group has already been detected.
[0045] If the determination in step S2, i.e., the determination in the first determination process, is affirmative, the operation proceeds to step S3, where the CPU of the determination device returns a determination result instructing the OSS / NFVO to perform auto-scaling (executes the second command transmission process). In this way, auto-scaling is performed.
[0046] The determination in step S2 will be explained in more detail with reference to Fig. 10. As shown in Fig. 10, it is assumed that multiple VNFs / CNFs 1a to 1c are executed by one NFVI1, multiple VNFs / CNFs 2a to 2c are executed by one NFVI2, and multiple VNFs / CNFs 3a to 3c are executed by one NFVI3. It is also assumed that VNFs / CNFs 1a, 2a, and 3a belong to the same group and are load balanced, VNFs / CNFs 1b, 2b, and 3b belong to the same group and are load balanced, and VNFs / CNFs 1c, 2c, and 3c belong to the same group and are load balanced.
[0047] If the performance of all VNFs / CNFs belonging to a load-balanced group (e.g., VNFs / CNF1a, 2a, 3a) is degraded, it is unlikely that there is a physical problem with the hardware of all of these VNFs / CNFs, or a problem with the software, such as firmware, running on the hardware. Rather, it is assumed that the load on that group has become excessive. Therefore, if the determination in step S2, i.e., the determination in the first determination process, is affirmative, the operation proceeds to step S3, where auto-scaling is performed. In this case, the OSS / NFVO may send a command for auto-scaling to multiple components (multiple VIMs, multiple VNFMs, multiple EMSs) corresponding to multiple VNFs / CNFs whose performance has deteriorated.
[0048] On the other hand, if the performance of only one VNF / CNF (e.g., VNF / CNF2a) among VNFs / CNFs (e.g., VNF / CNF1a, 2a, 3a) belonging to a load-balanced group is degraded, it is likely that the load on the group is not excessive, but that only that VNF / CNF (e.g., VNF / CNF2a) is abnormal and that a silent failure has occurred in that VNF / CNF. In this case, it is highly likely that it is appropriate to perform auto-healing while excluding the hardware resources related to the performance degradation. If the determination in step S2, i.e., the determination in the first determination process, is negative, the operation proceeds to step S4. In step S4, the CPU of the determination device determines whether or not a determination request signal has already been received regarding a performance degradation of another VNF / CNF operating on the NFVI on which the VNF / CNF with degraded performance operates. That is, a second determination process is executed to determine whether or not a degradation exceeding a threshold in performance of another VNF / CNF operating on the NFVI on which the VNF / CNF with degraded performance operates, related to the determination request signal in step S1, has already been detected.
[0049] If the determination in step S4, i.e., the determination in the second determination process, is affirmative, the operation proceeds to step S5, and the CPU of the determination device executes identification processing to identify a hardware resource related to an NFVI that has a problem in multiple VNFs / CNFs under it. That is, in step S5, a hardware resource related to an NFVI that has a problem is identified as a hardware resource related to the degradation of the performance of the VNFs / CNFs.
[0050] Then, in step S6, the CPU of the determination device returns a determination result instructing the OSS / NFVO to perform auto-healing (executes a first command transmission process). If step S5 is executed immediately before step S6, the determination result instructing the OSS / NFVO to perform auto-healing includes information identifying the hardware resource identified in step S5. In this way, auto-healing is performed so that hardware resources other than the hardware resource identified in step S5 are deployed for the VNF / CNF whose performance has deteriorated.
[0051] If the determination in step S4, i.e., the determination in the second determination process, is negative, the operation proceeds to step S7, where the CPU of the determination device executes an identification process to identify hardware resources related to the performance degradation of a single VNF / CNF whose performance has degraded, which is related to the determination request signal in step S1. That is, in step S7, hardware resources related only to that VNF / CNF are identified as hardware resources related to the degradation of the VNF / CNF's performance.
[0052] Then, in step S6, the CPU of the determination device returns a determination result instructing the OSS / NFVO to perform auto-healing (executes a first command transmission process). If step S7 is executed immediately before step S6, the determination result instructing the OSS / NFVO to perform auto-healing includes information that identifies the hardware resource identified in step S7. In this way, auto-healing is performed so that hardware resources other than the hardware resource identified in step S7 are deployed for the VNF / CNF whose performance has deteriorated.
[0053] The determination in step S4 will be explained in more detail with reference to Figure 10. Multiple VNFs / CNFs operating on the same NFVI (server) correspond to different applications operating on the same server. For example, VNFs / CNF2a to 2c operating on NFVI2 correspond to different applications operating on the same server. If performance is degraded in different applications running on the same server, it is unlikely that there is a problem with all of those applications. Rather, it is assumed that the server is abnormal. For example, if performance is degraded in all of VNFs / CNF2a to 2c running on NFVI2, it is assumed that NFVI2 is abnormal rather than that there is a problem with VNFs / CNF2a to 2c. Therefore, if the judgment in step S4, i.e., the judgment of the second judgment process, is positive, the operation proceeds to step S5, and hardware resources related to the NFVI that has problems with multiple subordinate VNFs / CNFs are identified as hardware resources that should be excluded by auto-healing.
[0054] On the other hand, if the performance of other applications running on the same server is not degraded, it is presumed that only the VNF / CNF corresponding to the application whose performance has degraded is abnormal. Therefore, if the determination in step S4, i.e., the determination in the second determination process, is negative, the operation proceeds to step S7, where the hardware resources related only to that VNF / CNF are identified as hardware resources to be excluded by autohealing.
[0055] As described above, in this embodiment, when a degradation in performance of a VNF / CNF occurs that could result in a silent failure in a network implemented in a virtualized environment, auto-healing is appropriately performed, thereby automatically minimizing adverse effects on services as quickly as possible. Furthermore, when it is estimated that implementing auto-scaling is appropriate, auto-scaling is performed. Furthermore, when it is estimated that the degradation in performance of a VNF / CNF is caused by NFVI, hardware resources related to NFVI are excluded by auto-healing, and when it is estimated that the degradation in performance of a VNF / CNF is caused only by that VNF / CNF, hardware resources related to that VNF / CNF are excluded by auto-healing.
[0056] However, steps S2 and S3 may be omitted. In this case, if the determination in step S1 becomes positive, the operation proceeds directly to step S4. Also, steps S4 and S5 may be omitted. In this case, if the determination in step S2 is negative, the operation proceeds directly to step S7 and then to step S6.
[0057] While the present disclosure has been shown and described with reference to preferred embodiments thereof, it will be understood by those skilled in the art that changes in form and detail may be made therein without departing from the scope of the appended claims, and such changes, modifications and alterations are intended to be within the scope of the present disclosure.
[0058] Aspects of the present disclosure are also described in the following numbered clauses:
[0059] [1] a receiving process for receiving a signal that is provided when a degradation in performance of a VNF or CNF beyond a threshold is detected in a network virtualization environment; an identification process that, upon receiving the signal, identifies a hardware resource associated with the performance degradation; a first command sending process for sending a command to perform auto-healing so as to deploy a hardware resource other than the hardware resource for the VNF or CNF whose performance has deteriorated; At least one processor to run A network management device comprising:
[0060] [2] The processor: After the receiving process and before the identifying process, if the VNF or CNF whose performance has deteriorated belongs to a group consisting of multiple load-balanced VNFs or CNFs, a first determination process is performed to determine whether a deterioration in performance of other VNFs or CNFs belonging to the group that exceeds a threshold has been detected; a second command transmission process for transmitting a command to perform auto-scaling when the determination of the first determination process is affirmative; Execute The network management device according to [1].
[0061] [3] The processor: After the receiving process and before the identifying process, a second determination process is performed to determine whether a performance degradation exceeding a threshold of another VNF or CNF operating on an NFVI on which the VNF or CNF with degraded performance operates has been detected; If the determination of the second determination process is affirmative, the identification process identifies hardware resources associated with the NFVI; If the determination in the second determination process is negative, the identification process identifies a hardware resource associated with the VNF or CNF whose performance has deteriorated. The network management device according to [1] or [2],
[0062] [4] receiving a signal in a network virtualization environment that is provided when a degradation in performance of a VNF or CNF beyond a threshold is detected; Upon receiving the signal, identifying a hardware resource associated with the performance degradation; and sending an instruction to perform autohealing so as to deploy hardware resources other than the hardware resource for the VNF or CNF whose performance has degraded. Network management methods, including: [Explanation of symbols]
[0063] Mo...Mobile network, 1,2,3...NFVI, 1a-1c...VNF / CNF, 2a-2c...VNF / CNF, 3a-3c...VNF / CNF
Claims
1. a receiving process for receiving a signal that is provided when a degradation in performance of a VNF or CNF beyond a threshold is detected in a network virtualization environment; an identification process that, upon receiving the signal, identifies a hardware resource associated with the performance degradation; a first command transmission process for transmitting a command to execute auto-healing so as to deploy a hardware resource other than the hardware resource for the VNF or CNF whose performance has deteriorated; At least one processor that runs A network management device comprising:
2. The processor: a first determination process for determining whether or not a performance degradation exceeding a threshold value of another VNF or CNF belonging to a group consisting of a plurality of load-balanced VNFs or CNFs is detected when the VNF or CNF whose performance has degraded belongs to the group after the reception process and before the identification process; a second command transmission process for transmitting a command to perform auto-scaling when the determination in the first determination process is affirmative; Execute 2. The network management device according to claim 1.
3. The processor: After the receiving process and before the identifying process, a second determination process is performed to determine whether a performance degradation exceeding a threshold of another VNF or CNF operating on an NFVI on which the VNF or CNF with degraded performance operates is detected; If the determination in the second determination process is affirmative, the identification process identifies a hardware resource associated with the NFVI; If the determination in the second determination process is negative, the identification process identifies a hardware resource associated with the VNF or CNF whose performance has deteriorated.
3. The network management device according to claim 1, wherein:
4. receiving a signal that is provided when a degradation in performance of a VNF or CNF beyond a threshold is detected in a network virtualization environment; Upon receiving the signal, identifying a hardware resource associated with the performance degradation; and sending a command to perform auto-healing so as to deploy a hardware resource other than the hardware resource for the VNF or CNF whose performance has deteriorated. Network management methods, including:
Citation Information
Patent Citations
Virtual network function management device, system, healing method, and program
WO2016121830A1