Log analysis when failure occurs in virtualized environment

The network management device automates log analysis in virtualized networks by using a log dictionary to quickly identify failure causes and impact scope, addressing access and definition challenges across different organizational logs.

JP2025114132APending Publication Date: 2025-08-05RAKUTEN MOBILE INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024008619
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-24
Publication Date
2025-08-05

Smart Images

  • Figure 2025114132000001_ABST
    Figure 2025114132000001_ABST
Patent Text Reader

Abstract

To quickly perform primary analysis of a log when a failure occurs in virtualized environment.SOLUTION: A network management device executes storage processing, memory processing, retrieval processing, and presentation processing. The storage processing is processing to store logs of a plurality of components constituting a virtualized environment of a network. The memory processing is processing to store correspondence information in which error logs of a plurality of components related to failures are associated with each failure that can occur in a network. When a failure occurs in the network, the retrieval processing is processing to retrieve an error log of a second component related to the failure from the stored logs, using the correspondence information, based on an error log of a first component among the plurality of components. The presentation processing is processing to present the result of the retrieval to a user.SELECTED DRAWING: Figure 14
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to log analysis when a failure occurs in a virtualized environment. [Background technology]

[0002] With the improvement in performance of general-purpose servers and the expansion of network infrastructure, cloud computing (hereafter referred to as "cloud"), which uses virtualized computing resources on physical resources such as servers on demand, has become widespread. NFV (Network Function Virtualization), which virtualizes network functions and provides them on the cloud, is also well known. NFV is a technology that uses virtualization and cloud technologies to separate the hardware and software of various network services that previously ran on dedicated hardware, and runs the software on a virtualized platform. This is expected to lead to more advanced operations and cost reductions. In recent years, virtualization has also been progressing in mobile networks. The European Telecommunications Standards Institute (ETSI) NFV defines the architecture of NFV (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] International Publication No. 2016 / 121802 Summary of the Invention [Problem to be solved by the invention]

[0004] Recent telecom networks are virtualized networks that run applications that make up Network Function Virtualization (VNF) on virtualization infrastructure servers (Compute Nodes). In such virtualized networks, if a failure occurs in the virtualization infrastructure, it can also affect the applications running on it. When a failure occurs in a network, a primary analysis is first performed to identify the location where the failure is thought to have occurred (suspect location) and the extent of the failure's impact. To perform this primary analysis, it is necessary to analyze the logs of each component.

[0005] However, in large-scale networks such as telecom networks, the person in charge of the virtualization infrastructure (infrastructure owner) and the person in charge of the application (application owner) are separated into different organizations (different departments, different companies), and system development is carried out under a division of labor system. As a result, the environment may be such that each owner cannot access each other's logs, and the definitions of logs for each component may differ from developer to developer. This makes it difficult for infrastructure owners to determine the extent to which a failure in the virtualization platform affects their applications. It also makes it difficult for application owners to determine whether the cause of an application failure is in the virtualization platform or the application itself. As a result, communication costs increase and log analysis and problem resolution take time.

[0006] Therefore, an object of the present disclosure is to quickly perform a primary analysis of logs when a failure occurs in a virtualized environment. [Means for solving the problem]

[0007] A network management device according to one aspect of the present disclosure includes one or more processors, and at least one of the one or more processors executes a save process, a storage process, a search process, and a presentation process. The save process is a process of saving logs of multiple components constituting a virtualized environment of a network. The storage process is a process of storing correspondence information that associates, for each failure that may occur in the network, error logs of the multiple components related to the failure. The search process is a process of, when a failure occurs in the network, searching the saved logs for an error log of a second component related to the failure using the correspondence information based on the error log of a first component of the multiple components. The presentation process is a process of presenting the search results to a user.

[0008] A network management method according to one aspect of the present disclosure includes saving logs of multiple components that constitute a virtualized environment of a network, storing correspondence information that associates, for each failure that may occur in the network, error logs of the multiple components related to the failure, and when a failure occurs in the network, using the correspondence information based on the error log of a first component of the multiple components, searching the saved logs for an error log of a second component related to the failure, and presenting the results of the search to a user.

[0009] A network management system according to one aspect of the present disclosure includes one or more processors, and at least one of the one or more processors executes a save process, a storage process, a search process, and a presentation process. The save process is a process of saving logs of multiple components constituting a virtualized environment of a network. The storage process is a process of storing correspondence information that associates, for each failure that may occur in the network, error logs of the multiple components related to the failure. When a failure occurs in the network, the search process is a process of searching the saved logs for an error log of a second component related to the failure using the correspondence information based on the error log of a first component of the multiple components. The presentation process is a process of presenting the search results to a user. [Effects of the Invention]

[0010] According to one aspect of the present disclosure, it is possible to quickly perform a primary analysis of a log when a failure occurs in a virtualized environment. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a diagram showing an example of the configuration of a mobile network including a network management device according to this embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of the internal configuration of the network management system. [Figure 3] FIG. 3 is a functional block diagram of the log management unit. [Figure 4] FIG. 4 shows an example of the configuration of the log dictionary database. [Figure 5] Figure 5 shows an example of the configuration of a log storage database. [Figure 6] FIG. 6 is a sequence diagram showing the operation of registering a record in a dictionary. [Figure 7] FIG. 7 is a sequence diagram showing the dictionary update operation. [Figure 8] FIG. 8 is a sequence diagram showing a dictionary reference operation. [Figure 9] FIG. 9 is a sequence diagram showing a dictionary deletion operation. [Figure 10] FIG. 10 is a sequence diagram showing the operation of registering a new log record. [Figure 11] FIG. 11 is a sequence diagram showing a log update operation. [Figure 12] FIG. 12 is a sequence diagram showing a log reference operation. [Figure 13] FIG. 13 is a sequence diagram showing a log deletion operation. [Figure 14] FIG. 14 is a sequence diagram showing the log analysis operation. [Figure 15] FIG. 15 shows an example of an application log reference screen. [Figure 16] FIG. 16 shows an example of a screen for explaining the details of an application log. [Figure 17] Figure 17 shows an example of a screen for viewing logs for a virtualization platform. [Figure 18] FIG. 18 shows an example of a detailed explanation screen for the virtualization platform log. [Figure 19] FIG. 19 is a block diagram illustrating an example of a hardware configuration of a network management device. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Among the components disclosed below, components having the same functions are designated by the same reference numerals, and their description will be omitted. Note that the embodiment disclosed below is one form of the present disclosure, and should be appropriately modified or changed depending on the configuration of the device and various conditions, and is not limited to the following embodiment. Furthermore, not all of the combinations of features described in the present embodiment are necessarily essential to solving the above-mentioned problems.

[0013] Hereinafter, a case will be described in which the network management device according to this embodiment has a log management function for managing logs when a failure occurs in a mobile network built on a virtualization platform and for performing a primary analysis of the logs. The network management device in this embodiment has a log sharing function and a log dictionary function as log management functions. Specifically, the network management device in this embodiment stores logs of multiple components that make up the network virtualization environment in a database for log sharing. The network management device also stores a log dictionary, which is correspondence information that associates error logs of multiple components related to each failure that may occur in the network. When a failure occurs in the network, the network management device uses the log dictionary to search the logs stored in the database for an error log of a second component related to the failure, based on the error log of a first component of the multiple components.

[0014] Here, the correspondence information (log dictionary) may be information that associates keywords that represent the contents of error logs of multiple components related to the failure. In this case, the network management device may refer to the log dictionary based on the keywords that represent the contents of the error log of the first component to search for keywords that represent the contents of the error log of the second component, and then search for the error log of the second component from the logs stored in the database based on the searched keywords. This enables automatic primary analysis, such as isolating the cause of a failure and identifying the scope of impact when a failure occurs.

[0015] FIG. 1 is a diagram showing an example of a network configuration of a mobile network 100 including a network management device according to this embodiment. In the mobile network 100 shown in Figure 1, a mobile communication-enabled terminal such as a smartphone communicates wirelessly with a radio access network (RAN), and the information is relayed through a backhaul network (mobile backhaul: MBH) and sent to a core network for processing, allowing the terminal to connect to the Internet 200 or to connect to another company's network to make voice calls.

[0016] Specifically, the mobile network 100 is configured to include a base station 11 and multiple accommodating stations 12 to 14. Here, the accommodating station 12 is an edge data center, the accommodating station 13 is a regional data center (RDC), and the accommodating station 14 is a central data center (CDC). A backhaul network is configured between the edge data center 12 and the central data center 14. The mobile network 100 in this embodiment is a virtualized network built on a virtualization platform. In this mobile network 100, everything from the backbone network switching equipment to the base station wireless access functions is implemented by software on a general-purpose server.

[0017] The base station 11 includes an antenna, a distribution board, a battery, and the like. The edge data center 12 is installed near the base station 11 and is connected to each of the multiple base stations 11 via optical fiber cables, etc. The edge data center 12 implements a RAN-related wireless access function. The regional data center 13 is connected to multiple edge data centers 12 located in the target region. In this regional data center 13, firewall / NAT (Network Address Translation), CDN (Content Distribution Network), and various applications for edge computing are realized by software. The central data center 14 is connected to a plurality of regional data centers 13. The central data center 14 implements core functions such as EPC (Evolved Packet Core) and IMS (IP Multimedia Subsystem).

[0018] The number of data centers (accommodation stations) such as the edge data center 12, the regional data center 13, and the central data center 14 is not limited to the number shown in Fig. 1. For example, although Fig. 1 shows only one regional data center 13 and one central data center 14, multiple regional data centers 13 and multiple central data centers 14 may be installed.

[0019] FIG. 2 is a diagram showing an example of the internal configuration of a network management system that constitutes the mobile network 100. As shown in FIG. Each of the components shown in Fig. 2 has a reference point. Lines connecting the components shown in Fig. 2 indicate that they can send and receive information to and from each other. The NFVI (NFV Infrastructure) 110 is a network function virtualization infrastructure and is configured to include physical resources, a virtualization layer, and virtualized resources. The physical resources include hardware resources such as computational resources, storage resources, and transmission resources. The virtualization layer is a virtualization layer such as a hypervisor that virtualizes the physical resources and provides them to the VNF (Network Function Virtualization) 120. The virtualized resources are virtualized infrastructure resources provided to the VNF 120.

[0020] In other words, NFVI110 is a platform that enables the hardware resources of a physical server (hereinafter simply referred to as a "server"), such as computing, storage, and network functions, to be flexibly handled as virtualized hardware resources, such as virtualized computing, virtualized storage, and virtualized networks, which are virtualized using a virtualization layer such as a hypervisor.

[0021] A plurality of servers constituting the NFVI 110 are collectively arranged in data centers (accommodation stations) 12 to 14. The number of servers arranged in each of the data centers 12 to 14, their locations, wiring, etc. are predetermined depending on the type of data center (accommodation station type). In each of the data centers 12 to 14, the arranged servers are connected by an internal network, allowing them to send and receive information to and from each other. In addition, the data centers are connected by a network, allowing servers provided in different data centers to send and receive information to and from each other via the network.

[0022] The VNF 120 corresponds to an application that runs on a virtual machine (VM) on a server, and realizes a network function in software. Although not specifically illustrated, a management function called an EM (Element Manager) may be provided for each VNF 120. The virtualized environment is composed of the NFVI 110 and VNF 120 in Figure 2. In other words, the virtualized environment is composed of three layers, from the bottom up: hardware, virtualization layer, and virtual machine. In the following description, the NFVI110, which is a component that constitutes the virtualized environment, is referred to as the "virtualization infrastructure," and the VNF120, which is a component that constitutes the virtualized environment, is referred to as the "application."

[0023] The MANO (Management and Orchestration) 130 has a management function and an orchestration function for a virtualized environment. The MANO 130 includes an NFVO (NFV-Orchestrator) 131, a VNFM (VNF-Manager) 132, and a VIM (Virtualized Infrastructure Manager) 133. The NFVO 131 orchestrates NFVI resources, manages the lifecycle of network services, and performs integrated operation and management of the entire system. The NFVO 131 can perform processing in response to instructions from the OSS / BSS (Operation Support System / Business Support System) 140, which will be described later.

[0024] The VNFM 132 manages the life cycle of the VNFs 120. Note that the VNFM 132 may be arranged in the MANO 130 as a dedicated VNFM corresponding to each VNF 120. Alternatively, one VNFM 132 may manage the life cycle of two or more VNFs 120. In this case, the VNFM 132 may be a general-purpose VNFM corresponding to VNFs 120 provided by different vendors. The VIM 133 performs operation management of the resources used by the VNF 120 .

[0025] The OSS / BSS 140 is an integrated management system for the mobile network 100 . Here, OSS refers to the systems (equipment, software, mechanisms, etc.) required to build and operate a service, and BSS refers to the information systems (equipment, software, mechanisms, etc.) used for charging usage fees, billing, customer support, etc.

[0026] The log management unit 150 realizes a log management function that performs a primary analysis of a log when a failure occurs. This log management unit 150 constitutes the network management device according to this embodiment. In the mobile network 100, if a failure occurs in the virtualization platform, it may affect the applications running on that platform. When a failure occurs in the network, it is necessary to perform a primary analysis by analyzing the logs of each component to identify the location where the failure is thought to have occurred (suspect location) and the extent of the impact of the failure. However, in the past, this primary analysis required a lot of man-hours.

[0027] One of the reasons for this is that the people responsible for developing, building, and operating the virtualization infrastructure (infrastructure owners) and the people responsible for developing, building, and operating the applications that make up the VNF (application owners) are separate organizations (different departments, different companies), and system development is carried out under a division of labor structure.In this case, the log definitions on the virtualization infrastructure side and the application side differ, making it difficult for the infrastructure owner and application owner to analyze each other's logs. Another reason is that application owners are in an environment where they cannot easily access logs on the virtualization platform. Most logs on the virtualization platform are stored in the local area of the servers that make up the virtualization platform. Furthermore, because VNFs are deployed on the servers, which are shared resources, access rights to the local area of the servers are restricted due to security issues and other factors.

[0028] In an environment with access restrictions like the one described above, if a problem occurs on the application side, the application owner will need to ask the infrastructure owner to analyze the logs on the virtualization platform to determine whether the cause of the problem is on the application side or the virtualization platform side, which increases the infrastructure owner's workload. Furthermore, if a failure occurs on the virtualization platform, the infrastructure owner needs to identify the extent of the impact that the failure will have on the application. However, because log definitions differ between organizations, log analysis is not possible, making it difficult to pinpoint the problem on the application side. This time-consuming log analysis and problem resolution increases network outage times and reduces performance.

[0029] Therefore, in this embodiment, the log management unit 150 automatically performs a primary analysis of the log when a failure occurs. Specifically, the log management unit 150 uses a log dictionary to search a database for related error logs on the virtualization platform based on the error log on the application side, and automatically determines whether a failure that occurred in an application is due to a problem on the application side or whether there is a possibility that it is a problem on the virtualization platform side. In this way, when an application error occurs, the problem is automatically isolated. Furthermore, the log management unit 150 uses a log dictionary to search a database for related application error logs based on the virtualization platform error log, and automatically identifies the application that is affected by the failure that occurred on the virtualization platform. In this way, the scope of the impact when a virtualization platform error occurs is automatically identified.

[0030] 2, the log management unit 150 is not limited to being an external function of the OSS / BSS 140 or the MANO 130. The log management unit 150 may be provided inside the OSS / BSS 140 or inside the MANO 130. In this case, the log management function of the log management unit 150 becomes part of the function of the OSS / BSS 140 or the MANO 130.

[0031] FIG. 3 is a functional block diagram of the log management unit 150. 3, the log management unit 150 includes a log dictionary management unit 151, a log storage management unit 152, and a search result presentation unit 153. Furthermore, the log management unit 150 includes a log dictionary database (DB) 150a and a log storage database (DB) 150b. The log dictionary database 150a stores correspondence information (log dictionary) that associates keywords representing the contents of the virtualization platform error log and the application error log related to each failure that may occur in the mobile network 100. Here, the keywords may be words or simple sentences that allow workers who are not familiar with log definitions or log analysis to easily understand the contents of the error log.

[0032] Fig. 4 is a diagram showing an example of the configuration of the log dictionary database 150a. As shown in Fig. 4, the log dictionary database 150a can store keywords (application keywords) that represent the contents of the error log of an application and keywords (infrastructure keywords) that represent the contents of the error log of the virtualization infrastructure in association with each other. Furthermore, as shown in Fig. 4, the log dictionary database 150a may store detailed information (detailed explanation) of each failure.

[0033] For example, the log dictionary database 150a can store a keyword "memory error" on the application side and a keyword "physical memory failure" on the virtualization platform side in association with each other. If an error log corresponding to the application keyword "memory error" is found on the application side, a dictionary search based on the application keyword "memory error" will yield the infrastructure keyword "physical memory failure." In this case, it is possible that the failure occurring on the application side is caused by a physical memory failure on the virtualization infrastructure side.

[0034] The log dictionary database 150a can register events (failures) that are known to be likely to occur in advance at the system design stage. Furthermore, the log dictionary database 150a may also be registered when verifying the operation of the system in a staging environment (test environment). For example, if a failure occurs in a test environment, the infrastructure owner and the application owner can each verify the logs on the virtualization infrastructure side and the logs on the application side, and, working together, define and register keywords corresponding to the failure that has occurred. Furthermore, the log dictionary database 150a can also register (add) failures that have occurred in an actual production environment (operational environment).

[0035] 3, the log storage database 150b stores the logs on the virtualization platform side and the logs on the application side along with their respective keywords, which correspond to the keywords registered in the log dictionary database 150a. Fig. 5 is a diagram showing an example of the configuration of log storage database 150b. As shown in Fig. 5, log storage database 150b stores the error logs of each component in association with keywords. Also, as shown in Fig. 5, log storage database 150b may store information about the target component (target application name, target infrastructure name) and log information (log type, log name). Furthermore, in addition to the information shown in Fig. 5, log storage database 150b may also store information indicating the storage destination of the error log, for example. In this embodiment, we will explain the case where logs from the virtualization platform side and logs from the application side are stored together in log storage database 150b, but the database that stores logs from the virtualization platform side and the database that stores logs from the application side may be separate.

[0036] Returning to FIG. 3, the log dictionary management unit 151 includes a dictionary DB operation unit 154 and a dictionary search unit 155. The dictionary DB operation unit 154 performs operations on the log dictionary database 150a based on instructions from a user. Here, the user may be, for example, an infrastructure owner or an application owner. The operations may include registering a record in a dictionary, updating a dictionary, referencing a dictionary, and deleting a dictionary. The dictionary search unit 155 refers to the log dictionary database 150a based on keywords representing the contents of the error log of either the virtualization infrastructure or the application (first component) and searches for keywords of the other component (second component) of the virtualization infrastructure or the application.

[0037] The log storage management unit 152 includes a storage DB operation unit 156 , an error log reading unit 157 , and a keyword search unit 158 . The log storage database operation unit 156 performs operations on the log storage database 150b based on instructions from a user. Here, the user may be, for example, an infrastructure owner or an application owner. The operations may include registering a new log record, referencing a log, and deleting a log.

[0038] The error log reading unit 157 reads out the error log and keywords to be analyzed from the log storage database 150b. Based on the keyword of the second component searched by the dictionary search unit 155, the keyword search unit 158 searches the logs stored in the log storage database 150b for an error log of the second component that corresponds to the keyword. The search result presentation unit 153 presents the results of the search processes in the log dictionary management unit 151 and the log storage management unit 152 to the user.

[0039] Note that the functional block configuration of the log management unit 150 shown in FIG. 3 is an example, and multiple functional blocks may form one functional block, or any functional block may be divided into blocks that perform multiple functions. Furthermore, the multiple functions of the log management unit 150 may be divided into external functions of the OSS / BSS 140 and MANO 130 of the network management system shown in FIG. 2, internal functions of the OSS / BSS 140, and internal functions of the MANO 130.

[0040] An outline of the operation of the log dictionary database 150a in the dictionary DB operation unit 154 will be described below. 6 is a sequence diagram showing the dictionary record registration operation, which is an operation for registering a new dictionary record in the log dictionary database 150a. First, in step S1, the user 300 transmits a dictionary registration command to the OSS 140. Here, the user 300 may be, for example, an application owner or an infrastructure owner. The dictionary registration command may include the number of records in the dictionary to be newly registered.

[0041] When the OSS 140 receives a dictionary registration command from the user 300, in step S2, the OSS 140 transmits a dictionary registration request to the log dictionary management unit 151 based on the information included in the dictionary registration command. In step S3, the log dictionary management unit 151 receives a dictionary registration request from the OSS 140 and registers a record in the log dictionary database 150a. When the record is successfully registered (step S4), the log dictionary management unit 151 transmits a dictionary registration successful completion notification to the user 300 via the OSS 140 (steps S5 and S6).

[0042] 7 is a sequence diagram showing the dictionary update operation, which is an operation for updating the contents of records registered in the log dictionary database 150a. First, in step S11, the user 300 performs dictionary input to the OSS 140. Here, the input information input by the user 300 is information to be recorded in a record registered in the log dictionary database 150a, and may include the application keywords, infrastructure keywords, and detailed explanation of the fault shown in FIG.

[0043] When the OSS 140 receives dictionary input from the user 300, in step S12, it transmits a dictionary update request to the log dictionary management unit 151 based on the input information. In step S13, the log dictionary management unit 151 receives a dictionary update request from the OSS 140 and updates the dictionary record in the log dictionary database 150a. When the record update is successfully completed (step S14), the log dictionary management unit 151 transmits a notification of successful completion of the dictionary update to the user 300 via the OSS 140 (steps S15 and S16).

[0044] 8 is a sequence diagram showing a dictionary reference operation, which refers to the contents of records registered in the log dictionary database 150a. First, in step S21, the user 300 transmits a dictionary read command to the OSS 140. Here, the dictionary read command may include information capable of identifying the record to be read, such as a keyword registered in the dictionary or a record number.

[0045] When the OSS 140 receives a dictionary read command from the user 300, in step S22, the OSS 140 transmits a dictionary read request to the log dictionary management unit 151 based on the information included in the dictionary read command. In step S23, the log dictionary management unit 151 receives a dictionary read request from the OSS 140 and reads the target record from the log dictionary database 150a. If the record is read successfully (step S24), the log dictionary management unit 151 transmits the read result to the user 300 via the OSS 140 (steps S25 and S26).

[0046] 9 is a sequence diagram showing the dictionary deletion operation, which is an operation for deleting records registered in the log dictionary database 150a. First, in step S31, the user 300 transmits a dictionary deletion command to the OSS 140. Here, the dictionary deletion command may include information capable of identifying the record to be deleted, such as a keyword registered in the dictionary or a record number.

[0047] When the OSS 140 receives a dictionary deletion command from the user 300, in step S32, the OSS 140 transmits a dictionary deletion request to the log dictionary management unit 151 based on the information included in the dictionary deletion command. In step S33, the log dictionary management unit 151 receives the dictionary deletion request from the OSS 140 and deletes the target record from the log dictionary database 150a. If the record is successfully deleted (step S34), the log dictionary management unit 151 transmits a dictionary deletion successful completion notification to the user 300 via the OSS 140 (steps S35 and S36).

[0048] Next, an overview of the operation of the log storage DB 150b in the storage DB operation unit 156 will be described. 10 is a sequence diagram showing the operation of registering a new log record. The basic operation is the same as the operation of registering a dictionary record shown in FIG. First, in step S41, the user 300 sends a log registration command to the OSS 140. The log registration command may include the number of records in the log to be newly registered, etc. The log registration command may also include information on the target component, log information, and error log keywords shown in FIG.

[0049] When the OSS 140 receives a log registration command from the user 300, in step S42, the OSS 140 transmits a log registration request to the log storage management unit 152 based on the information included in the log registration command. In step S43, the log storage management unit 152 receives a log registration request from the OSS 140 and registers a new log record in the log storage database 150b. When the new log record is successfully registered (step S44), the log storage management unit 152 transmits a successful log registration notification to the user 300 via the OSS 140 (steps S45 and S46).

[0050] FIG. 11 is a sequence diagram showing a log update operation. The logs to be stored in the log storage database 150b are periodically transmitted from the virtualization infrastructure (NFVI 110) or the VNF 120 to the OSS 140 (step S51). Upon receiving the log, the OSS 140 transmits a log update request to the log storage management unit 152 in step S52. In step S53, the log storage management unit 152 updates the log in the log storage database 150b in response to a log update request from the OSS 140. When the log update is completed successfully (step S54), the log storage management unit 152 transmits a successful log update completion notification to the user 300 via the OSS 140 (steps S55 and S56).

[0051] In the sequence diagram shown in Figure 11, the specifications are such that logs are sent from the virtualization infrastructure and VNF to the log storage management unit 152 via OSS 140, but the specifications may also be such that logs are sent directly from the virtualization infrastructure and VNF to the log storage management unit 152. Furthermore, the logs transmitted to the log storage management unit 152 are not limited to logs transmitted from the virtualization platform and the VNF. For example, logs may be transmitted to the log storage management unit 152 via the OSS 140 or directly from management devices (not shown) that monitor the virtualization platform and the VNF, respectively.

[0052] 12 is a sequence diagram showing the log reference operation, which is basically the same as the dictionary reference operation shown in FIG. First, in step S61, the user 300 transmits a log read command to the OSS 140. Here, the log read command may include information capable of identifying the log to be read. When the OSS 140 receives a log read command from the user 300, in step S62, the OSS 140 transmits a log read request to the log storage management unit 152 based on the information included in the log read command. In step S63, the log storage management unit 152 receives a log read request from the OSS 140 and reads the target log from the log storage database 150b. If the log is read successfully (step S64), the log storage management unit 152 transmits the read result to the user 300 via the OSS 140 (steps S65 and S66).

[0053] 13 is a sequence diagram showing a log deletion operation, the basic operation of which is the same as the dictionary deletion operation shown in FIG. First, in step S71, the user 300 transmits a log deletion command to the OSS 140. Here, the log deletion command may include information capable of identifying the log to be deleted. When the OSS 140 receives a log deletion command from the user 300, in step S72, the OSS 140 transmits a log deletion request to the log storage management unit 152 based on the information included in the log deletion command. In step S73, the log storage management unit 152 receives a log deletion request from the OSS 140 and deletes the target log from the log storage database 150b. If the log deletion is successful (step S74), the log storage management unit 152 transmits a successful log deletion notification to the user 300 via the OSS 140 (steps S75 and S76).

[0054] 14 is a sequence diagram showing a log analysis operation, which can be executed when a failure occurs in the mobile network 100. First, in step S101, the user 300 sends an error log analysis command to the OSS 140. Here, the user 300 may be, for example, an application owner or an infrastructure owner. The following description will be given assuming that the user 300 is an application owner. The error log analysis command may include information capable of identifying the log to be analyzed, such as an application name or a log name.

[0055] When the OSS 140 receives an error log analysis command from the user 300, it sends an error log read request to the log storage management unit 152 in step S102. In step S103, the log storage management unit 152 receives an error log read request from the OSS 140 and reads the target error log and the keyword of the error log (target keyword) from the log storage database 105b. The read result is provided from the log storage management unit 152 to the OSS 140 (steps S104 and S105).

[0056] Next, in step S106, the OSS 140 transmits a dictionary search request for the error log to the log dictionary management unit 151. This dictionary search request includes the target keyword included in the read result. In step S107, the log dictionary management unit 151 refers to the log dictionary database 150a based on the target keyword included in the dictionary search request to search for keywords (related keywords) associated with the target keyword. These related keywords are error keywords of the virtualization platform that are related to the application error. Also, in step S107, the log dictionary management unit 151 may refer to the log dictionary database 150a to acquire detailed information about the failure associated with the target keyword. The searched related keywords and detailed information about the failure are provided from the log dictionary management unit 151 to the OSS 140 (steps S108 and S109).

[0057] Then, in step S110, the OSS 140 transmits a keyword search request to the log storage management unit 152. This keyword search request includes the related keywords provided by the log dictionary management unit 151. In step S111, the log storage management unit 152 searches for error logs of the virtualization infrastructure from the logs stored in the log storage database 150b based on the related keywords included in the keyword search request from the OSS 140. The search results are provided to the user 300 from the log storage management unit 152 via the OSS 140 (steps S112 to S114). For example, if an error log of a virtualization platform associated with a related keyword is found in the keyword search in step S111, the virtualization platform is presented to the user 300 as a possible cause of the application error.

[0058] In this way, when user 300 is the application owner, the log analysis operation shown in Figure 14 performs a dictionary search based on the application error log, searches the error log on the related virtualization platform side, and automatically determines possible causes of the failure. For example, suppose an error occurs in an application (APP1) and an error log "ERR:aaaa" is generated. In this case, upon receiving a command from the application owner to analyze the error log of the application "APP1," the log storage management unit 152 reads the error log "ERR:aaaa" and the keyword "KEY_1" from the log storage database 150b shown in FIG. 5 (steps S101 to S105 in FIG. 14).

[0059] Then, the log dictionary management unit 151 searches the log dictionary database 150a shown in FIG. 4 based on the keyword “KEY_1” and obtains “KEY_A”, “KEY_B”, and “KEY_C” as keywords of the related virtualization infrastructure error log (steps S106 to S109 in FIG. 14). Next, the log storage management unit 152 searches the log storage database 150b shown in Fig. 5 for an error log of SERVER1 that corresponds to the searched related keywords "KEY_A," "KEY_B," and "KEY_C" (steps S110 to S112 in Fig. 14). If an error log is found, the log storage management unit 152 presents the searched virtualization platform (SERVER1) to the user as a possible cause of the failure (steps S113 and S114 in Fig. 14).

[0060] FIG. 15 shows an example of an application log reference screen. The application owner can check the dictionary search results, for example, from a log reference screen 400 shown in FIG. The log reference screen 400 may include, for example, a log display area 410 that displays the original log, and a details display area 420 that displays detailed information about the error log. On the log reference screen 400, it is possible to refer to logs, for example, for each pod or each application. In the log display area 410, the error log 411 may be highlighted. Furthermore, in the details display area 420, a log details explanation screen 421 corresponding to each error log 411 may be displayed.

[0061] FIG. 16 shows an example of the log details explanation screen 421. 16, a keyword (for example, KEY_1) of the application error log may be displayed on the log detail explanation screen 421. This keyword is the target keyword read from the log storage database 150b in step S103 of FIG. Furthermore, a description of the application error and keywords of the virtualization platform error log related to the application error (e.g., KEY_A, KEY_B, KEY_C) may be displayed on the log detail description screen 421. The description and keywords are the detailed information and related keywords searched from the log dictionary database 150a in step S107 of FIG.

[0062] Furthermore, a link to a virtualization infrastructure log related to the application error may be displayed on the log detail explanation screen 421. This link is a link for guiding the user to the storage destination of the error log searched from the log storage database 150b in step S111 of FIG. This allows the app owner to easily check possible causes of application errors.

[0063] On the other hand, if an error occurs in an application (APP2) and an error log "ERR:bbbb" is output, the log storage management unit 152, upon receiving a command from the application owner to analyze the error log of the application "APP2," reads out the keyword "KEY_2" of the error log "ERR:bbbb" from the log storage database 150b shown in Fig. 5 (steps S101 to S105 in Fig. 14).Then, the log dictionary management unit 151 searches the log dictionary database 150a shown in Fig. 4 based on the keyword "KEY_2," and confirms that the keyword of the related virtualization platform error log has not been registered (steps S106 to S109 in Fig. 14). In this case, the application owner is notified that there are no related keywords. That is, the keywords for the related virtualization platform error log are left blank on the log details explanation screen 421 shown in Fig. 16. This allows the application owner to determine that the cause of the failure is not on the virtualization platform side, that is, that the failure is due to a problem on the application side.

[0064] Furthermore, when user 300 is the infrastructure owner, the log analysis operation shown in FIG. 14 performs a dictionary search based on the error log of the virtualization platform, searches the error log of the related application, and automatically determines the scope of the impact of the virtualization platform failure. For example, suppose an error occurs in the virtualization platform (SERVER1) and an error log "ERR:cccc" is generated. In this case, upon receiving a command from the infrastructure owner to analyze the error log of the virtualization platform "SERVER1," the log storage management unit 152 reads the error log "ERR:cccc" and the keyword "KEY_A" from the log storage database 150b shown in FIG. 5 (steps S101 to S105 in FIG. 14).

[0065] Then, the log dictionary management unit 151 searches the log dictionary database 150a shown in FIG. 4 based on the keyword "KEY_A" and obtains "KEY_1" as a keyword of the related application error log (steps S106 to S109 in FIG. 14). Next, based on the searched related keyword "KEY_1," the log save management unit 152 searches the log save database 150b shown in Fig. 5 for an error log of APP1 that corresponds to the searched related keyword (steps S110 to S112 in Fig. 14). If an error log is found, the log save management unit 152 presents the searched application (APP1) to the user as an application that has been affected by the virtualization infrastructure failure (steps S113 and S114 in Fig. 14).

[0066] Figure 17 shows an example of a screen for viewing logs for a virtualization platform. The infrastructure owner can check the dictionary search results, for example, from a log reference screen 500 shown in FIG. The log reference screen 500 may include a log display area 510 and a details display area 520, similar to the application log reference screen 400 shown in FIG. 15. On the log reference screen 500, it is possible to refer to logs, for example, for each pod or each server. In the log display area 510, an error log 511 may be highlighted. Furthermore, in the details display area 520, a log details explanation screen 521 corresponding to each error log 511 may be displayed.

[0067] FIG. 18 shows an example of the log details explanation screen 521. The log details explanation screen 521, similar to the application log details explanation screen 421 shown in Figure 16, can display keywords for the virtualization platform error log, a description of the virtualization platform error, keywords for the application error log related to the virtualization platform error, and a link to the application log related to the virtualization platform error. This allows infrastructure owners to easily see the extent to which virtualization infrastructure errors affect their applications.

[0068] As described above, log management unit 150, which is a network management device in this embodiment, includes log storage database 150b and stores logs of multiple components that make up the virtualized environment of mobile network 100. Log management unit 150 also includes log dictionary database 150a and stores correspondence information (log dictionary) that associates, for each failure that may occur in mobile network 100, error logs of multiple components related to the failure. When a failure occurs in mobile network 100, log management unit 150 searches logs stored in log storage database 150b for an error log of a second component related to the failure that has occurred, based on the error log of a first component among the multiple components, using the correspondence information stored in log dictionary database 150a, and presents the search results to user 300. Here, the multiple components may include a virtualization platform and an application on the virtualization platform.

[0069] In this way, the log management unit 150 in this embodiment stores the error log of the virtualization platform and the error log of the application in association with each other, and when an error occurs in an application, it can search the error log of the virtualization platform related to the error and present the result to the user (application owner) 300. Similarly, when an error occurs in the virtualization platform, the log management unit 150 can search the error log of the application related to the error and present the result to the user (infrastructure owner) 300.

[0070] This allows application owners to easily check whether an error log related to an application failure has been output on the virtualization platform side when the failure occurs in the application, allowing them to easily determine whether the failure is due to a problem on the application side or whether there is a possibility that the problem is on the virtualization platform side. Furthermore, when a failure occurs in the virtualization platform, the infrastructure owner can easily check whether an error log related to the failure has also been output on the application side, allowing the infrastructure owner to easily understand whether the failure that occurred in the virtualization platform is affecting the application.

[0071] Therefore, even in an environment where owners cannot access each other's logs, or where the log definitions for each component differ depending on the developer, it is possible to quickly and accurately detect the problem, isolate the cause of the problem, and identify the scope of the problem's impact.

[0072] Specifically, when the first component is an application and the second component is a virtualization platform, if an error log of the virtualization platform related to an application failure is searched from the logs stored in the log storage database 150b, the log management unit 150 can present the virtualization platform as a possible cause of the application failure. In this case, the application owner can request the infrastructure owner in charge of the virtualization platform, which is the suspected cause of the failure, to analyze the cause and perform recovery work.In this way, the application owner can properly isolate the problem and then request work from the infrastructure owner, thereby minimizing the infrastructure owner's work.

[0073] Furthermore, if no virtualization infrastructure error log related to the application failure is found from the logs stored in the log storage database 150b, the log management unit 150 can present the application as the cause of the failure. In this case, the app owner can quickly start work such as debugging, thereby shortening the time it takes to resolve the problem.

[0074] Furthermore, when the first component is a virtualization platform and the second component is an application, if an error log of an application related to a failure in the virtualization platform is searched from the logs stored in the log storage database 150b, the log management unit 150 can present the application as being within the scope of the impact of the failure. In this case, the infrastructure owner can properly identify which application is affected by the failure that occurred in the virtualization platform, and can properly communicate the status of the response to the failure to the application owner in charge of the application affected by the failure.

[0075] In this way, when a failure occurs, the infrastructure owner and the application owner can share the same understanding of the problem and work together to quickly resolve it.

[0076] Furthermore, the log management unit 150 in this embodiment can store keywords representing the contents of error logs of multiple components in association with each other in the log dictionary database 150a, as shown in Fig. 4. In this case, the log management unit 150 can search for keywords (related keywords) in the error log of a second component based on the keywords in the error log of a first component by referring to the correspondence information stored in the log dictionary database 150a, and can search for the error log of the second component from the logs stored in the log storage database 150b based on the searched related keywords.

[0077] In this way, by performing a search process using a keyword, it is possible to easily search for related keywords and related error logs. Furthermore, the log management unit 150 may present the keywords of the error log of the first component and the keywords of the error log of the second component as search results to the user 300. This allows the user 300 to easily understand the details of the failure occurring in the first component and the details of the error in the second component related to the failure.

[0078] Furthermore, the log management unit 150 may present the user 300 with a link to the error log of the second component searched from the logs stored in the log storage database 150b as a search result. This allows the user 300 to easily access the error log of the second component related to the failure occurring in the first component and check its contents.

[0079] As described above, in this embodiment, when a failure occurs in a virtualized environment, the primary analysis of the log can be performed quickly, thereby shortening the time from the occurrence of the failure to the recovery of the system, thereby shortening the time the failure affects users of the mobile network 100, and improving performance.

[0080] The network management device according to this embodiment may be implemented in any general-purpose server that constitutes the backhaul network, core network, etc. of the mobile network 100. The network management device may also be implemented in a dedicated server. The network management device may also be implemented on a single computer or multiple computers. 19, the network management device 1 may include a CPU 2, a ROM 3, a RAM 4, a HDD 5, an input unit (keyboard, pointing device, etc.) 6, a display unit (monitor, etc.) 7, a communication I / F 8, etc. The network management device 1 may also include an external memory.

[0081] The CPU 2 is configured with one or more processors and performs overall control of the operations of the network management device 1. At least some of the functions of each element of the log management unit 150 shown in Fig. 3 can be realized by the CPU 2 executing a program. Note that the program may be stored in a non-volatile memory such as the ROM 3 or the HDD 5, or in an external memory such as a removable storage medium (not shown).

[0082] However, at least some of the elements of the log management unit 150 shown in Fig. 3 may be configured to operate as dedicated hardware. In this case, the dedicated hardware operates under the control of the CPU 2. For functions implemented by hardware, a dedicated circuit can be automatically generated on an FPGA from a program for implementing the function of each functional module using, for example, a predetermined compiler. Alternatively, a gate array circuit can be formed in the same way as an FPGA and implemented as hardware. Alternatively, the functions can be implemented using an ASIC (Application Specific Integrated Circuit).

[0083] Aspects of the present disclosure may include a computer-readable storage medium storing a program, wherein the program includes instructions that, when executed by a CPU 2 (at least one of one or more processors) of the network management device 1, cause the network management device 1 to perform at least one of the methods described above.

[0084] Although specific embodiments have been described above, these embodiments are merely examples and are not intended to limit the scope of the present disclosure. The devices and methods described herein may be embodied in forms other than those described above. Furthermore, appropriate omissions, substitutions, and modifications may be made to the above-described embodiments without departing from the scope of the present disclosure. Such omissions, substitutions, and modifications are included within the scope of the claims and their equivalents, and belong to the technical scope of the present disclosure.

[0085] (Embodiments of the present disclosure) The present disclosure includes the following embodiments. [1] A network management device comprising one or more processors, wherein at least one of the one or more processors executes the following: a storage process for storing logs of multiple components that constitute a virtualized environment of a network; a memory process for storing correspondence information that associates, for each failure that may occur in the network, error logs of the multiple components related to the failure; a search process for, when a failure occurs in the network, searching the stored logs for an error log of a second component related to the failure using the correspondence information based on the error log of a first component of the multiple components; and a presentation process for presenting the results of the search to a user.

[0086] [2] The network management device described in [1], characterized in that the storage process stores the correspondence information that associates keywords that represent the contents of the error logs of the multiple components related to the failure, and the search process searches for keywords that represent the contents of the error log of the second component by referring to the correspondence information based on the keywords that represent the contents of the error log of the first component, and searches for the error log of the second component from the stored logs based on the searched keywords.

[0087] [3] The network management device described in [2] is characterized in that the presentation process presents, as a result of the search, information including keywords representing the contents of the error log of the first component and keywords representing the contents of the error log of the second component searched by referring to the correspondence information.

[0088] [4] A network management device according to any one of [1] to [3], characterized in that, in the search process, if an error log of the second component is searched from the stored logs, the presentation process presents information including a link to the searched error log of the second component as a result of the search.

[0089] [5] A network management device according to any one of [1] to [4], characterized in that the plurality of components include a virtualization platform and an application on the virtualization platform.

[0090] [6] The network management device described in [5], wherein the first component is the application, the second component is the virtualization platform, and the presentation process, when an error log of the virtualization platform related to the failure is searched from the stored logs in the search process, presents the virtualization platform as a possible cause of the failure.

[0091] [7] A network management device according to [5] or [6], characterized in that the first component is the application, the second component is the virtualization platform, and the presentation process presents the application as the cause of the failure if the search process does not find an error log of the virtualization platform related to the failure from the stored logs.

[0092] [8] A network management device according to any one of [5] to [7], characterized in that the first component is the virtualization platform, the second component is the application, and the presentation process presents the application as being within the scope of impact of the failure if an error log of the application related to the failure is searched for from the stored log in the search process.

[0093] [9] A network management method characterized by storing logs of multiple components that constitute a virtualized environment of a network, storing correspondence information that associates, for each failure that may occur in the network, the error logs of the multiple components related to the failure, and when a failure occurs in the network, using the correspondence information based on the error log of a first component of the multiple components, searching the stored logs for the error log of a second component related to the failure, and presenting the results of the search to a user.

[0094]

[10] A network management system comprising one or more processors, wherein at least one of the one or more processors executes the following: a storage process for storing logs of multiple components that constitute a virtualized environment of a network; a memory process for storing correspondence information that associates, for each failure that may occur in the network, error logs of the multiple components related to the failure; a search process for, when a failure occurs in the network, searching the stored logs for an error log of a second component related to the failure using the correspondence information based on the error log of a first component of the multiple components; and a presentation process for presenting the results of the search to a user. [Explanation of symbols]

[0095] 11...base station, 12...edge data center, 13...regional data center, 14...central data center, 100...mobile network, 110...NFVI, 120...VNF, 130...MANO, 131...NFVO, 132...VNFM, 133...VIM, 140...OSS / BSS, 150...log management unit, 151...log dictionary management unit, 152...log storage management unit, 153...search result presentation unit, 150a...log dictionary database (DB), 150b...log storage database (DB)

Claims

1. one or more processors; by at least one of the one or more processors, A storage process for storing logs of multiple components that make up the network virtualization environment; a storage process for storing correspondence information in which, for each failure that may occur in the network, error logs of the plurality of components related to the failure are associated with the failure; a search process for searching, when a failure occurs in the network, an error log of a second component related to the failure from the stored logs using the correspondence information based on an error log of a first component among the plurality of components; a presentation process for presenting the search results to a user; is executed, A network management device comprising:

2. the storage process stores the correspondence information in which keywords representing the contents of the error logs of the plurality of components related to the failure are associated with each other; the search process refers to the correspondence information and searches for keywords that represent the content of the error log of the second component based on keywords that represent the content of the error log of the first component, and searches the stored logs for the error log of the second component based on the searched keywords; 2. The network management device according to claim 1.

3. The presentation process includes: presenting, as a result of the search, information including a keyword representing the content of the error log of the first component and a keyword representing the content of the error log of the second component searched with reference to the correspondence information; 3. The network management device according to claim 2.

4. The presentation process includes: If the error log of the second component is searched for from the stored logs in the search process, information including a link to the searched error log of the second component is presented as a result of the search.

2. The network management device according to claim 1.

5. The plurality of components include a virtualization platform and an application on the virtualization platform.

2. The network management device according to claim 1.

6. the first component is the application; the second component is the virtualization infrastructure; The presentation process includes: In the search process, if an error log of the virtualization platform related to the failure is found from the stored logs, the virtualization platform is presented as a candidate for the cause of the failure.

6. The network management device according to claim 5.

7. the first component is the application; the second component is the virtualization infrastructure; The presentation process includes: If an error log of the virtualization platform related to the failure is not found from the stored logs in the search process, the application is presented as the cause of the failure.

6. The network management device according to claim 5.

8. the first component is the virtualization infrastructure; the second component is the application; The presentation process includes: If an error log of the application related to the failure is found from the stored logs in the search process, the application is presented as being within the scope of the impact of the failure.

6. The network management device according to claim 5.

9. It stores logs of multiple components that make up the network virtualization environment, storing correspondence information that associates, for each failure that may occur in the network, error logs of the plurality of components related to the failure; When a failure occurs in the network, based on an error log of a first component among the plurality of components, search the stored logs for an error log of a second component related to the failure using the correspondence information; presenting the results of said search to the user; A network management method comprising:

10. one or more processors; by at least one of the one or more processors, A storage process for storing logs of multiple components that make up the network virtualization environment; a storage process for storing correspondence information in which, for each failure that may occur in the network, error logs of the plurality of components related to the failure are associated with the failure; a search process for searching, when a failure occurs in the network, an error log of a second component related to the failure from the stored logs using the correspondence information based on an error log of a first component among the plurality of components; a presentation process for presenting the search results to a user; is executed, A network management system comprising:

Citation Information

Patent Citations

  • Virtualization management / orchestration apparatus, virtualization management / orchestration method, and program

    WO2016121802A1